EDBT 2026 Demo / reviewers in the wild / expert
Robert K. Montoye
dblp:91/5357
· DBLP profile ↗
13ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Hardware accelerators and domain-specific architectures · 38% Processor architecture and microarchitecture · 32% Energy-efficient computing · 22% | |
| Computer networks
1 paper |
Cellular and mobile networks · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.5 | 1 | 2021 | Energy Efficiency Boost in the AI-Infused POWER10 Processor · ISCA 2021 |
Processor architecture and microarchitecture › microprocessor design
processor core design |
0.5 | 1 | 2021 | Energy Efficiency Boost in the AI-Infused POWER10 Processor · ISCA 2021 |
Cellular and mobile networks › mobile networks › mobile network architecture › cellular network architecture
base station architecture |
0.1 | 1 | 2012 | Architectural perspectives of future wireless base stations based on the IBM PowerEN™ processor · HPCA 2012 |
Cellular and mobile networks
radio access networks |
0.1 | 1 | 2012 | Architectural perspectives of future wireless base stations based on the IBM PowerEN™ processor · HPCA 2012 |
Energy-efficient computing
power-performance tradeoff |
0.1 | 1 | 2010 | Practical Strategies for Power-Efficient Computing Technologies · Proc. IEEE 2010 |
Energy-efficient computing
voltage scaling |
0.1 | 1 | 2010 | Practical Strategies for Power-Efficient Computing Technologies · Proc. IEEE 2010 |
Electronic design automation
logic synthesis |
0.1 | 2 | 2008 | Custom is from Venus and synthesis from Mars · DAC 2008 Area-time efficient addition in charge based technology · DAC 1981 |
Processor architecture and microarchitecture
multithreading |
0.0 | 1 | 2012 | Architectural perspectives of future wireless base stations based on the IBM PowerEN™ processor · HPCA 2012 |
Memory systems
cache |
0.0 | 1 | 2010 | Practical Strategies for Power-Efficient Computing Technologies · Proc. IEEE 2010 |
Parallel and multicore computing
parallel algorithms |
0.0 | 1 | 1982 | A Practical Algorithm for the Solution of Triangular Systems on a Parallel Processing System · IEEE Trans. Computers 1982 |
Processor architecture and microarchitecture
arithmetic unit |
0.0 | 1 | 1981 | Area-time efficient addition in charge based technology · DAC 1981 |
Integrated circuit design
digital circuit design |
0.0 | 1 | 1981 | Area-time efficient addition in charge based technology · DAC 1981 |
Processor architecture and microarchitecture › SIMD
SIMD machine |
0.0 | 1 | 1982 | A Practical Algorithm for the Solution of Triangular Systems on a Parallel Processing System · IEEE Trans. Computers 1982 |
Electronic design automation › design optimization
area-time optimization |
0.0 | 1 | 1981 | Area-time efficient addition in charge based technology · DAC 1981 |
Methods — techniques the papers use, named apart from their topics
pre-silicon methodology · 0.5reconfigurable computing · 0.3SIMD · 0.3device-circuit-system codesign · 0.1charge-based technology · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Energy Efficiency Boost in the AI-Infused POWER10 ProcessorabstractWe present the novel micro-architectural features, supported by an innovative and novel pre-silicon methodology in the design of POWER10. The resulting projected energy efficiency boost over POWER9 is 2.6x at core level (for SPECint) and up to 3x at socket level. In addition, a new feature supporting inline AI acceleration was added to the POWER ISA and incorporated into the POWER10 processor core design. The resulting boost in SIMD/AI socket performance is projected to be up to 10x for FP32 and 21x for INT8 models of ResNet-50 and BERT-Large. In this paper, we describe the novel methodology deployed and used not only to obtain these efficiency boosts for traditional workloads, but also to infuse AI/ML/HPC capability directly into the POWER10 core. Brian W. Thompto, Dung Q. Nguyen, José E. Moreira, Ramon Bertran Monfort, Hans M. Jacobson, Richard J. Eickemeyer, Rahul M. Rao, Michael Goulet, Marcy Byers, Christopher J. Gonzalez, Karthik Swaminathan, Nagu R. Dhanwada, Silvia M. Müller, Satish Kumar Sadasivam, Robert K. Montoye, William J. Starke, Christian G. Zoellin, Michael S. Floyd, Jeffrey Stuecheli, Nandhini Chandramoorthy, John-David Wellman, Alper Buyuktosunoglu, Matthias Pflanz, Balaram Sinharoy, Pradip Bose |
ISCA | 16 |
| 2017 | Very Low Voltage (VLV) DesignabstractThis paper is a tutorial-style introduction to a special session on: Effective Voltage Scaling in the Late CMOS Era. It covers the fundamental challenges and associated solution strategies in pursuing very low voltage (VLV) designs. We discuss the performance and system reliability constraints that are key impediments to VLV. The associated trade-offs across power, performance and reliability are helpful in inferring the optimal operational voltage-frequency point. This work was performed under the auspices of an ongoing DARPA program (named PERFECT) that is focused on maximizing system-level energy efficiency. Ramon Bertran Monfort, Pradip Bose, David Brooks 0001, Jeff Burns, Alper Buyuktosunoglu, Nandhini Chandramoorthy, Eric Cheng, Martin Cochet, Schuyler Eldridge, Daniel J. Friedman, Hans M. Jacobson, Rajiv V. Joshi, Subhasish Mitra, Robert K. Montoye, Arun Paidimarri, Pritish Parida, Kevin Skadron, Mircea R. Stan, Karthik Swaminathan, Augusto Vega, Swagath Venkataramani, Christos Vezyrtzis, Gu-Yeon Wei, John-David Wellman, Matthew M. Ziegler |
ICCD | 14 |
| 2014 | Matrix-matrix multiplication on a large register file architecture with indirectionabstractDense matrix-matrix multiply is an important kernel in many high performance computing applications including the emerging deep neural network based cognitive computing applications. Graphical processing units (GPU) have been very successful in handling dense matrix-matrix multiply in a variety of applications. However, recent research has shown that GPUs are very inefficient in using the available compute resources on the silicon for matrix multiply in terms of utilization of peak floating point operations per second (FLOPS). In this paper, we show that an architecture with a large register file supported by “indirection ” can utilize the floating point computing resources on the processor much more efficiently. A key feature of our proposed in-line accelerator is a bank-based very-large register file, with embedded SIMD support. This processor-in-regfile (PIR) strategy is implemented as local computation elements (LCEs) attached to each bank, overcoming the limited number of register file ports. Because each LCE is a SIMD computation element, and all of them can proceed concurrently, the PIR approach constitutes a highly-parallel super-wide-SIMD device. We show that we can achieve more than 25% better performance than the best known results for matrix multiply using GPUs. This is achieved using far lesser floating point computing units and hence lesser silicon area and power. We also show that architecture blends well with the Strassen and Winograd matrix multiply algorithms. We optimize the selective data parallelism that the LCEs enable for these algorithms and study the area-performance trade-offs. Dheeraj Sreedhar, Jeff H. Derby, Robert K. Montoye, Charles L. Johnson |
HiPC | 3 |
| 2013 | Processor architecture for software implementation of multi-sector G-RAKE receivers for HSUPA wireless infrastructureabstractThe high speed uplink packet access (HSUPA) wireless standard requires extremely high-performance signal processing in the baseband receiver, the most challenging being the chip rate rake receiver. In this paper we describe the architectural enhancements on the IBM's PowerEN processor, to enable it to support the computational requirements of the rake receiver in a fully programmable and scalable fashion. A key feature of these enhancements is a bank-based very-large register file, with embedded single instruction multiple data (SIMD) support. This processor-in-regfile (PIR) strategy is implemented as local computation elements (LCEs) attached to each bank. This overcomes the limitation on the number of register file ports and at the same time enables high degree of parallelism. We show that these enhancements enable the integration of multi-sector HSUPA G-RAKE receivers on a single processor. Dheeraj Sreedhar, Jeff H. Derby, Augusto Vega, Brian Rogers, Charles L. Johnson, Robert K. Montoye |
ICASSP | 6 |
| 2012 | Architectural perspectives of future wireless base stations based on the IBM PowerEN™ processorabstractIn wireless networks, base stations are responsible for operating on large amounts of traffic at high speed rates. With the advent of new standards, as 4G, further pressure is put in the hardware requirements to satisfy speeds of up to 1 Gbps. In this work, we study the applicability and potential benefits of the IBM PowerEN processor (a multi-core, massively multithreaded platform) in the realm of base stations for the 3G and 4G standards. The approach involves exploiting the throughput computation capabilities of the PowerEN processor, replacing the bus-attached special-function accelerators with a layer of in-line universal acceleration support, incorporated within the cores. A key feature of this in-line accelerator is a bank-based very-large register file, with embedded SIMD support. This processor-in-regfile (PIR) strategy is implemented as local computation elements (LCEs) attached to each bank, overcoming the limited number of register file ports. Because each LCE is a SIMD computation element, and all of them can proceed concurrently, the PIR approach constitutes a highly-parallel super-wide-SIMD device. To target a broad spectrum of applications for base stations, we also consider a PIR-based architecture built upon reconfigurable LCEs. In this paper, we evaluate the in-line universal accelerator and the PIR strategy focusing on two specific applications for base stations: FFT and Turbo Decoding. Augusto Vega, Pradip Bose, Alper Buyuktosunoglu, Jeff H. Derby, Michele Franceschini, Robert K. Montoye |
HPCA | 7 |
| 2010 | Practical Strategies for Power-Efficient Computing TechnologiesabstractAfter decades of continuous scaling, further advancement of silicon microelectronics across the entire spectrum of computing applications is today limited by power dissipation. While the trade-off between power and performance is well-recognized, most recent studies focus on the extreme ends of this balance. By concentrating instead on an intermediate range, an ~ 8× improvement in power efficiency can be attained without system performance loss in parallelizable applications-those in which such efficiency is most critical. It is argued that power-efficient hardware is fundamentally limited by voltage scaling, which can be achieved only by blurring the boundaries between devices, circuits, and systems and cannot be realized by addressing any one area alone. By simultaneously considering all three perspectives, the major issues involved in improving power efficiency in light of performance and area constraints are identified. Solutions for the critical elements of a practical computing system are discussed, including the underlying logic device, associated cache memory, off-chip interconnect, and power delivery system. The IBM Blue Gene system is then presented as a case study to exemplify several proposed directions. Going forward, further power reduction may demand radical changes in device technologies and computer architecture; hence, a few such promising methods are briefly considered. Leland Chang, David J. Frank, Robert K. Montoye, Steven J. Koester, Brian L. Ji, Paul Coteus, Robert H. Dennard, Wilfried Haensch |
Proc. IEEE | 3 |
| 2008 | Custom is from Venus and synthesis from MarsabstractDue to ever increasing cost of doing design, design productivity and more specifically, cost of design has become a major bottleneck in large scale design projects. Due to the cost crunch, automated synthesis techniques have been gaining ground on manual and cost intensive (although potentially yielding higher quality of results) custom techniques. The push towards system level performance with multiple lower performance cores as opposed to single high performance core is also making the pursuit of frequency at any cost meaningless. Synthesis vs Custom has traditionally been a lively debate in microprocessor companies but now it is becoming more widespread with the desire of fabless companies to more fully utilize the technology node in order to compete with IDMs by utilizing more custom techniques. Custom hotshots argue that we cannot afford to leave performance or power on table when you are spending 3billion $s in fabs. Industry needs to get maximum benefits possible. Synthesis fanatics will argue that designers should get over the last few ps and last few nW, as time to market, cost, and system performance are the new drivers and synthesis can not only meet the challenge but beat it many times with lower power solutions. This panel will present this internal debate in a public forum. Ruchir Puri, William H. Joyner, Shekhar Borkar, Ty Garibay, Jonathan Lotz, Robert K. Montoye |
DAC | 6 |
| 2005 | Testing and debugging delay faults in dynamic circuitsabstractWe propose novel design for test and debug techniques to apply two patterns for delay fault test and debug in dynamic circuits. Dynamic circuits, which have traditionally been difficult to test, pose new challenges for AC tests due to the presence of a reset phase between applications of any two patterns, which impedes delay fault testing of such circuits. We present two sets of design for test and debug techniques. The first set facilitates application of two patterns to dynamic circuits in general, overcoming the issue of reset phase, and reduces the problem of test generation for dynamic circuits to test generation for pull down paths of static CMOS circuits. The second set enables application of two patterns to scan based dynamic circuits. The proposed techniques reduce the problem of delay test generation for scan based dynamic circuits to that of delay test generation for static CMOS circuits with complete accessibility to all primary inputs. The techniques have minimal area overhead and also provide significant reduction in power during scan operation Ramyanshu Datta, Sani R. Nassif, Robert K. Montoye, Jacob A. Abraham |
ITC | 3 |
| 2004 | The four degrees of 3DabstractNo abstract available. Robert K. Montoye |
ISPD | 1 |
| 2004 | The four degrees of 3DabstractNo abstract available. Robert K. Montoye |
ISPD | 1 |
| 1989 | IBM second-generation RISC machine organizationabstractA highly concurrent second-generation RISC (reduced-instruction-set computer) that combines a powerful RISC architecture with sophisticated hardware design techniques to achieve a short cycle time and a low cycles-per-instruction (CPI) ratio is described. Like earlier RISC processors, this design uses a register-oriented instruction set, the CPU is hardwired rather than microcoded, and it features a pipelined implementation. Unlike earlier RISC processors, however, several advanced architectural and implementation features are used, including separate instruction and data caches, zero-cycle branches, multiple-instruction dispatch, and simultaneous execution of fixed- and floating-point instructions. The CPU has a four-word data bus to main memory, a four-word instruction-fetch bus from the I-cache arrays, and a two-word data bus between the D-cache and floating-point unit. The CPU has a full 64-b floating-point engine, and thirty-two 64-b floating point registers in addition to thirty-two 32-b fixed-point registers. In a single cycle, four instructions can be executed simultaneously.> H. B. Bakoglu, Gregory F. Grohoski, L. E. Thatcher, James A. Kahle, Charles R. Moore, David P. Tuttle, Warren E. Maule, W. R. Hardell Jr., Dwain A. Hicks, M. Nguyenphu, Robert K. Montoye, W. T. Glover, Sudhir Dhawan |
ICCD | 11 |
| 1982 | A Practical Algorithm for the Solution of Triangular Systems on a Parallel Processing SystemabstractAn algorithm is presented for a more efficient and implementable solution of triangular systems on a parallel (SIMD) computer which requires 0(log (N)) fewer processing cycles than the best previous results, where N is the system size. We will also show that the data can be accessed and aligned in the same order of time using as many memory units as processors and Ω networks for data alignment. (Previous results dealing with this type of algorithm have not dealt in any detail with the problem of data access and alignment.) Robert K. Montoye, Duncan H. Lawrie |
IEEE Trans. Computers | 1 |
| 1981 | Area-time efficient addition in charge based technology
Robert K. Montoye |
DAC | 1 |