VLDB 2026 Research / reviewers in the wild / expert
Pong-Fei Lu
dblp:08/6179
· DBLP profile ↗
11ranked-venue papers
2as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Hardware accelerators and domain-specific architectures · 52% Emerging computing paradigms · 24% Electronic design automation · 11% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 2 | 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021 Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator |
0.5 | 1 | 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021 |
Emerging computing paradigms
approximate computing |
0.4 | 1 | 2020 | Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Hardware accelerators and domain-specific architectures
approximate computing accelerator |
0.4 | 1 | 2020 | Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Emerging computing paradigms › approximate computing
cross-layer approximate computing |
0.4 | 1 | 2020 | Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Electronic design automation
physical design |
0.2 | 2 | 2014 | Row Based Dual-VDD Island Generation and Placement · DAC 2014 A Semi-Custom Design Flow in High-Performance Microprocessor Design · DAC 2001 |
Integrated circuit design
low-power circuit design |
0.2 | 1 | 2014 | Row Based Dual-VDD Island Generation and Placement · DAC 2014 |
Electronic design automation › physical design
placement |
0.2 | 1 | 2014 | Row Based Dual-VDD Island Generation and Placement · DAC 2014 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.1 | 1 | 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021 |
Machine learning › Deep learning architectures and training
neural network inference |
0.1 | 1 | 2020 | Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Energy-efficient computing
power management |
0.1 | 1 | 2014 | Row Based Dual-VDD Island Generation and Placement · DAC 2014 |
Integrated circuit design › digital circuit design › CMOS circuit design
digital CMOS circuits |
0.0 | 1 | 1998 | SOI for digital CMOS VLSI: design considerations and advances · Proc. IEEE 1998 |
Integrated circuit design › semiconductor devices
silicon-on-insulator |
0.0 | 1 | 1998 | SOI for digital CMOS VLSI: design considerations and advances · Proc. IEEE 1998 |
Integrated circuit design › ASIC design
semi-custom design |
0.0 | 1 | 2001 | A Semi-Custom Design Flow in High-Performance Microprocessor Design · DAC 2001 |
Memory systems
DRAM |
0.0 | 1 | 1998 | SOI for digital CMOS VLSI: design considerations and advances · Proc. IEEE 1998 |
Memory systems › random-access memory
SRAM |
0.0 | 1 | 1998 | SOI for digital CMOS VLSI: design considerations and advances · Proc. IEEE 1998 |
Methods — techniques the papers use, named apart from their topics
quantization · 0.9pruning · 0.9mixed-precision arithmetic · 0.9custom number representation · 0.9performance modeling · 0.5island generation · 0.2physical design tools · 0.0automated tuning · 0.0device simulation · 0.0circuit design · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and InferenceabstractThe growing prevalence and computational demands of Artificial Intelligence (AI) workloads has led to widespread use of hardware accelerators in their execution. Scaling the performance of AI accelerators across generations is pivotal to their success in commercial deployments. The intrinsic error-resilient nature of AI workloads present a unique opportunity for performance/energy improvement through precision scaling. Motivated by the recent algorithmic advances in precision scaling for inference and training, we designed RaPiD1, a 4-core AI accelerator chip supporting a spectrum of precisions, namely, 16 and 8-bit floating-point and 4 and 2-bit fixed-point. The 36mm2RaPiD chip fabricated in 7nm EUV technology delivers a peak 3.5 TFLOPS/W in HFP8 mode and 16.5 TOPS/W in INT4 mode at nominal voltage. Using a performance model calibrated to within 1% of the measurement results, we evaluated DNN inference using 4-bit fixed-point representation for a 4-core 1 RaPiD chip system and DNN training using 8-bit floating point representation for a 768 TFLOPs AI system comprising 4 32-core RaPiD chips. Our results show INT4 inference for batch size of 1 achieves 3 - 13.5 (average 7) TOPS/W and FP8 training for a mini-batch of 512 achieves a sustained 102 - 588 (average 203) TFLOPS across a wide range of applications. Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang 0333, Sanchari Sen, Ankur Agrawal, Monodeep Kar, Shubham Jain 0004, Alberto Mannari, Hoang Tran, Eri Ogawa, Kazuaki Ishizaki, Hiroshi Inoue, Marcel Schaal, Mauricio J. Serrano, Jungwook Choi, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Allison Allain, James Bonanno, Nianzheng Cao, Robert Casatuta, Matthew Cohen, Bruce M. Fleischer, Michael Guillorn, Howard Haynie, Jinwook Jung, Mingu Kang, Kyu-Hyoun Kim, Siyu Koswatta, Sae Kyu Lee, Martin Lutz, Silvia M. Müller, Jinwook Oh, Ashish Ranjan 0001, Zhibin Ren, Scot Rider, Kerstin Schelm, Michael Scheuermann, Joel Silberman, Vidhi Zalani, Xin Zhang 0025, Ching Zhou, Matthew M. Ziegler, Vinay Shah, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Leland Chang, Kailash Gopalakrishnan |
ISCA | 50 |
| 2020 | Efficient AI System Design With Cross-Layer Approximate ComputingabstractAdvances in deep neural networks (DNNs) and the availability of massive real-world data have enabled superhuman levels of accuracy on many AI tasks and ushered the explosive growth of AI workloads across the spectrum of computing devices. However, their superior accuracy comes at a high computational cost, which necessitates approaches beyond traditional computing paradigms to improve their operational efficiency. Leveraging the application-level insight of error resilience, we demonstrate how approximate computing (AxC) can significantly boost the efficiency of AI platforms and play a pivotal role in the broader adoption of AI-based applications and services. To this end, we present RaPiD, a multi-tera operations per second (TOPS) AI hardware accelerator core (fabricated at 14-nm technology) that we built from the ground-up using AxC techniques across the stack including algorithms, architecture, programmability, and hardware. We highlight the workload-guided systematic explorations of AxC techniques for AI, including custom number representations, quantization/pruning methodologies, mixed-precision architecture design, instruction sets, and compiler technologies with quality programmability, employed in the RaPiD accelerator. Swagath Venkataramani, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Jungwook Choi, Mingu Kang, Ankur Agarwal, Jinwook Oh, Shubham Jain 0004, Tina Babinsky, Nianzheng Cao, Thomas W. Fox, Bruce M. Fleischer, George Gristede, Michael Guillorn, Howard Haynie, Hiroshi Inoue, Kazuaki Ishizaki, Michael J. Klaiber, Shih-Hsien Lo, Gary W. Maier, Silvia M. Müller, Michael Scheuermann, Eri Ogawa, Marcel Schaal, Mauricio J. Serrano, Joel Silberman, Christos Vezyrtzis, Wei Wang 0333, Fanchieh Yee, Matthew M. Ziegler, Ching Zhou, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Vijayalakshmi Srinivasan, Leland Chang, Kailash Gopalakrishnan |
Proc. IEEE | 35 |
| 2018 | Across the Stack Opportunities for Deep Learning AccelerationabstractThe combination of growth in compute capabilities and availability of large datasets has led to a re-birth of deep learning. Deep Neural Networks (DNNs) have become state-of-the-art in a variety of machine learning tasks spanning domains across vision, speech, and machine translation. Deep Learning (DL) achieves high accuracy in these tasks at the expense of 100s of ExaOps of computation; posing significant challenges to efficient large-scale deployment in both resource-constrained environments and data centers. Vijayalakshmi Srinivasan, Bruce M. Fleischer, Sunil Shukla, Matthew M. Ziegler, Joel Silberman, Jinwook Oh, Jungwook Choi, Silvia M. Müller, Ankur Agrawal, Tina Babinsky, Nianzheng Cao, Chia-Yu Chen, Pierce Chuang, Thomas W. Fox, George Gristede, Michael Guillorn, Howard Haynie, Michael J. Klaiber, Dongsoo Lee, Shih-Hsien Lo, Gary W. Maier, Michael Scheuermann, Swagath Venkataramani, Christos Vezyrtzis, Naigang Wang, Fanchieh Yee, Ching Zhou, Pong-Fei Lu, Brian W. Curran, Leland Chang, Kailash Gopalakrishnan |
ISLPED | 28 |
| 2016 | Synthesis design strategies for energy-efficient microprocessorsabstractA detailed synthesis study has been performed on a functional unit from a recent IBM microprocessor to explore the voltage-frequency space for energy-efficient design points across a wide performance spectrum ranging from 625 MHz at 0.48V to 5.6 GHz at 0.95V. It is found that the optimal operating voltage depends strongly on frequency for an energy-efficient design. Circuit characteristics, as represented by the combination of the average gate width, effective VT, and buffering scheme, differ significantly between designs optimized for low voltage-frequency and for high voltage-frequency operations and suggest a distinct application dependence in the selection of standard cell images and optimal design points. In particular, for optimal energy efficiency at a given frequency, low voltage designs should utilize smaller gate width and lower VT. Though a design energy-optimized near 1V is more scalable over a wide frequency range when operating at low voltages, designs optimized at a lower voltage-frequency point can be leveraged to offer better solutions in both performance and energy efficiency within a narrow frequency range near the design point. Ching Zhou, Yu-Shiang Lin, Pong-Fei Lu, Bruce M. Fleischer, David J. Frank, Leland Chang |
ICCD | 3 |
| 2015 | A Ring-Oscillator-Based Reliability Monitor for Isolated Measurement of NBTI and PBTI in High-k/Metal Gate TechnologyabstractRing-oscillator-based test structures that can separately measure the negative bias temperature instability (NBTI) and positive bias temperature instability (PBTI) degradation effects in digital circuits are presented for high-k metal gate devices. The mathematical derivation also shows that the structure for frequency degradation measurement can directly be used for estimating the portion of the NBTI and PBTI in the conventional ring oscillator. The proposed test structures including frequency degradation sensing circuitry have been implemented in an experimental high-k/metal gate SoI process. Tony Tae-Hyoung Kim, Pong-Fei Lu, Keith A. Jenkins, Chris H. Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Row Based Dual-VDD Island Generation and PlacementabstractPower consumption has become a major consideration in nanometer chip design. Since the dynamic power is proportional to V dd2, and the static power is proportional to V dd, lowering power supply voltage is an efficient method to reduce the power usage. Hua Xiang 0001, Haifeng Qian, Ching Zhou, Yu-Shiang Lin, Fanchieh Yee, Andrew Sullivan, Pong-Fei Lu |
DAC | 7 |
| 2012 | Design of ring oscillator structures for measuring isolated NBTI and PBTIabstractRing oscillator based test structures that can separately measure the NBTI and PBTI degradation effects in digital circuits are presented for high-k metal-gate devices. The proposed test structures enable simultaneous stress of all devices under test in either NBTI or PBTI mode and measure frequency or threshold voltage shifts. The mathematical derivation also shows that the structure for frequency degradation measurement can directly be used for estimating the portion of the NBTI and PBTI in the conventional ring oscillator. The proposed test structures including beat frequency sensing circuitry have been designed in a 0.9V, 45nm SOI technology. Tony Tae-Hyoung Kim, Pong-Fei Lu, Chris H. Kim |
ISCAS | 2 |
| 2006 | A pulsed low-voltage swing latch for reduced power dissipation in high-frequency microprocessorsabstractWe have reported previously [1] a low-swing latch (LSL) with superior performance-power tradeoff compared to the conventional pass-gate master-slave latch. In this paper, hardware results are presented for the proposed LSL with pulsed clock waveforms. The motivation is to combine low-voltage swing with pulsed signals to further reduce overall system power in high-frequency microprocessors. We have designed a 65-bit accumulator loop experiment to mimic a microprocessor pipeline stage. The local clock buffer design features a mode switch to toggle between two-phase (c1/c2) master-slave clocking and one-phase pulsed (c2 only) clocking. Our data show that 15-25% system power saving can be achieved in pulsed mode compared to non-pulsed mode. Power contribution from individual components is also presented. Pong-Fei Lu, Nianzheng Cao, Leon J. Sigal, Pieter Woltgens, Raphael Robertazzi, David F. Heidel |
ISLPED | 1 |
| 2001 | A Semi-Custom Design Flow in High-Performance Microprocessor DesignabstractIn this paper we present techniques shown to significantly enhance the custom circuit design process typical of high-performance microprocessors. This methodology combines flexible custom circuit design with automated tuning and physical design tools to provide new opportunities to optimized design throughout the development cycle. Gregory A. Northrop, Pong-Fei Lu |
DAC | 2 |
| 1998 | SOI for digital CMOS VLSI: design considerations and advancesabstractThis paper reviews the recent advances of silicon-on-insulator (SOI) technology for complementary metal-oxide-semiconductor (CMOS) very-large-scale-integration memory and logic applications. Static random access memories (SRAMs), dynamic random access memories (DRAMs), and digital CMOS logic circuits are considered. Particular emphases are placed on the design issues and advantages resulting from the unique SOI device structure. The impact of floating-body in partially depleted devices on the circuit operation, stability, and functionality are addressed. The use of smart-body contact to improve the power and delay performance is discussed, as are global design issues. Ching-Te Chuang, Pong-Fei Lu, Carl J. Anderson |
Proc. IEEE | 2 |
| 1996 | Floating body effects in partially-depleted SOI CMOS circuitsabstractThis paper presents a detailed study on the impact of floating body in partially-depleted (PD) SOI MOSFET on various digital VLSI CMOS circuit families. The parasitic bipolar effect resulting from the floating body is shown to degrade the circuit noise margin and stability in general. In certain dynamic circuits and wide multiplexers, the parasitic bipolar effect is shown to cause logic state error if not properly accounted for. Pong-Fei Lu, Jin Ji, Ching-Te Chuang, Lawrence F. Wagner, Chang-Ming Hsieh, Jente B. Kuang, L. Hsu, Mario M. Pelella Jr., Shao-Fu Sanford Chu, Carl J. Anderson |
ISLPED | 1 |