Pingqiang Zhou

dblp:54/4046 · DBLP profile ↗
← Back
58ranked-venue papers
9as first author
37since 2021 · last 2026
0000-0001-9515-9302ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 54 · 8 first-author · 35 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2026 RL-Guided Thermal-Aware Quantization for Efficient and Robust ReRAM CIM Systems
abstract
Resistive RAM (ReRAM)-based Computing-inMemory (CIM) systems present significant advantages in energy efficiency and computational throughput for neural network acceleration. However, their performance is highly constrained by thermal-induced conductance drift, especially under aggressive quantization strategies. This work presents a reinforcement learning (RL)-guided thermal-aware layer-wise quantization framework optimized for ReRAM-based CIM systems. The proposed method encourages sparse and low-magnitude weight representations, while adaptively exploring layer-wise bit-width configurations guided by direct hardware evaluation feedback and thermal-aware reward. Experiments on CIFAR-10 and ImageNet benchmarks with ResNet and VGG show that the proposed method achieves up to 10.2% peak temperature reduction, and an average top- 1 accuracy improvements of 46.37% over fixed 8bit baselines. Compared to prior layer-wise quantization methods without thermal considerations, our approach improves accuracy by $\mathbf{6. 3 6 \% - 1 5. 0 6 \%}$ on average.
Lihua An, Pingqiang Zhou
ASP-DAC3
2026 GNN-Based Timing Yield Prediction From Statistical Static Timing Analysis
Chenbo Xi, Biwei Xie, Pingqiang Zhou
ASP-DAC4
2026 CNN-Assisted Low-Power Clock Tree Synthesis for 3D ICs
abstract
In this work, a convolutional neural network (CNN)-assisted low-power clock topology generation method for 3D clock tree synthesis (CTS) is proposed. Our approach considers both local and global costs in each merging step to prevent getting stuck in local optima. To explore the trade-off between local and global costs, we use CNN to set two important weighting factors to obtain an optimized clock tree that reduces power consumption. Compared with the conventional NNG-based method, experimental results on ISPD09 benchmarks show that our approach can reduce wirelength by 5.83% and power consumption by 4.73% on average. We also demonstrate the transferability of our method on larger-scale ISPD10 benchmarks.
Chenbo Xi, Jindong Zhou, Pingqiang Zhou
ASP-DAC3
2026 Efficient RF Passive Components Modeling with Bayesian Online Learning and Uncertainty Aware Sampling
abstract
Conventional radio frequency (RF) passive components modeling based on machine learning requires extensive electromagnetic (EM) simulations to cover geometric and frequency design spaces, creating computational bottlenecks. In this paper, we introduce an uncertainty-aware Bayesian online learning framework for efficient parametric modeling of RF passive components, which includes: 1) a Bayesian neural network with reconfigurable heads for joint geometric-frequency domain modeling while quantifying uncertainty; 2) an adaptive sampling strategy that simultaneously optimizes training data sampling across geometric parameters and frequency domain using uncertainty guidance. Validated on three RF passive components, the framework achieves accurate modeling while using only 2.86% EM simulation time compared to traditional ML-based flow, achieving a $35 \times$ speedup.
Huifan Zhang, Pingqiang Zhou
ASP-DAC2
2026 A High-Performance Neural Rendering Accelerator Based on Novel Multi-Level Ray Scheduling and Dual-Process Backend
abstract
Neural rendering enables photorealistic scene re-construction but remains difficult to deploy on edge devices due to intensive computation, redundant sampling, and memory bandwidth constraints. This work presents a high-performance neural rendering accelerator for real-time embedded rendering. The proposed design integrates: (1) a dual-process backend with fused micro-MLPs to significantly improve sample processing efficiency, (2) multi-resolution spatial partitioning with adaptive ray clustering to exploit sparsity and achieve over 95% cache hit rate, and (3) a multi-level scheduling framework with proactive prefetching to reduce MLP stalls. Implemented on FPGA, the prototype achieves 94.7 FPS at 800×800 resolution with 6.4 W power consumption. An ASIC implementation in 28 nm technology sustains 440 FPS at 268 mW. Experimental results demonstrate state-of-the-art performance and energy efficiency while preserving rendering quality above 30 dB PSNR.
Wenkai Zhou, Yuefeng Zhang, Binzhe Yuan, Junsheng Chen, Luntian Zhang, Xiangyu Zhang 0002, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001
DATE8
2026 RapidPnR: Accelerating the physical design for FPGAs via design-level parallelism
Wanzheng Weng, Pingqiang Zhou
Integr.2
2026 Few-shot learning GNN-EQL model with gm/ID method for analog integrated circuit design
Hongjian Zhou, Pingqiang Zhou
Integr.3
2026 An Energy-Efficient Edge Coprocessor for Neural Rendering With Explicit Data Reuse Strategies
abstract
Neural radiance fields (NeRFs) have transformed 3-D reconstruction and rendering, facilitating photorealistic image synthesis from sparse viewpoints. This work introduces an explicit data reuse neural rendering (EDR-NR) architecture, which reduces frequent external memory accesses (EMAs) and cache misses by exploiting the spatial locality from three phases, including rays, ray packets (RPs), and samples. The EDR-NR architecture features a four-stage scheduler that clusters rays on the basis of$Z$-order, prioritize lagging rays when ray divergence happens, reorders RPs based on spatial proximity, and issues samples out-of-orderly (OoO) according to the availability of on-chip feature data. In addition, a four-tier hierarchical RP marching (HRM) technique is integrated with an axis-aligned bounding box (AABB) to facilitate spatial skipping (SS), reducing redundant computations and improving throughput. Moreover, a balanced allocation strategy for feature storage is proposed to mitigate SRAM bank conflicts. Fabricated using a 40-nm process with a die area of 10.5 mm2, the EDR-NR chip demonstrates a$2.41\times $enhancement in normalized energy efficiency, a$1.21\times $improvement in normalized area efficiency, a$1.20\times $increase in normalized throughput, and a 53.42% reduction in on-chip SRAM consumption compared with state-of-the-art accelerators.
Binzhe Yuan, Xiangyu Zhang 0002, Yuefeng Zhang, Haochuan Wan, Zhechen Yuan, Junsheng Chen, Yunxiang He, Junran Ding, Chaolin Rao, Wenyan Su, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001
IEEE Trans. Very Large Scale Integr. Syst.13
2025 Time-Domain 3D Electromagnetic Fields Estimation Based on Physics-Informed Deep Learning Framework
abstract
Electromagnetic simulation is important and time-consuming in RF/microwave circuit design. Physics-informed deep learning is a promising method to learn a family of parametric partial differential equations. In this work, we propose a physics-informed deep learning framework to estimate time-domain 3D electromagnetic fields. Our method leverages physics-informed loss functions to model Maxwell's equations which govern electromagnetic fields. Our post-trained model produces accurate results with over 200 × speedup over the FDTD simulation. We reduce the mean square error by at least 14% and 15%, with respect to purely data-driven learning and the Fourier operator learning method FNO. In order to optimize data and physical loss simultaneously, we introduce a self-adaptive scaling factors updating algorithm, which has 8.4% less error than the loss balancing method ReLoBRaLo. Our codes are open-source at this link.
Huifan Zhang, Pingqiang Zhou
DATE3
2025 RapidPnR: Accelerating the Physical Design for FPGAs via Design-Level Parallelism
abstract
The runtime of physical design has become a critical issue for FPGA development as the scale and complexity of circuit designs surge with the increasing logic capacity of FPGA devices. The time-consuming process of physical design significantly extends the cycle of design iteration, which heavily impacts the efficiency of debugging and architecture optimization. To address this issue, this work proposes a generic, fully-automated and split-and-parallel physical design flow to accelerate the deployment of large-scale circuits on FPGAs. Specifically, our flow automatically partitions the synthesized netlist into multiple smaller pieces, performs parallel physical design of each piece, and then merges them into the complete design. Evaluated on a set of real circuit benchmarks, our flow reduces the runtime by more than 50% and ensures nearly the same design frequency compared to the physical design flow provided by the commercial tool Vivado.
Wanzheng Weng, Pingqiang Zhou
FCCM2
2025 Defending Side-Channel Attacks in Convolutional Neural Networks with Channel-Level Parallelization
abstract
Side-channel attacks (SCAs) pose significant threats to the security of neural networks (NNs) deployed on hardware platforms, especially in cloud Field-Programmable Gate Array (FPGA) environments. This paper presents a novel approach to enhance the security of convolutional layers in NNs against SCAs by introducing a channel-level parallel structure. Compared with the original structure and the state-of-the-art masking technique, the channel-level parallel structure significantly reduces the success rate of SCAs (from 97.13% to 5.46% on average) and is able to be optimized for either low resource overhead (83.64% reduction) or good timing performance (83.01% improvement).
Yankun Zhu, Ranxi Lin, Pingqiang Zhou
FCCM3
2025 Clock-Wirelength-Driven Detailed Placement
Ziang Ge, Yikai Liu, Jindong Zhou, Pingqiang Zhou
ACM Great Lakes Symposium on VLSI4
2025 Inductance-aware Clock Network Synthesis Considering Hierarchical Interconnects in 3D ICs
Jindong Zhou, Ziang Ge, Chenbo Xi, Pingqiang Zhou
ACM Great Lakes Symposium on VLSI4
2025 A thermal-aware layer-wise quantization framework for ReRAM-Based DNN CIM systems
Lihua An, Pingqiang Zhou
Integr.2
2025 A Neural Rendering Coprocessor With Optimized Ray Representation and Marching
abstract
Neural rendering, a transformative approach for 3-D scene reconstruction and rendering, has advanced rapidly in recent years. This article introduces an energy-efficient neural rendering coprocessor that implements the popular and widely used instant neural graphics primitive (Instant-NGP) algorithm. In particular, we address the challenges of limited resources for deploying Instant-NGP on edge by proposing a dedicated architecture, which incorporates three main innovations: 1) we optimize occupancy grid queries in the ray marching module by partitioning the grid and decoupling the query process from sampling point generation, which improves both efficiency and memory usage; 2) we introduce a bilinked list-based ray switching strategy, which ensures continuous pipeline utilization to overcome the inefficiencies caused by sequential processing; and 3) we optimize the hash encoding process by incorporating quantization-aware training (QAT), enabling the hash table to fit into on-chip memory, thereby improving performance on resource-constrained devices. To demonstrate the effectiveness of our architecture, we design and fabricate a proof-of-concept chip using 40-nm CMOS technology and develop a testing system to evaluate its performance. Measurement results validate the advantages of the proposed design, showing that our chip achieves superior energy efficiency compared to both server and edge graphics processing units (GPUs), as well as other state-of-the-art neural rendering chip designs.
Zhechen Yuan, Binzhe Yuan, Chaolin Rao, Yiren Zhu, Yunxiang He, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2024 A Transferable GNN-based Multi-Corner Performance Variability Modeling for Analog ICs
abstract
Performance variability appears strong-nonlinear in analog ICs due to large process variations in advanced technologies. To capture such variability, a vast amount of data is required for learning-based accurate models. On the other hand, yield estimation across multiple PVT corners exacerbates data dimensionality further. In this paper, we propose a graph neural network (GNN)-based performance variability modeling method. The key idea is to leverage GNN techniques to extract variations-related local mismatch in analog circuits, and data efficiency is benefited by the ability of knowledge transfer among different PVT corners. Demonstrated upon three circuits in a commercial 65nm CMOS process and compared with the state-of-the-art modeling techniques, our method can achieve higher modeling accuracy while utilizing significantly less training data.
Hongjian Zhou, Pingqiang Zhou
ASPDAC4
2024 ZeroTetris: A Spacial Feature Similarity-based Sparse MLP Engine for Neural Volume Rendering
abstract
Neural Volume Rendering (NVR), a novel paradigm for the longstanding problem of photo-realistic rendering of virtual worlds, has developed explosively in the past three years. The unique and substantial computational requirements of NVR pose challenge on deploying NVR to existing dedicated accelerator for neural networks. In this work, we propose ZeroTetris, a spacial feature similarity-based sparse multilayer perceptron (MLP) hardware accelerator for NVR. By leveraging the unique similarity-based sparsity between adjacent sampling points in NVR models, ZeroTetris efficiently bypass the computation of zero activations, thereby enhancing energy efficiency. Evaluation results affirm the effectiveness of the proposed design, showcasing ZeroTetris's superior performance in both area and power efficiency compared to other dedicated sparse matrix multiplication or MLP accelerator designs.
Haochuan Wan, Linjie Ma, Antong Li, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001
DAC4
2024 LDL-SCA: Linearized Deep Learning Side-Channel Attack Targeting Multi-tenant FPGAs✱
abstract
In recent years, deep-learning side-channel attacks (DL-SCA) have gained increasing attention due to their enhanced efficacy against cryptographic modules. This paper explores that traditional non-profiled DL-SCA is unable to discern correct cipher keys in multi-tenant Field Programmable Gate Array (FPGA) scenarios due to the low correlation between power traces and cipher keys. To address this challenge, we propose Linearized Deep Learning Side-Channel Attack (LDL-SCA). Through modifying the output layer and integrating K-means clustering, LDL-SCA is capable of capturing linear features regardless of the low correlation between input and label. Moreover, we introduce new evaluation metrics derived from R2 and Cohen-kappa score. Our experiments show that LDL-SCA generate results with improved distinguishability, which has about ten times smaller standard deviation and ten times larger peak differences compared with Correlation Power Analysis (CPA) and Linear Regression Analysis (LRA).
Yankun Zhu, Siting Liu 0001, Liyu Yang, Pingqiang Zhou
ACM Great Lakes Symposium on VLSI4
2024 An Efficient Hardware Volume Renderer for Convolutional Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) has attracted growing attention in the fields of 3D reconstruction and rendering. However, straightforward NeRF algorithms encounter challenges in accurately capturing complex surface details with rich high-frequency information. A recent development known as Convolutional Neural Radiance Field (ConvNeRF) has demonstrated state-of-the-art results for these tasks. But it comes with substantial irregular computational requirements, particularly in the convolutional volume rendering phase. In this paper, we introduce a hardware accelerator designed to enhance the efficiency of convolutional volume rendering in ConvNeRF. Our approach includes the creation of specialized computation modules and corresponding on-chip memory system optimized for seamless support of gated convolutions and skip connections in ConvNeRF. To validate our design, we implement it in VerilogHDL and build a prototype using Field Programmable Gate Array (FPGA). We also map our design to 40nm CMOS technology. The evaluation results underscore the superiority of our accelerator in terms of energy efficiency when compared to an implementation on an NVIDIA 2080Ti GPU, offering approximately 84.6× more frames per watt.
Xuexin Wang, Yunxiang He, Xiangyu Zhang 0002, Pingqiang Zhou, Xin Lou 0001
ISCAS4
2024 Spiking-NeRF: Spiking Neural Network for Energy-Efficient Neural Rendering
abstract
Artificial Neural Networks (ANNs) have achieved remarkable performance in many artificial intelligence tasks. As the application scenarios become more sophisticated, the computation and energy consumption of ANNs are also constantly increasing, which poses a challenge for deploying ANNs on energy-constrained devices. Spiking Neural Networks (SNNs) provide a promising solution to build energy-efficiency neural networks. However, the current training methods of SNNs cannot output values as precise as ANNs. This limits the applications of SNNs to relatively simple image classification tasks. In this article, we extend the application of SNNs to neural rendering tasks and propose an energy-efficient spiking neural rendering model, called Spiking-NeRF (Spiking Neural Radiance Fields). We first analyze the ANN-to-SNN conversion theory and propose an output scheme for SNNs to obtain the precise scene property values. Then we customize the parameter normalization method for the special network architecture of neural rendering. Furthermore, we present an early termination strategy (ETS) based on the discrete nature of spikes to reduce energy consumption. We evaluate the performance of Spiking-NeRF on both realistic and synthetic scenes. Experimental results show that Spiking-NeRF can achieve comparable rendering performance to ANN-based NeRF with up to \(2.27\times\) energy reduction.
Ziwen Li 0004, Jindong Zhou, Pingqiang Zhou
ACM J. Emerg. Technol. Comput. Syst.4
2024 Ray Reordering for Hardware-Accelerated Neural Volume Rendering
abstract
Neural Volume Rendering (NVR) has advanced explosively since the advent of Neural Radiance Field (NeRF), a technique for novel view synthesis of complex scenes based on a finite set of input views. Existing ray casting-based NVR approaches process rays concurrently to leverage parallelism but fails to consider its impact on cache locality, which ultimately undermines the efficiency of corresponding dedicated hardware accelerator designs. We further observed that there exhibits spatial correspondence between features and voxels in NVR that can be exploited by processing in the order of voxel, not ray. This paper introduces a novel approach to meticulously reorder the execution of rays, ensuring that rays with similar memory access patterns are processed in parallel, thereby enhancing cache locality. On the basis of that, we also propose an efficient backend architecture and a corresponding memory subsystem, facilitating accurate data prefetching to hide off-chip memory latency. To validate the proposed architecture, we implement our design in VerilogHDL and evaluate the performance by post-synthesis simulation with real scene data. The evaluation results demonstrate that our design markedly enhances the efficiency of NVR processing, achieving a considerable speedup ($1.62\times $) compared to the state-of-the-art NVR accelerator, while necessitating significantly less silicon area ($5.12\times $) and power ($32.79\times $).
Junran Ding, Yunxiang He, Binzhe Yuan, Zhechen Yuan, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Detecting Adversarial Examples Utilizing Pixel Value Diversity
abstract
In this article, we introduce two novel methods to detect adversarial examples utilizing pixel value diversity. First, we propose the concept of pixel value diversity (which reflects the spread of pixel values in an image) and two independent metrics (UPVR and RPVR) to assess the pixel value diversity separately. Then we propose two methods to detect adversarial examples based on the threshold method and Bayesian method respectively. Experimental results show that compared to an excellent prior method LID, our proposed methods achieve better performances in detecting adversarial examples. We also show the robustness of our proposed work against an adaptive attack method.
Jinxin Dong, Pingqiang Zhou
ACM Trans. Design Autom. Electr. Syst.2
2024 Protecting Parallel Data Encryption in Multi-Tenant FPGAs by Exploring Simple but Effective Clocking Methodologies
abstract
Capitalizing on their versatility and high-performance attributes within heterogeneous designs, increasingly number of field-programmable gate arrays (FPGAs) are integrated into cloud data centers by cloud service providers (CSPs). While CSPs intend to reduce the cost by sharing one board among multiple users (called multi-tenant FPGA), hardware security problems such as side-channel attacks restrict it from spreading commercially. Existing research works have underscored the feasibility of remote side-channel attacks targeting a singular advanced encryption standard (AES) module on multi-tenant FPGAs, but they have not looked into the scenario of parallel data encryption on multiple AES modules for a single tenant, which is possible due to the small resource consumption of one AES module. In this work, we scrutinize correlation power analysis (CPA)-based side-channel attacks on parallel data encryption modules and develop two simple yet effective protective methods based on clocking methodologies—clocking phase shift and small frequency shift. The former technique adopts an identical clock frequency but with distinctive clocking phase to parallel encryption modules while the latter implements slightly different clock frequencies for parallel encryption modules. Experimental results show that both the methods can effectively increase the minimum required power traces for successful CPA, thus instituting a natural protective barrier for parallel data encryption.
Yankun Zhu, Pingqiang Zhou
IEEE Trans. Very Large Scale Integr. Syst.2
2023 A Speed- and Energy-Driven Holistic Training Framework for Sparse CNN Accelerators
abstract
Sparse convolution neural network (CNN) accelera-tors have shown to achieve high processing speed and low energy consumption by leveraging zero weights or activations, which can be further optimized by finely tuning the sparse activation maps in training process. In this paper, we propose a CNN training frame-work targeting at reducing energy consumption and processing cycles in sparse CNN accelerators. We first model accelerator's energy consumption and processing cycles as functions of layer-wise activation map sparsity. Then we leverage the model and propose a hybrid regularization approximation method to further sparsify activation maps in the training process. The results show that our proposed framework can reduce the energy consumption of Eyeriss by 31.33%, 20.6% and 26.6% respectively on MobileNet-V2, SqueezeNet and Inception-V3. In addition, the processing speed can be increased by$\boldsymbol{1.96\times, 1.4\times}$and$\boldsymbol{1.65}\times$respectively.
Yuanchen Qu, Pingqiang Zhou
DATE3
2023 Exploring Remote Power Attacks Targeting Parallel Data Encryption On Multi-Tenant FPGAs
abstract
Cloud service providers (CSPs) are increasingly incorporating Field Programmable Gate Arrays (FPGAs) into their cloud data centers due to the benefits of their flexibility and high performance in heterogeneous designs. However, the optimization of hardware resource utilization through multi-tenancy presents new security concerns. Prior research has demonstrated that remote side-channel attacks represent a significant security threat in the case of a single Advanced Encryption Standard (AES) module. However, it remains an open question whether parallel encryption can offer natural protection against Correlation Power Analysis (CPA). Our research focuses on side-channel attacks on parallel data encryption modules. We implemented delay-line based power sensors to collect mixed power traces and conducted CPA to steal the cipher key. Our results show that clocking methodology would have a significant influence on data protection. If parallel modules work at the same frequency without difference in clocking phase, the mixed voltage drops would contain sufficient information for attackers to decrypt the cipher key. Nevertheless, once the victim applies unique clocking phase to each module, he would convert voltage fluctuations from other modules into noises that offer a natural protection mechanism for parallel data encryption.
Yankun Zhu, Jindong Zhou, Pingqiang Zhou
ACM Great Lakes Symposium on VLSI3
2023 The study of TSV-induced and strained silicon-enhanced stress in 3D-ICs
Jindong Zhou, Youliang Jing, Pingqiang Zhou
Integr.4
2023 A Mapping Method Tolerating SAF and Variation for Memristor Crossbar Array Based Neural Network Inference on Edge Devices
abstract
There is an increasing demand for running neural network inference on edge devices. Memristor crossbar array (MCA) based accelerators can be used to accelerate neural networks on edge devices. However, reliability issues in memristors, such as stuck-at faults (SAF) and variations, lead to weight deviation of neural networks and therefore have a severe influence on inference accuracy. In this work, we focus on the reliability issues in memristors for edge devices. We formulate the reliability problem as a 0–1 programming problem, based on the analysis of sum weight variation (SWV) . In order to solve the problem, we simplify the problem with an approximation - different columns have the same weights, based on our observation of the weight distribution. Then we propose an effective mapping method to solve the simplified problem. We evaluate our proposed method with two neural network applications on two datasets. The experimental results on the classification application show that our proposed method can recover 95% accuracy considering SAF defects and can increase by up to 60% accuracy with variation σ =0.4. The results of the neural rendering application show that our proposed method can prevent render quality reduction.
Linfeng Zheng, Pingqiang Zhou
ACM J. Emerg. Technol. Comput. Syst.3
2023 Guest Editorial Special Issue on the Asian Hardware Oriented Security and Trust Symposium (AsianHOST 2022)
abstract
Asian Hardware Oriented Security and Trust Symposium (AsianHOST) is an annual symposium that aims to facilitate the rapid growth of hardware-based security research and development. Hardware security is a fashionable research area in both industry and academia. Its scope is consistently growing to embrace secure design, manufacturing, and deployment of modern and emerging interoperable computing, communication, storage devices, and circuits and systems. The 7th Asian Hardware Oriented Security and Trust Symposium (AsianHOST 2022) was held in hybrid mode on December 14–16 in Singapore. Among all the accepted contributions presented at the conference, a subset of top-rated articles was selected and invited for this Special Issue. The invited articles included extended new technical contributions and results and went through a peer-review process consisting of expert reviewers in the related topics. A brief description of the selected articles is as follows.
Chip-Hong Chang, Pingqiang Zhou, Yuan Cao 0003, Qiang Liu 0011
IEEE Trans. Circuits Syst. I Regul. Pap.2
2023 Analysis and Design of Precision-Scalable Computation Array for Efficient Neural Radiance Field Rendering
abstract
Neural Radiance Field (NeRF), a disruptive method for 3D representation and rendering, is extremely popular in the field of computer graphics and computer vision in the past three years. The most distinctive feature of NeRF models is their scene representation property, making it possible to quantize the models according to the complexity of the representing scenes. This paper proposes a novel approach to improve the efficiency of NeRF rendering by adopting precision-scalable computation. We first analyze and validate the idea of scene-dependent quantization for NeRF models. Based on that, we further propose look-up table (LUT) processing element (PE)-based precision-scalable computation unit designs. To evaluate the performance of different precision-scalable computing units, we implement these designs and compare the corresponding area, power, speed and energy efficiency. We also compare the proposed designs with existing approaches as well as the fixed precision approach for NeRF rendering tasks. The comparison results show that energy efficiency can be significantly improved by using precision-scalable computation for NeRF.
Kangjie Long, Chaolin Rao, Yunxiang He, Zhechen Yuan, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 An Energy-Efficient Accelerator for Medical Image Reconstruction From Implicit Neural Representation
abstract
This work presents an energy-efficient accelerator for medical image reconstruction from implicit neural representation (INR). The accelerator implements an INR-based algorithm to deliver high-quality medical image reconstruction with arbitrary resolution from a compact implicit format. In particular, we propose a dedicated hardware architecture based on an optimized computation flow for the INR-based reconstruction algorithm, which co-designs data reuse and computation load. The proposed architecture takes in the coordinate of the intersection of three scans and outputs all the voxel intensities, minimizing the data movement between on-chip and off-chip. To validate the proposed accelerator, we build a proof-of-concept prototype demonstration system using field programmable gate array (FPGA). We also map our design to 40nm CMOS technology to measure the performance of the proposed accelerator. The implementation results show that, running at 400MHz, the proposed accelerator is capable of processing medical images with$256\times 256$resolution in real-time at 26.3 frames per second (FPS), with a power consumption of only 795 mW. Comparison results show that the performance, as well as the energy efficiency of the proposed accelerator, outperforms the central processing unit (CPU)-based and graphic processing unit (GPU)-based implementations.
Chaolin Rao, Qing Wu 0001, Pingqiang Zhou, Jingyi Yu 0001, Yuyao Zhang 0005, Xin Lou 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 Thermal-Aware Layout Optimization and Mapping Methods for Resistive Neuromorphic Engines
abstract
Resistive neuromorphic engines can accelerate spiking neural network tasks with memristor crossbars. However, the stored weight is influenced by the temperature, which leads to accuracy and endurance degradation. The higher the temperature is, the larger the influence is. In this work, we propose a cross-array mapping method and a layout optimization method to reduce the thermal effect with the consideration of input distribution, weight value and layout of memristor crossbars. Experimental results show that our method reduces the peak temperature up to 10.4K and improves the endurance up to 1.72×.
Chengrui Zhang, Pingqiang Zhou
ASP-DAC3
2022 Two-Stage Energy Efficiency Optimization of Switched-Capacitor Converters for IoT Systems
abstract
IoT systems operate at low power domain to reduce energy consumption. Switched-capacitor converters can achieve high energy conversion efficiency while meeting the load requirements. Our proposed work optimizes the conversion efficiency of the switched-capacitor converters in two stages: in the early planning stage, we optimize the capacitance allocation among multiple SCCs considering the load requirements in both the active state and sleep state; in the runtime stage, we propose to finely tune the switching frequency and switch width according to load requirements of the active state. Results show that our proposed work could improve the energy efficiency by around 13% and 3% in two stages, respectively.
Yuanchen Qu, Qingfu Xu, Pingqiang Zhou
ISCAS4
2022 Defending against Adversarial Attacks in Deep Learning with Robust Auxiliary Classifiers Utilizing Bit-plane Slicing
abstract
Deep Neural Networks (DNNs) have been widely used in variety of fields with great success. However, recent research indicates that DNNs are susceptible to adversarial attacks, which can easily fool the well-trained DNN-based classifiers without being detected by human eyes. In this article, we propose to integrate the target DNN model with our robust bit-plane classifiers to defend against adversarial attacks. The bit-plane classifiers take bit-planes of input images for convolution, which is motivated by our observation that successful attacks aim to generate imperceptible perturbations, and they mainly affect the low-order bits of pixels in clean images when adding the perturbations. We also propose two metrics, bit-plane perturbation rate and channel modification rate, to further explain the robustness of bit-plane classifiers. We discuss potential adaptive attack and find that our defense can be effective as long as the adversarial examples are qualified. We conduct experiments on dataset CIFAR-10 and GTSRB under white-box attack and black-box attack. The results show that our defense method can effectively increase the average model accuracy from 16.23% to 83.53% under white-box attack and from 40.65% to 88.14% under black-box attack on CIFAR-10 without sacrificing the accuracy of clean images.
Jinxin Dong, Pingqiang Zhou
ACM J. Emerg. Technol. Comput. Syst.3
2022 ICARUS: A Specialized Architecture for Neural Radiance Fields Rendering
abstract
The practical deployment of Neural Radiance Fields (NeRF) in rendering applications faces several challenges, with the most critical one being low rendering speed on even high-end graphic processing units (GPUs). In this paper, we present ICARUS, a specialized accelerator architecture tailored for NeRF rendering. Unlike GPUs using general purpose computing and memory architectures for NeRF, ICARUS executes the complete NeRF pipeline using dedicated plenoptic cores (PLCore) consisting of a positional encoding unit (PEU), a multi-layer perceptron (MLP) engine, and a volume rendering unit (VRU). A PLCore takes in positions & directions and renders the corresponding pixel colors without any intermediate data going off-chip for temporary storage and exchange, which can be time and power consuming. To implement the most expensive component of NeRF, i.e., the MLP, we transform the fully connected operations to approximated reconfigurable multiple constant multiplications (MCMs), where common subexpressions are shared across different multiplications to improve the computation efficiency. We build a prototype ICARUS using Synopsys HAPS-80 S104, a field programmable gate array (FPGA)-based prototyping system for large-scale integrated circuits and systems design. We evaluate the power-performancearea (PPA) of a PLCore using 40nm LP CMOS technology. Working at 400 MHz, a single PLCore occupies 16.5 mm 2 and consumes 282.8 mW, translating to 0.105 uJ/sample. The results are compared with those of GPU and tensor processing unit (TPU) implementations.
Chaolin Rao, Huangjie Yu, Haochuan Wan, Jindong Zhou, Yueyang Zheng, Minye Wu, Anpei Chen, Binzhe Yuan, Pingqiang Zhou, Xin Lou 0001, Jingyi Yu 0001
ACM Trans. Graph.10
2021 Efficient Techniques for Training the Memristor-based Spiking Neural Networks Targeting Better Speed, Energy and Lifetime
abstract
Speed and energy consumption are two important metrics in designing spiking neural networks (SNNs). The inference process of current SNNs is terminated after a preset number of time steps for all images, which leads to a waste of time and spikes. We can terminate the inference process after proper number of time steps for each image. Besides, normalization method also influences the time and spikes of SNNs. In this work, we first use reinforcement learning algorithm to develop an efficient termination strategy which can help find the right number of time steps for each image. Then we propose a model tuning technique for memristor-based crossbar circuit to optimize the weight and bias of a given SNN. Experimental results show that the proposed techniques can reduce about 58.7% crossbar energy consumption and over 62.5% time consumption and double the drift lifetime of memristor-based SNN.
Pingqiang Zhou
ASP-DAC2
2021 A Quantized Training Framework for Robust and Accurate ReRAM-based Neural Network Accelerators
abstract
Neural networks (NN), especially deep neural networks (DNN), have achieved great success in lots of fields. ReRAM crossbar, as a promising candidate, is widely employed to accelerate neural network owing to its nature of processing MVM. However, ReRAM crossbar suffers high conductance variation due to many non-ideal effects, resulting in great inference accuracy degradation. Recent works use uniform quantization to enhance the tolerance of conductance variation, but these methods still suffer high accuracy loss with large variation. In this paper, firstly, we analyze the impact of the quantization and conductance variation on the accuracy. Then, based on two observation, we propose a quantized training framework to enhance the robustness and accuracy of the neural network running on the accelerator, by introducing a smart non-uniform quantizer. This framework consists of a robust trainable quantizer and a corresponding training method, and needs no extra hardware overhead and compatible with a standard neural network training procedure. Experimental results show that our proposed method can improve inference accuracy by 10% ~ 30% under large variation, compared with uniform quantization method.
Pingqiang Zhou
ASP-DAC2
2021 Tolerating Stuck-at Fault and Variation in Resistive Edge Inference Engine via Weight Mapping
abstract
There is an increasing demand for running neural network inference on edge devices. Memristor crossbar array (MCA) based accelerators can be used to accelerate neural networks on edge devices. However, reliability issues in memristors, such as stuck-at faults (SAF) and variations, lead to weight deviation of neural networks and therefore have severe influence on inference accuracy. In this work, we focus on reliability issues for edge devices. We formulate the reliability problem as a 0-1 programming problem, based on the analysis of sum weight variation (SWV). In order to solve the problem, we simplify the problem with an approximation - different columns have the same weights - based on our observation of the weight distribution. Then we propose an effective mapping method to solve the simplified problem. The experimental results show that our proposed method can recover 95% accuracy considering SAF defects and can increase by up to 60% accuracy in variation σ=0.4.
Linfeng Zheng, Pingqiang Zhou
ACM Great Lakes Symposium on VLSI3
2020 How Secure Is Split Manufacturing in Preventing Hardware Trojan?
abstract
With the trend of outsourcing fabrication, split manufacturing is regarded as a promising way to both acquire the high-end nodes in untrusted external foundries and protect the design from potential attackers. However, in this article, we show that split manufacturing is not inherently secure, that a hardware Trojan attacker can still recover necessary information with a proximity-based or a simulated-annealing-based mapping approach together with a probability-based or net-based pruning method at the placement level. We further propose a defense approach by moving the insecure gates away from their easily attacked candidate locations. Results on benchmark circuits show the effectiveness of our proposed methods.
Yajun Yang, Tsung-Yi Ho, Yier Jin, Pingqiang Zhou
ACM Trans. Design Autom. Electr. Syst.6
2019 Optimizing the Energy Efficiency of Power Supply in Heterogeneous Multicore Chips with Integrated Switched-Capacitor Converters
abstract
Energy efficiency is a major concern in heterogeneous multi-core chips. Due to the switching-capacitor converter (SCC) has wide output voltages and high potential ratio efficiency, they are widely used in multi-core chips. In this paper we propose the optimization of Metal-Insulator-Metal (MIM) capacitance resource allocation and converter ratio selection for SCCs to improve the power efficiency by transforming the mixed integer nonlinear programming (MINLP) problems into a series of convex problems. The experimental results show that our approach can achieve a 9%-13% improvement in power efficiency and can be applied to more complicated heterogeneous multicore scenarios.
Leilei Wang, Dejia Shang, Cheng Zhuo, Pingqiang Zhou
DATE5
2019 An orchestrated NoC prioritization mechanism for heterogeneous CPU-GPU systems
Xiangwei Cai, Jieming Yin, Pingqiang Zhou
Integr.3
2019 Reliability- and performance-driven mapping for regular 3D NoCs using a novel latency model and Simulated Allocation
Wei Gao 0024, Zhiliang Qian, Pingqiang Zhou
Integr.3
2019 Run-time demand estimation and modulation of on-chip decaps at system level for leakage power reduction in multicore chips
Leilei Wang, Cheng Zhuo, Pingqiang Zhou
Integr.3
2019 A Cross-Layer Framework for Temporal Power and Supply Noise Prediction
abstract
In modern microprocessor and SoC designs, supply noise margin has been significantly reduced due to the continuously decreasing supply voltage level. On the other hand, with increasing current density, chips may see larger supply noise variations on various spots and from time to time. As a result, chip robustness and reliability are inevitably deteriorated with more frequent supply noise emergencies. It is therefore crucial to have an efficient supply noise prediction method to enhance design robustness. The state-of-art solutions either try to build a spatial noise estimation framework at the layout-level using the limited distributed physical noise sensors or attempt to develop emergency predictors at the architecture-level thus ignore back-end power delivery details. In this paper, we propose a cross-layer framework for temporal supply noise prediction. Our method not only accounts for the temporal characteristics of workload execution at micro-architecture-level but also incorporates the power delivery model at the circuit-level into such system-level prediction. In order to enable the capability of on-the-fly noise prediction, we first bridge the gap between system-level workload and micro-architectural-level power by employing an ordinary least square-based power estimation model and an adaptive auto-regressive integrated moving average model (ARIMA)-based power prediction model. Then a layout-level supply noise model is developed to explore the correlations between micro-architectural-level power and layout-level supply noise. Compared with existing methods, the proposed ARIMA-based power model improves the prediction performance by up to 37.5%/63.0% in X86/ARM. Moreover, compared with SPICE simulation, our framework is able to estimate present supply noise with an average error of 0.005% and predict future supply noise with an average error of 1.58%/1.17% for X86/ARM architecture.
Cheng Zhuo, Pingqiang Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 Dependable Visual Light-Based Indoor Localization with Automatic Anomaly Detection for Location-Based Service of Mobile Cyber-Physical Systems
abstract
Indoor localization has become popular in recent years due to the increasing need of location-based services in mobile cyber-physical systems (CPS). The massive deployment of light emitting diodes (LEDs) further promotes the indoor localization using visual light. As a key enabling technique for mobile CPS, accurate indoor localization based on visual light communication remains nontrivial due to various non-idealities such as attenuation induced by unexpected obstacles. The anomalies of localization can potentially reduce the dependability of location-based services. In this article, we develop a novel indoor localization framework based on relative received signal strength. Most importantly, an efficient method is derived from the triangle inequality to automatically detect the abnormal LED lamps that are blocked by obstacles. These LED lamps are then ignored by our localization algorithm so that they do not bias the localization results, which improves the dependability of our localization framework. As demonstrated by the simulation results, the proposed techniques can achieve superior accuracy over the conventional approaches, especially when there exist abnormal LED lamps.
Yang Liu 0064, Xiaoming Chen 0003, Dileep Kadambi, Ajinkya Bari, Xin Li 0001, Shiyan Hu 0001, Pingqiang Zhou
ACM Trans. Cyber Phys. Syst.7
2017 Machine Learning for Noise Sensor Placement and Full-Chip Voltage Emergency Detection
abstract
Power supply fluctuation can be potential threat to the correct operations of processors, in the form of voltage emergency that happens when supply voltage drops below a certain threshold. Noise sensors (with either analog or digital outputs) can be placed in the nonfunction area of processors to detect voltage emergencies by monitoring the runtime voltage fluctuations. Our work addresses two important problems related to building a sensor-based voltage emergency detection system: 1) offline sensor placement, i.e., where to place the noise sensors so that the number and locations of sensors are optimized in order to strike a balance between design cost and chip reliability and 2) online voltage emergency detection, i.e., how to use these placed sensors to detect voltage emergencies in the hotspot locations. In this paper, we propose integrated solutions to these two problems, respectively, for analog and digital (more specifically, binary) sensor outputs, by exploiting the voltage correlation between the sensor candidate locations and the hotspot locations. For the analog case, we use the Group Lasso and an ordinary least squares approach; for the binary case, we integrate the Group Lasso and the SVM approach. Experimental results show that, our approach can achieve 2.3X-2.7X better voltage emergency detection results on average for analog outputs when compared to the state-of-the-art work; and for the binary case, on average our methodology can achieve up to 21% improvement in prediction accuracy compared to an approach called max-probability-no-prediction.
Shupeng Sun, Xin Li 0001, Haifeng Qian, Pingqiang Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2015 A statistical methodology for noise sensor placement and full-chip voltage map generation
abstract
Noise margin violation, also known as voltage emergency induced by continuously reducing noise margin and increasing magnitude of current swings, is becoming a severe threat to the correct execution of applications in processors. Noise sensors can be placed in the non-function area of processors to detect such emergencies by monitoring runtime voltage fluctuations. In this work, we aim to accurately predict the voltage droops using a small set of sensors. We achieve our goal in two steps: We first propose a methodology via group lasso approach to select the optimal set of noise sensors, then build a practical model via ordinary least-squares fitting approach to predict the voltage in the function area of the chip, using the selected sensors in non-function area. Experiment results show that when compared to the full-chip voltage transient simulation, the prediction error of our model is much less than 0.01, and compared to prior work, our approach can achieve better error rates of voltage emergency detection (less than half).
Shupeng Sun, Pingqiang Zhou, Xin Li 0001, Haifeng Qian
DAC3
2014 Energy-Efficient Time-Division Multiplexed Hybrid-Switched NoC for Heterogeneous Multicore Systems
abstract
NoCs are an integral part of modern multicore processors, they must continuously support high-throughput low-latency on-chip data communication under a stringent energy budget when system size scales up. Heterogeneous multicore systems further push the limit of NoC design by integrating cores with diverse performance requirements onto the same die. Traditional packet-switched NoCs, which have the flexibility of connecting diverse computation and storage devices, are facing great challenges to meet the performance requirements within the energy budget due to latency and energy consumption associated with buffering and routing at each router. In this paper, we take advantage of the diversity in performance requirements of on-chip heterogeneous computing devices by designing, implementing, and evaluating a hybrid-switched network that allows the packet-switched and circuit-switched messages to share the same communication fabric by partitioning the network through time-division multiplexing (TDM). In the proposed hybrid-switched network, circuit-switched paths are established along frequently communicating nodes. Our experiments show that utilizing these paths can improve system performance by reducing communication latency and alleviating network congestion. Furthermore, better energy efficiency is achieved by reducing buffering in routers and in turn enabling aggressive power gating.
Jieming Yin, Pingqiang Zhou, Sachin S. Sapatnekar, Antonia Zhai
IPDPS2
2014 Distributed On-Chip Switched-Capacitor DC-DC Converters Supporting DVFS in Multicore Systems
abstract
Dynamic voltage and frequency scaling (DVFS) is a powerful technique to reduce power consumption in a chip multiprocessor. To support DVFS in the multicore power delivery network, we integrate on-chip switched-capacitor (SC) dc-dc converters that can work with multiple conversion ratios to provide varying levels of Vdd supplies. We study the application of such SC converters in multicore chips by simulation. Our results show that distributed SC converters can significantly reduce the voltage droop seen by the local core loads by providing better localized power regulation. Considering the fact that the current distribution in a multicore chip is unbalanced, we further develop computer-aided design techniques to automate the design (size) and distribution (number and location) of these SC converters, using the efficiency of the whole power delivery system as the optimization metric. This is a major concern, but has not been addressed at the system level in prior research. We develop models for the power loss of such a system as a function of size and distribution of the SC converters, then propose an approach to optimize the SC converters to maximize the efficiency of the system, while considering all the possible conversion ratios an SC converter can work with. We verify the accuracy of our models for the power loss in the power delivery system, and demonstrate the efficiency of our techniques to optimize the SC converters on both homogenous and heterogenous multicore chips.
Pingqiang Zhou, Ayan Paul, Chris H. Kim, Sachin S. Sapatnekar
IEEE Trans. Very Large Scale Integr. Syst.1
2013 Placement optimization of power supply pads based on locality
abstract
This paper presents an efficient algorithm for the placement of power supply pads in flip-chip packaging for high-performance VLSI circuits. The placement problem is formulated as a mixed-integer linear program (MILP), subject to the constraints on mean-time-to-failure (MTTF) for the pads and the voltage drop in the power grid. To improve the performance of the optimizer, the pad placement problem is solved based on the divide-and-conquer principle, and the locality properties of the power grid are exploited by modeling the distant nodes and sources coarsely, following the coarsening stage in multi-grid-like approach. An accurate electromigration (EM) model that captures current crowding and Joule heating effects is developed and integrated with our C4 placement approach. The effectiveness of the proposed approach is demonstrated on several designs adapted from publicly released benchmarks.
Pingqiang Zhou, Vivek Mishra, Sachin S. Sapatnekar
DATE1
2012 Circuit reliability: From Physics to Architectures: Embedded tutorial paper
abstract
In the period of extreme CMOS scaling, reliability issues are becoming a critical problem. These problems include issues related to device reliability, in the form of bias temperature instability, hot carrier injection, time-dependent dielectric breakdown of gate oxides, as well as interconnect reliability concerns such as electromigration and TSV stress in 3D integrated circuits. This tutorial surveys these effects, and discusses methods for mitigating them at all levels of design.
Jianxin Fang, Saket Gupta, Sanjay V. Kumar, Sravan K. Marella, Vivek Mishra, Pingqiang Zhou, Sachin S. Sapatnekar
ICCAD6
2012 Optimization of on-chip switched-capacitor DC-DC converters for high-performance applications
abstract
On-chip switched-capacitor (SC) DC-DC converters have recently been demonstrated in silicon for high-performance applications such as multicore processors. The efficiency of the power delivery system using SC converters is a major concern, but this has not been addressed at the system level in prior research. This work develops models for the efficiency of such a system as a function of size and layout of the SC converters, and proposes an approach to optimize the size and layout of the SC converter to minimize power loss. The efficiency of these techniques is demonstrated on both homogenous and heterogenous multicore chips.
Pingqiang Zhou, Won Ho Choi, Bongjin Kim, Chris H. Kim, Sachin S. Sapatnekar
ICCAD1
2012 Energy-efficient non-minimal path on-chip interconnection network for heterogeneous systems
abstract
Network-on-Chips (NoCs) in heterogeneous systems containing both CPU and GPU cores must be designed to satisfy the performance requirements of both latency-sensitive CPU traffic and throughput-intensive GPU traffic. DVFS and adaptive routing can potentially improve NoC energy and performance efficiency. We further notice that GPU traffic can sometimes tolerate a slack defined as the number of cycles a packet can be delayed without causing performance penalty. In this work, we take advantage of the slack in GPU packets to route packets through non-minimal path, so that routers can operate at a lower frequency without suffering performance penalty.
Jieming Yin, Pingqiang Zhou, Anup Holey, Sachin S. Sapatnekar, Antonia Zhai
ISLPED2
2012 Optimized 3D Network-on-Chip Design Using Simulated Allocation
abstract
Three-dimensional (3D) silicon integration technologies have provided new opportunities for Network-on-Chip (NoC) architecture design in Systems-on-Chip (SoCs). In this article, we consider the application-specific NoC architecture design problem in a 3D environment. We present an efficient floorplan-aware 3D NoC synthesis algorithm based on simulated allocation (SAL), a stochastic method for traffic flow routing, and accurate power and delay models for NoC components. We demonstrate that this method finds greatly improved solutions compared to a baseline algorithm reflecting prior work. To evaluate the SAL method, we compare its performance with the widely used simulated annealing (SA) method and show that SAL is much faster than SA for this application, while providing solutions of very similar quality. We then extend the approach from a single-path routing to a multipath routing scheme and explore the trade-off between power consumption and runtime for these two schemes. Finally, we study the impact of various factors on the network performance in 3D NoCs, including the TSV count and the number of 3D tiers. Our studies show that link power and delay can be significantly improved when moving from a 2D to a 3D implementation, but the improvement flattens out as the number of 3D tiers goes beyond a certain point.
Pingqiang Zhou, Ping-Hung Yuh, Sachin S. Sapatnekar
ACM Trans. Design Autom. Electr. Syst.1
2011 NoC frequency scaling with flexible-pipeline routers
Pingqiang Zhou, Jieming Yin, Antonia Zhai, Sachin S. Sapatnekar
ISLPED1
2010 Application-specific 3D Network-on-Chip design using simulated allocation
abstract
Three-dimensional (3D) silicon integration technologies have provided new opportunities for Network-on-Chip (NoC) architecture design in Systems-on-Chip (SoCs). In this paper, we consider the application-specific NoC architecture design problem in a 3D environment. We present an efficient floorplan-aware 3D NoC synthesis algorithm, based on simulated allocation, a stochastic method for traffic flow routing, and accurate power and delay models for NoC components. We demonstrate that this method finds greatly improved topologies for various design objectives such as NoC power (average savings of 34%), network latency (average reduction of 35%) and chip temperature (average reduction of 20%).
Pingqiang Zhou, Ping-Hung Yuh, Sachin S. Sapatnekar
ASP-DAC1
2009 Congestion-aware power grid optimization for 3D circuits using MIM and CMOS decoupling capacitors
abstract
In three-dimensional (3D) chips, the amount of supply current per package pin is significantly more than in two-dimensional (2D) designs. Therefore, the power supply noise problem, already a major issue in 2D, is even more severe in 3D. CMOS decoupling capacitors (decaps) have been used effectively for controlling power grid noise in the past, but with technology scaling, they have grown increasingly leaky. As an alternative, metal-insulator-metal (MIM) decaps, with high capacitance densities and low leakage current densities, have been proposed. In this paper, we explore the tradeoffs between using MIM decaps and traditional CMOS decaps, and propose a congestion-aware 3D power supply network optimization algorithm to optimize this tradeoff. The algorithm applies a sequence-of-linear-programs based method to find the optimum tradeoff between MIM and CMOS decaps. Experimental results show that power grid noise can be more effectively optimized after the introduction of MIM decaps, with lower leakage power and little increase in the routing congestion, as compared to a solution using CMOS decaps only.
Pingqiang Zhou, Karthikk Sridharan, Sachin S. Sapatnekar
ASP-DAC1
2007 Thermal Effects with Leakage Power Considered in 2D/3D Floorplanning
abstract
Leakage power is becoming a key design challenge in current and future CMOS designs. Due to technology scaling, the leakage power is rising so quickly that it largely elevates the die temperature. In this paper, we deeply investigate the impact of leakage power on thermal profile in 2D and 3D floorplanning. Our results show that chip temperature can increase by about 11 V in 2D design and 68 V for 3D case with leakage power considered. Then we propose a thermal-driven floorplanning flow integrated with an iterative leakage-aware thermal analysis process to optimize chip temperature and save leakage power consumption. Experimental results show that for 2D design, the max chip temperature can be reduced by about 8 "C and the proportion of leakage power to total power can be reduced from 19.17% to 11.12%. The corresponding results for 3D are 60 degC temperature reduction and 16.3% less leakage power proportion.
Pingqiang Zhou, Yuchun Ma, Qiang Zhou 0001, Xianlong Hong
CAD/Graphics1
2007 3D-STAF: scalable temperature and leakage aware floorplanning for three-dimensional integrated circuits
abstract
Thermal issues are a primary concern in the threedimensional (3D) integrated circuit (IC) design. Temperature, area, and wire length must be simultaneously optimized during 3D floorplanning, significantly increasing optimization complexity. Most existing floorplanners use combinatorial stochastic optimization techniques, hampering performance and scalability when used for 3D floorplanning. In this work, we propose and evaluate a scalable, temperature-aware, force-directed floorplanner called 3D-STAF. Force-directed techniques, although efficient at reacting to physical information such as temperature gradients, must eventually eliminate overlap. This can cause significant displacement when used for heterogeneous blocks. To smooth the transition from an unconstrained 3D placement to a legalized, layer-assigned floorplan, we propose a three-stage force-directed optimization flow combined with new legalization techniques that eliminate white spaces and block overlapping during multi-layer floorplanning. A temperature-dependent leakage model is used within 3D-STAF to permit optimization based on the feedback loop connecting thermal profile and leakage power consumption. 3D-STAF has good performance that scales well for large problem instances. Compared to recently published 3D floorplanning work, 3D-STAF improves the area by 6%, wire length by 16%, via count by 22%, peak temperature by 6% while running nearly 4× faster on average.
Pingqiang Zhou, Yuchun Ma, Zhuoyuan Li 0003, Robert P. Dick, Hai Zhou 0001, Xianlong Hong, Qiang Zhou 0001
ICCAD1