VLDB 2026 Research / reviewers in the wild / expert
Yuanqing Cheng
dblp:05/10696
· DBLP profile ↗
52ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0003-2477-314XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 48 · 5 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Full-Chip Thermal Map Estimation by Multimodal Data Fusion via Denoising DiffusionabstractAs the integration density on-chip increases, thermal challenges become more prominent, and the key to addressing these challenges lies in effective and efficient thermal analysis. Current thermal estimation methods face a fundamental paradigm limitation: they rely on single data source that captures only partial aspects of complex thermal behavior. Performance counter-based methods capture computational activity but lacks of accurate thermal readouts, while sensor-based approaches provide local temperature measurements but lack of comprehensive spatial coverage. This single-source paradigm has created an insurmountable accuracy ceiling in thermal map estimation, limiting the effectiveness of modern thermal management systems. We introduce the first multimodal data fusion framework for full-chip thermal estimation, leveraging denoising diffusion models to synergistically combine performance counters and thermal sensors. Our approach treats thermal mapping as a conditional generation problem, where complementary data modalities guide the reconstruction process through progressive denoising. The key insight is that thermal behavior is inherently multimodal-requiring both activity context (performance counters) and temperature ground truth (thermal sensors) for accurate estimation. Our multimodal fusion delivers transformative results: $90 \%+$ improvement over single-source methods (0.347 vs 4.862-6.302 average RMSE), 41.8% average RMSE improvement over state-of-the-art approaches, and robust performance across diverse sensor configurations (9-25 sensors). Furthermore, our framework effectively captures hotspot locations, with errors typically within 0.5 K. More importantly, this work establishes multimodal data fusion as a new paradigm for thermal analysis, opening new research directions and enabling next-generation thermal management systems with unprecedented accuracy and reliability, fundamentally changing how we approach thermal analysis in modern processors. Yuquan Sun, Yuanqing Cheng |
ASP-DAC | 4 |
| 2026 | Φ-BO: Physics-Informed Bayesian Optimization for Multi-Port Decoupling Capacitor Placement in 2.5-D ChipletsabstractPower distribution network (PDN) optimization in 2.5-D chiplet architectures represents a critical bottleneck as designs scale to 100+ integrated chiplets, where decoupling capacitor placement becomes a multi-port optimization challenge requiring millions of expensive electromagnetic (EM) simulations. Current state-of-the-art (SOTA) methods - from genetic algorithms (GA) to reinforcement learning (RL) - treat PDN as black-box functions, failing to exploit inherent physical structure and scaling exponentially with problem complexity. We introduce Φ-BO, the first physics-informed Bayesian optimization (BO) framework specifically designed for multi-port decoupling capacitor placement in 2.5-D chiplet PDN. Our key innovation systematically integrates EM field theory into machine learning (ML) optimization through novel spatial feature transformations and Multi-Port Aware Transformation (MPAT), enabling a paradigm shift from black-box to physics-aware optimization. This approach captures spatial dependencies and port coupling effects, dramatically reducing effective problem dimensionality while enabling intelligent exploration of discrete placement configurations. Demonstrated on a 22-chiplet RISC-V processor design, Φ-BO achieves 23% impedance improvement, and 3 × faster convergence compared to SOTA methods. Quansen Wang, Yuchuan Lin, Zhuohua Liu, Ning Xu 0006, Yuanqing Cheng |
ASP-DAC | 7 |
| 2026 | ConvGA: Convolution Network-Guided Genetic Algorithm for Optimal PDN Decoupling DesignabstractPower supply noise has emerged as a critical bottleneck in modern integrated circuit design, where increasing current densities and higher operating frequencies pose significant challenges to system reliability. While decoupling capacitors (decaps) serve as the primary solution for suppressing power delivery network (PDN) noise, determining their optimal values and placement remains computationally prohibitive using traditional methods. This aritcle introduces ConvGA, a novel framework that seamlessly integrates convolutional neural networks (CNN) with genetic algorithms to revolutionize PDN decap optimization. At the heart of ConvGA is a specialized CNN architecture trained on comprehensive boundary element method (BEM) simulations, enabling ultra-fast impedance prediction for arbitrary PCB configurations. Our CNN achieves remarkable accuracy while reducing impedance computation time from hours to mere milliseconds—a 500× speedup over conventional BEM calculations. This acceleration enables the genetic algorithm to efficiently explore vast design spaces through adaptive population control and dynamic constraint mechanisms, systematically minimizing both the number of required capacitors and the deviation from target impedance. Extensive experiments on industrial-scale PDNs demonstrate that ConvGA achieves a 15× reduction in optimization time while requiring 30% fewer capacitors compared to state-of-the-art methods, consistently producing high-quality solutions across diverse PDN configurations. Yuchuan Lin, Ning Xu 0006, Wei W. Xing, Yuanqing Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Folivora: Ultralow-Power Microprocessor Design With Nano-Electromechanical Relay and Nanotube MemoryabstractIn post-Moore era, CMOS technology scaling has encountered enormous design and fabrication challenges. “Power Wall” limits the further increase of integration density. Emerging AI computing and data center deployments aggravate the power consumption problem further. In the pursuit of efficient computing paradigm, Nano Electro-mechanical (NEM) relay and Nanotube Random Access Memory (NRAM) technology have attracted enormous attention and have ultra-low power consumption compared to CMOS counterparts. NEM relay is a kind of device based on electronic and mechanical interaction switching, characterized by remarkably low power consumption. This article explores the application of NEM relay and NRAM technology to build a complex RISC processor, aiming to achieve much lower power without degrading performance. The controller and data path can be implemented with primitive logic gates made of NEM relays, and on-chip cache can be implemented with NRAM. Experimental results show that the energy efficiency of the processor design based on NEM relay and NRAM can be improved by 88.2% and 78.9% compared with CMOS technology based in-order and out-of-order microprocessors, respectively. Meanwhile, the performance can be improved by 42.9% and the instruction execution time can be reduced by more than 17.9%, which implies the potentials of NEM relay and NRAM for emerging ultra-low power applications. Yuanqing Cheng, Ying Wang 0001, Rui Wang 0014 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | ASAP: Accelerating Corner-Based Timing Analysis With Bayesian Active Self-Attention Neural ProcessabstractWith the advancement of modern nanoscale technology nodes, Static Timing Analysis (STA) has become an indispensable technique for ensuring circuit reliability and performance across diverse process conditions. However, traditional STA methods scale poorly to the explosion of process corners in the nanoscale fabrication technology. Despite some seminal works in using AI to accelerate such processes, they either lack reliability or stability. To this end, we introduce ASAP, a novel approach addressing this challenge by combining both the latest deep learning methods and the classical Bayesian models to deliver scalable and accurate predictions with a self-calibration strategy to ensure reliability. Technically, the ASAP novelly integrates self-attention to help identify and prioritize crucial features under various input conditions and employs Neural Process to make confidence-based predictions for the final timing results. Furthermore, ASAP is equipped with Active Learning for self-refinement and self-correction. Experimental evaluations on benchmark circuits demonstrate that our method surpasses state-of-the-art work in STA accuracy by 18% in terms of prediction accuracy. Longze Wang, Wei W. Xing, Zhelong Wang, Christos P. Sotiriou, Nikolaos Sketopoulos, Ning Xu 0006, Yuanqing Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | Cool3D: Cost-Optimized and Efficient Liquid Cooling for 3D Integrated CircuitsabstractCMOS scaling faces challenges due to lithography and device physics issues, leading to increased costs and difficulties in expanding chip footprint. 3D integration technology offers increased integration density without increasing footprint, but elevated power density makes heat dissipation a significant challenge. Microchannel cooling effectively removes heat inside 3D chips. Traditional microchannel optimizations typically focus only on minimizing pump power within a limited parameter design space, leading to suboptimal cooling efficiency. Moreover, existing research rarely considers manufacturing costs, limiting practical application. To address these issues, we propose a high-dimensional non-uniform microchannel design scheme based on Segmented Sampling Bayesian Optimization (SSBO). This multi-parameter collaborative optimization framework comprehensively optimizes microchannel design. Our method reduces pump power by 70% compared to limited parameter design spaces. Additionally, we introduce a cost model for microchannel design, formulating a multi-objective optimization problem that considers both manufacturing cost and pump power consumption. By solving the multi-objective optimization problem by searching for the Pareto front, we demonstrate a balanced design between microchannel manufacturing cost and pump power and provide guidelines for key design parameters. Bingrui Zhang, Yuquan Sun, Yuanqing Cheng |
DATE | 5 |
| 2025 | A Comprehensive Inductance-Aware Modeling Approach to Power Distribution Network in Heterogeneous 3D Integrated CircuitsabstractHeterogeneous 3D integration technology is a cost-effective and high-performance alternative to planar integrated circuits (ICs). In this paper, we propose an on-chip power distribution network (PDN) modeling technique for heterogeneous 3D-ICs (H3D-ICs), which explicitly takes the effects of on-chip inductance into account. The proposed model facilitates efficient transient and AC simulations with integrated inductive effects, enabling accurate noise characterization at high frequencies and facilitating the exploration of early-stage PDN design. The model is validated via HSPICE simulations, demonstrating a maximum error below 1% and achieving average speedups of 1.5x in transient and 8.5x in AC simulations. Quansen Wang, Vasilis F. Pavlidis, Yuanqing Cheng |
DATE | 3 |
| 2025 | Inter-chip Clock Network Synthesis on Passive Interposer of 2.5D Chiplet Considering Transmission Line EffectabstractWith the slowdown of technology node scaling, 2.5D Chiplet technology has emerged as a promising approach to sustain Moore’s Law. When routing clock signal on the passive interposer layer, transmission line effect must be considered, as signal rise/fall delay becomes comparable to propagation delay, significantly impacting signal integrity and system performance. Additionally, since active devices cannot be placed on the passive interposer, buffers can only be inserted in the chiplets mounted on the interposer. This paper investigates the problem of clock network synthesis for a 2.5D Chiplet with a passive interposer. Firstly, we propose to use a transmission line model to evaluate inter-chip clock skew and delay accurately. Then, we propose a clock network synthesis method considering transmission line effect and impedance matching. Finally, we propose a buffer insertion method with minimum clock wire detouring on passive interposer in order to optimize the inter-chiplet clock network. Experimental results show that our fast transmission line model can calculate clock skew more accurately compared to conventional Elmore model and only results in 2.4% error compared to HSPICE simulations. The clock network syntheses on several 2.5D Chiplet benchmarks show that our proposed algorithm is effective to construct zero-skew clock network for a 2.5D Chiplet in terms of clock skew and buffer area. Tai Yan, Ning Xu 0006, Yuanqing Cheng |
VLSI-SoC | 5 |
| 2025 | Testing and fault tolerance techniques for carbon nanotube-based FPGAs
Kangwei Xu, Rui Wang 0014, Yuanqing Cheng |
Integr. | 5 |
| 2025 | Carbon Nanotube Interconnect Optimizations With Bayesian Neural Network and Bayesian OptimizationabstractAs Cu interconnects near their physical limits with continued technology scaling, carbon nanotube (CNT) interconnects have emerged as a promising alternative due to their excellent conductivity. However, fabrication immaturity introduces significant process variations, causing discrepancies between ideal and actual performance. This article presents a novel approach to optimize CNT interconnects considering process variations. We first develop a parameterized CNT interconnect model that accounts for process variations. Using this model, a Bayesian Neural Network (BNN) is proposed to predict performance distributions by leveraging its inherent uncertainty. We then introduce a Bayesian optimization framework that uses the BNN’s posterior to jointly optimize interconnect parameters and buffer insertion, targeting Area-Delay Product (ADPIn this work, we use area-delay product ziegler2001optimal as the performance metric, though other metrics can also be applied within our framework.) with process variations. Experimental results demonstrate the effectiveness of our approach. The proposed BNN model achieves over 95% prediction accuracy for interconnects performance distributions. Compared to existing methods, our method achieves an average ADP improvement of 20.3% over the state-of-the-art methods and 13% over the standard Monte Carlo method. Compared with Monte Carlo method, our method also achieves an average 8.8x acceleration. Moreover, the optimized CNT interconnects show an average improvement of 82.8% in ADP and 68.7% in delay compared to Cu interconnects. This work offers an effective method for optimizing CNT interconnects under process variations and highlights their potential as a viable alternative to Cu interconnects in future integrated circuits. Zhelong Wang, Wei W. Xing, Ning Xu 0006, Yuanqing Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | BoCNT: A Bayesian Optimization Framework for Global CNT Interconnect OptimizationabstractAs the prevailing copper interconnect technology advances to its fundamental physical limit, interconnect delay due to ever-increasing wire resistivity shows a significant impact on circuit performance. Bundled single-wall carbon nanotubes (SWCNTs) interconnects have emerged as a promising candidate technique to replace copper interconnects thanks to their superior conductivity and immunity to electromigration. To deliver satisfying performance within low power consumption, the CNT interconnect timing is optimized by adjusting either the interconnect geometry, e.g., CNT diameter and nanotube pitch, or the buffer insertion. These two operations are normally optimized separately, which leads to an inferior design that is not global optimum. To resolve this problem, we first propose a model that parameterizes SWCNT interconnects. We then leverage the known Bayesian optimization to optimize SWCNT global interconnect and buffer insertion simultaneously to promote interconnect performance. The proposed method is assessed based on a set of interconnect benchmarks at 22nm technology node. Compared to the state-of-the-art methods, the Bayesian co-optimization technique can reduce more than 17% power delay product (PDP). Additionally, we evaluate the SWCNT interconnect performance at 32nm, 22nm and 16nm technology nodes. Compared to the SOTA method with the same wire dimension, SWCNT interconnects optimized by the proposed method can further reduce delay and PDP relative to copper by 35% and 45% on average, which highlights the promising prospect of the SWCNT interconnect technology and the effectiveness of our proposed technique. Ning Xu 0006, Wei W. Xing, Yuanqing Cheng |
ASPDAC | 4 |
| 2024 | MAUnet: Multiscale Attention U-Net for Effective IR Drop PredictionabstractThe efficient analysis of power grids is a crucial yet computationally challenging task in integrated circuit (IC) design, given the shrinking power supply voltage of ultra deep-submicron VLSI design. Different from the conventional modified nodal analysis technique, this paper introduces MAUnet, an innovative machine-learning model that redefines state-of-the-art full-chip static IR drop prediction. MAUnet ingeniously integrates multi-scale convolutional blocks, attention mechanisms, and U-Net architecture to optimize prediction accuracy. The multi-scale convolutional blocks significantly enhance feature extraction from image-based data, while the attention mechanism precisely identifies hotspot regions. The U-Net architecture, on the other hand, enables scalable image-to-image prediction applicable to circuits of any size. Uniquely, MAUnet also incorporates a pioneering fusion method that synergies both power grids and image-based data. Additionally, we introduce a low-rank approximation transfer learning technique to extend MAUnet's applicability to unseen test cases. Benchmark tests validate MAUnet's superior performance, achieving an average error of less than 6% relative to the average IR drop on three benchmarks. The performance enhancements offered by our proposed method are substantial, outperforming the current state-of-the-art method, IREDGe, by considerable margins of 29%, 65%, and 68% in three canonical benchmarks. Transfer learning is validated to enable model to achieve effective improvement on real circuit test cases. Compared to commercial tools, which often require hours to deliver results, the proposed method provides orders of magnitude speed-up with negligible error in practice. Yuanqing Cheng, Yage Lin, Kelin Peng, Shunchuan Yang, Zhou Jin 0001, Wei W. Xing |
DAC | 2 |
| 2024 | ARO: Autoregressive Operator Learning for Transferable and Multi-fidelity 3D-IC Thermal Analysis With Active LearningabstractAs 3D integrated circuits (ICs) have emerged as a promising direction in the semiconductor industry, thermal issues in 3D-ICs have become increasingly prominent. In this work, we develop a novel machine learning (ML) thermal analysis framework, namely Autoregressive Operator (ARO), to address the pressing need for rapid yet highly accurate thermal predictions during the chip design process. Unlike traditional ML-based methods that can only deal with scenarios of well-defined input-output domains, ARO learns the thermal diffusion operator such that it can generalize to any unseen circuits and map the power traces to the steady-state/transient thermal spatial-temporal distributions. To further reduce the computational demand of data preparation, we equip ARO with multi-fidelity fusion to exploit the advantage of computationally cheap low-fidelity simulations and expensive high-fidelity simulations and active learning to guide the preparation of training data. Our results show that, for the unseen testing cases, a well-trained ARO can produce accurate results with about 1000× speedup compared to MTA. Moreover, equipped with active learning, ARO achieves at least 25% data reduction compared to pseudo-random strategies. Yuanqing Cheng, Weiheng Zeng, Zhenjie Lu, Vasilis F. Pavlidis |
ICCAD | 2 |
| 2024 | Multicorner Timing Analysis Acceleration for Iterative Physical Design of ICsabstractWe propose a multi-corner multi-stage timing analysis prediction framework using a generalized linear model with latent features. We then further improve such methods using kernel trick extension, transfer learning with knowledge from previous designs, and multi-output feature engineering to deliver state-of-the-art (SOTA) prediction accuracy with very limited training data. Most importantly, our method is equipped with a Bayesian decision strategy to deliver reliable predictions with accuracy close to 100%, pushing the frontier of the machine-learning-based STA for practical implementation in the industry environment, where reliability is highly desired. Experimental results show that the accuracy of our proposed method outperforms the SOTA competitors by up to 4x and can improve prediction accuracy to 100% with little extra STA executions. Wei W. Xing, Longze Wang, Zhelong Wang, Zhaoyu Shi, Ning Xu 0006, Yuanqing Cheng, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | PDG: A Prefetcher for Dynamic Graph UpdatingabstractDynamic graphs can be utilized to model many real-world applications like social media analysis in which the connections and entities evolve continuously. Hence, the processing of dynamic graphs is gaining increasing popularity. However, prior dynamic graph processing systems mainly focus on the optimization of graph analytics but overlook graph updating which manages the evolving graph structure and presents a unified view to graph analytics. Since graph updating operates on evolving graphs and involves a large number of irregular memory accesses, it poses a substantial influence on the performance of dynamic graph processing systems. In this work, we observe that graph updating is mainly bottlenecked by a frequent indirect memory access pattern *(*(BAi+offset)). The pattern is inherent to the typical graph updating from the incoming edge stream to the base data store organized with either an adjacent list or a compressed sparse row. With this observation, we propose a novel Prefetcher for Dynamic Graph updating abbreviated as PDG. PDG is a lightweight pipelined instruction-based prefetcher specialized for graph updating and it is also compatible with the irregular memory access pattern BAi widely used in graph analytics. In addition, it leverages a monitor of the instruction queue to decide the appropriate timing of prefetching to make the best use of the cache. According to our experiments, PDG achieves 1.60×, 1.26× and 1.30× performance speedup compared to three representative prefetchers respectively with negligible hardware overhead in graph updating. Xinmiao Zhang 0004, Cheng Liu 0008, Yuanqing Cheng, Lei Zhang 0008, Huawei Li 0001, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | TOTAL: Multi-Corners Timing Optimization Based on Transfer and Active LearningabstractIn modern advanced integrated circuit design, a design normally needs to be progressively optimized until the static timing analysis (STA) of full process corners meets the timing constraints. To improve efficiency, using machine learning to predict the path timings directly in order to reduce the extensive time-consuming SPICE simulations has become a promising technique to approach fast design closure. However, current methods lack both flexibility and reliability to be used in a practical industrial environment. To resolve these challenges, we propose TOTAL, which is constructed using a generalized linear model with latent features to effectively capture knowledge transferred from previous designs and delivers state-of-the-art (SOTA) prediction accuracy that is up to 6.6x improvement over the competitors in terms of mean absolute error (MAE). Most importantly, TOTAL is equipped with a Bayesian decision strategy to actively update uncertain predictions and deliver reliable predictions with accuracy close to 100%, pushing the frontier of the machine-learning-based STA for practical implementation. Wei W. Xing, Rongqi Lu, Zhelong Wang, Ning Xu 0006, Yuanqing Cheng, Weisheng Zhao 0001 |
DAC | 6 |
| 2023 | OPT: Optimal Proposal Transfer for Efficient Yield Optimization for Analog and SRAM CircuitsabstractYield optimization is one of the central challenges in submicrometer integrated circuit manufacture. However, yield optimization is computationally expensive due to intensive yield estimation and intractable optimization processes. In this work, we first reinvent the state-of-the-art all sensitivity adversarial importance sampling (ASAIS) yield optimization from a Laplace approximation perspective, which also reveals its limitations and suggests improvements. We then generalize it with infinite components and discover the key ingredient in yield optimization to be an effective proposal distribution transfer (OPT) procedure, which is captured using conditional normalizing flow (CNF). To deliver a reliable yield optimization pipeline that accounts for the uncertainty due to the lack of data, we propose sequential ensemble, the first empirical uncertainty estimation that enables tractable Bayesian yield optimization without introducing an extra surrogate for the first time. We conduct extensive experiments against five state-of-the-art baselines and show that the proposed method delivers superior performance: a speedup of 1.01x-11.94x (5.57x on average) with higher yield designs, and most importantly, excellent robustness and consistency in all our experiments on analog and SRAM circuits. Guohao Dai 0002, Yuanqing Cheng, Wang Kang 0001, Wei W. Xing |
ICCAD | 3 |
| 2022 | Fault Testing and Diagnosis Techniques for Carbon Nanotube-Based FPGAsabstractAs process technology shrinks into the nanometer-scale, the CMOS-based Field Programmable Gate Arrays (FPGAs) face big challenges in the scalability of performance and power consumption. Multi-walled Carbon Nanotube (MWCNT) serves as a promising candidate for Cu interconnects due to superior conductivity. Moreover, Carbon Nanotube Field Transistor (CNFET) also emerges as a prospective alternative to the conventional CMOS device because of its higher power efficiency and larger noise margin. However, the MWCNT interconnects exhibit significant variations due to an immature fabrication process, leading to delay faults. Furthermore, the non-ideal CNFET fabrication process may generate a few metallic-CNTs (m-CNTs), rendering correlated faulty blocks. In this paper, we propose a ring oscillator (RO) based testing technique to detect delay faults due to the process variations of MWCNT interconnects. In addition, a novel circuit design based on the lookup table (LUT) is applied to speed up the fault testing of CNT-based FPGAs. Finally, we propose a testing algorithm to detect m-CNTs in configurable logic blocks (CLBs). Experimental results show that the test application time for a 6-input LUT can be reduced by 35.49% compared to the conventional testing method, and the proposed algorithm can also achieve a high fault coverage with lower testing overheads. Kangwei Xu, Yuanqing Cheng |
ASP-DAC | 2 |
| 2022 | GIA: A Reusable General Interposer Architecture for Agile Chiplet Integrationabstract2.5D chiplet technology is gaining popularity for the efficiency of integrating multiple heterogeneous dies or chiplets on interposers, and it is also considered an ideal option for agile silicon system design by mitigating the huge design, verification, and manufacturing overhead of monolithic SoCs. Although it significantly reduces development costs by chiplet reuse, the design and fabrication of interposers also introduce additional high non-recurring engineering (NRE) costs and development cycles which might be prohibitive for application-specific designs having low volume. Fuping Li, Ying Wang 0001, Yuanqing Cheng, Yinhe Han 0001, Huawei Li 0001, Xiaowei Li 0001 |
ICCAD | 3 |
| 2022 | Emerging monolithic 3D integration: Opportunities and challenges from the computer system perspective
Yuanqing Cheng, Vasilis F. Pavlidis |
Integr. | 1 |
| 2022 | All-spin PUF: An Area-efficient and Reliable PUF Design with Signature Improvement for Spin-transfer Torque Magnetic Cell-based All-spin CircuitsabstractRecently, spin-transfer torque magnetic cell (STT-mCell) has emerged as a promising spintronic device to be used in Computing-in-Memory (CIM) systems. However, it is challenging to guarantee the hardware security of STT-mCell-based all-spin circuits. In this work, we propose a novel Physical Unclonable Function (PUF) design for the STT-mCell-based all-spin circuit (All-Spin PUF) exploiting the unique manufacturing process variation (PV) on STT-mCell write latency. A methodology is used to select appropriate logic gates in the all-spin chip to generate a unique identification key. A linear feedback shift register (LFSR) initiates the All-Spin PUF and simultaneously generates a 64-bit signature at each clock cycle. Signature generation is stabilized using an automatic write-back technique. In addition, a masking scheme is applied for signature improvement. The uniqueness of the improved signature is 49.61%. With ± 20% supply voltage and 5°C to 105°C temperature variations, the All-Spin PUF shows a strong resiliency. In comparison with state-of-the-art PUFs, our approach can reduce hardware overhead effectively. Finally, the robustness of the All-Spin PUF against emerging modeling attacks is verified as well. Kangwei Xu, Dongrong Zhang, Yuanqing Cheng, Patrick Girard 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2022 | Carbon Nanotube SRAM in 5-nm Technology Node Design, Optimization, and Performance Evaluation - Part I: CNFET Transistor OptimizationabstractIn this article, we propose a carbon nanotube (CNT) field-effect transistor (CNFET)-based static random access memory (SRAM) design at the 5-nm technology node that is optimized based on the tradeoff between performance, stability, and power efficiency. In addition to size optimization, physical model parameters including CNT density, CNT diameter, and CNFET flat band voltage are evaluated and optimized for CNFET SRAM performance improvement. Optimized CNFET SRAM is compared with state-of-the-art 7-nm FinFET SRAM cell based on Arizona State University [ASAP 7-nm FinFET predictive technology models (PTM)] library. We find that the read, write EDPs, and static power of the proposed CNFET SRAM cell are improved by 67.6%, 71.5%, and 43.6%, respectively, compared with the FinFET SRAM cell, with slightly better stability. CNT interconnects both inside and in-between CNFET SRAM cells are considered to compose an all-carbon-based SRAM (ACS) array which will be discussed in the Part II of this article. A 7-nm FinFET SRAM cell with copper interconnects is implemented and used for comparison. Rongmei Chen, Yuanqing Cheng, Souhir Elloumi, Kangwei Xu, Vihar P. Georgiev, Kai Ni 0004, Peter Debacker, A. Asenov, Aida Todri |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | Carbon Nanotube SRAM in 5-nm Technology Node Design, Optimization, and Performance Evaluation - Part II: CNT Interconnect OptimizationabstractThe size and parameter optimization for the 5-nm carbon nanotube field effect transistor (CNFET) static random access memory (SRAM) cell was presented in Part I of this article. Based on that work, we propose a carbon nanotube (CNT) SRAM array composed of the schematically optimized CNFET SRAM and CNT interconnects. We consider the interconnects inside the CNFET SRAM cell composed of metallic single-wall CNT (M-SWCNT) bundles to represent the metal layers 0 and 1 (M0 and M1). We investigate the layout structure of CNFET SRAM cell considering CNFET devices, M-SWCNT interconnects, and metal electrode Palladium with CNT (Pd-CNT) contacts. Two versions of cell layout designs are explored and compared in terms of performance, stability, and power efficiency. Furthermore, we implement a 16 Kbit SRAM array composed of the proposed CNFET SRAM cells, multiwall CNT (MWCNTs) inter-cell interconnects and Pd-CNT contacts. Such an array shows significant advantages, with the read and write overall energy-delay product (EDP), static power consumption, and core area of$0.28\times $,$0.52\times $, and$0.76\times $respectively to 7-nm FinFET-SRAM array with copper interconnects, whereas the read and write static noise margins are 6% and 12% respectively larger than the FinFET counterpart. Rongmei Chen, Yuanqing Cheng, Souhir Elloumi, Kangwei Xu, Vihar P. Georgiev, Kai Ni 0004, Peter Debacker, A. Asenov, Aida Todri |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | A Survey of Test and Reliability Solutions for Magnetic Random Access MemoriesabstractMemories occupy most of the silicon area in nowadays' system-on-chips and contribute to a significant part of system power consumption. Though widely used, nonvolatile Flash memories still suffer from several drawbacks. Magnetic random access memories (MRAMs) have the potential to mitigate most of the Flash shortcomings. Moreover, it is predicted that they could be used for DRAM and SRAM replacement. However, they are prone to manufacturing defects and runtime failures as any other type of memory. This article provides an up-to-date and practical coverage of MRAM test and reliability solutions existing in the literature. After some background on existing MRAM technologies, defectiveness and reliability issues are discussed, as well as functional fault models used for MRAM. This article is dedicated to a summarized description of existing test and reliability improvement methods developed so far for various MRAM technologies. The last part of this article gives some perspectives on this hot topic. Patrick Girard 0001, Yuanqing Cheng, Arnaud Virazel, Wei Zhao 0010, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
Proc. IEEE | 2 |
| 2021 | DOVA PRO: A Dynamic Overwriting Voltage Adjustment Technique for STT-MRAM L1 Cache Considering Dielectric Breakdown EffectabstractAs device integration density increases exponentially as predicted by Moore's law, power consumption becomes a bottleneck for system scaling where leakage power of on-chip cache occupies a large fraction of the total power budget. Spin transfer torque magnetic random access memory (STT-MRAM) is a promising candidate to replace static random access memory (SRAM) as an on-chip last level cache (LLC) due to its ultralow leakage power, high integration density, and nonvolatility. Moreover, with the prevalence of edge computing and Internet-of-Things (IoT) applications, it can be beneficial to build a total nonvolatile cache hierarchy, including the L1 cache. However, building an L1 cache with STT-MRAM still faces severe challenges particularly because reducing its relatively high write latency by increasing write voltage can accelerate oxide breakdown of the MTJ device and threaten the L1 cache lifetime significantly due to intensive accesses. In our previous work, we proposed a dynamic overwriting voltage adjustment (DOVA) technique to deal with this challenge. In this article, we improve this technique by a DOVA promotion (DOVA PRO) technique for the STT-MRAM L1 cache, considering the cache write endurance and performance simultaneously. A high write voltage is used for performance-critical cache lines, while a low write voltage is used for other cache lines to approach an optimal tradeoff between reliability and performance. Experimental results show that the proposed technique DOVA PRO can improve cache performance by 23.5%, on average, compared to the DOVA technique. In the meantime, the average degradation of cache lifetime remains almost unchanged compared with the DOVA technique on average. Furthermore, DOVA PRO can support flexible configurations to achieve various optimization targets, such as higher performance or a longer lifetime. Jinbo Chen 0002, Chengcheng Lu, Patrick Girard 0001, Yuanqing Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2020 | SIP: Boosting Up Graph Computing by Separating the Irregular Property DataabstractGraph analytics is an important class of applications and is one of the cornerstone of big-data workloads. Unfortunately, due to poor data locality in most graph applications, conventional general-purpose computer architectures are unable to perform the best of their processing abilities. The main source of poor locality comes from accessing vertex properties. Upper-level caches cannot hold data blocks long enough due to their limited capacity and the long reuse distance of vertex properties. Moreover, accesses to properties can evict other useful data with good locality, which causes more conflicting misses. In this work, a small cache is added exclusively for the properties to solve this problem. We further enhance this structure with prefetchers to increase the hit rate of properties and improve performance of system. Experimental results show that compared to two state-of-the-art prefetcher and accelerator for graph computing, our proposed architecture achieves 1.13x-2.54x and 1.04x-1.27x performance improvements. In the meanwhile, the energy consumptions can be saved by 6.41%-13.43% and 34.67%-43.92% respectively. Yuanqing Cheng |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | Zero-skew Clock Network Synthesis for Monolithic 3D ICs with Minimum WirelengthabstractClock network synthesis has traditionally been an important step of the physical design process, greatly affecting the performance of ICs. In this paper, we focus on the clock network design process for monolithic 3D (M3D) ICs. Firstly, we investigate the difference between Monolithic Inter-tier Via (MIV) and Through-Silicon Via (TSV) due to the different fabrication process and explore the ramifications of clock network design for monolithic 3D systems. Secondly, we develop a two step clock network synthesis algorithm (M3D-ZST) based on clustering and the deferred-merge embedding algorithm. The proposed algorithm considers the MIV characteristics and constructs a zero-skew clock tree considering wirelength optimization. Furthermore, we apply a look-ahead approach, thereby determining the optimal locations of the merging segments and MIVs such that the wirelength is reduced further (M3D-ZSTLA). Experimental results indicate that M3D-ZST algorithm reduces the total wirelength by 9.7% \textendash\ 19.7%, and reduces power by 9.4% \textendash\ 18.6% compared to the 3D-MMM algorithm over IBM benchmarks. The M3D-ZSTLA algorithm further decreases the total wirelength by about 3%, and reduces the power by about 2%. Vasilis F. Pavlidis, Yuanqing Cheng |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | Write Back Energy Optimization for STT-MRAM-based Last-level Cache with Data Pattern CharacterizationabstractTraditional memory technologies face severe challenges in meeting the ever-increasing power and memory bandwidth requirements for high-performance computing and big-data analyses. Several emerging memory technologies are promising as the replacements of SRAM or DRAM. Among them, STT-MRAM can be used to replace SRAM as the last-level cache (LLC). However, it suffers from high write energy and latency. In this article, we investigate data patterns written from SRAM-based upper-level cache to STT-MRAM-based LLC to explore the write energy reduction potential. Depending on the data layout within a cache line, redundant bits can be identified and eliminated from write back operations to save STT-MRAM write energy. We also propose a dynamic profiling method to accommodate different application characteristics. The extensive simulation results show that write energy can be saved by 37.05% ∼ 38.89% for static profiling and 19.76% ∼ 34.29% for dynamic profiling. Keren Liu, Bi Wu 0002, Weisheng Zhao 0001, Yuanqing Cheng, Ying Wang 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal ConsiderationabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. Spin transfer torque magnetic memory (STT-MRAM) is proposed as a promising solution for the low power cache design due to its high integration density and ultralow leakage power. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM, and observe that the temperature can affect the write delay and energy significantly. Then, we explore the nonuniform cache access (NUCA) design of the chip-multiprocessors with STT-MRAM-based last level cache (LLC). A thermal aware data migration policy, called “Thermosiphon,” which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions dynamically based on the thermal distribution monitored by thermal sensors available on-chip, and adaptively migrates write intensive data among different thermal regions considering the thermal gradient. Compared to the conventional NUCA design, our proposed design can save 41.2% write energy at most and 13.01% on average with negligible hardware overhead. Bi Wu 0002, Pengcheng Dai, Yuanqing Cheng, Ying Wang 0001, Jianlei Yang 0001, Zhaohao Wang, Dijun Liu, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | An Adaptive Thermal-Aware ECC Scheme for Reliable STT-MRAM LLC DesignabstractConsidering the insatiable demand for high-performance computing, on-chip cache capacity increases rapidly. Spin-transfer-torque magnetoresistive random-access memory (STT-MRAM) is a promising cache candidate due to ultralow standby power, high-access speed, and integration density. Unfortunately, when the feature size of magnetic tunnel junction (MTJ) scales down to 1 Xnm, read current approaches write current closely, which may result in read disturbance threatening the reliability of STT-MRAM. Furthermore, the elevating on-chip temperature reduces the thermal stability of STT-MRAM remarkably and aggravates the read disturbance. Error correction code (ECC) is an effective technique to enhance memory reliability. In this paper, we take advantage of the thermal dependence of STT-MRAM and propose a thermally adaptive ECC design, called “Chameleon,” that can adjust the ECC protection strength dynamically to reduce the ECC storage overhead and improve the cache access performance and energy efficiency. Experimental results show that compared to the conservative nonadaptive ECC scheme, our design can improve both cache performance and energy consumption effectively. Bi Wu 0002, Yuanqing Cheng, Ying Wang 0001, Dijun Liu, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | NEAR: A Novel Energy Aware Replacement Policy for STT-MRAM LLCsabstractAs the technology node shrinks, leakage power becomes a bottleneck for processor performance and memory capacity scalings. Spin Torque Transfer Magnetic Random Access Memory (STT-MRAM) has negligible leakage power, fast access speed, high integration density and non-volatility. Therefore, it is a promising candidate for the last level cache design. However, it suffers from high write energy and slow write speed. In the paper, we observe that the traditional cache replacement policy is not optimal when applied to STT-MRAM from the energy consumption perspective. So we propose a novel write energy aware cache replacement policy, which utilizes a MinHash function to identify the similarities between the cache line to be written back and candidates for the replacement. The cache line with the highest similarity is chosen as the victim. In addition, we propose a new metric for cache replacement considering both performance and write energy to improve the replacement policy further. The experimental results show that our proposed policy can reduce write energy by 33.6% on average compared to the state-of-the-art Least Recently Used (LRU) replacement policy with only 0.5% performance penalty and negligible hardware overhead. Yuanqing Cheng, Ying Wang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 2 |
| 2018 | Power Supply Noise Aware Task Scheduling on Homogeneous 3D MPSoCs Considering the Thermal Constraint
Yinglin Zhao, Jianlei Yang 0001, Weisheng Zhao 0001, Aida Todri, Yuanqing Cheng |
J. Comput. Sci. Technol. | 5 |
| 2018 | An Adaptive 3T-3MTJ Memory Cell Design for STT-MRAM-Based LLCsabstractThe STT-MRAM technology is a promising candidate for future on-chip cache memory because of its high density, low standby power, and nonvolatility. As the technology node scales, especially under 40-nm technology node, STT-MRAM cell design becomes a key issue to approach low power consumption, high access performance, and desirable reliability. The conventional 1T-1 magnetic tunnel junction (MTJ) and 2T-2MTJ cell designs cannot address these challenges efficiently. In this paper, we propose a novel 3T-3MTJ cell structure using the advanced perpendicular MTJ (p-MTJ) technology. It can store 2 bits with three MTJs. The differential sensing technique can be used to read out the most significant bit as fast as the 2T-2MTJ design. The sensing latency of 2 bits within the same cell is almost the same as the sensing latency of the 1T-1MTJ cell design. Therefore, the 3T-3MTJ cell can have the advantages of both 2T-2MTJ and 1T-1MTJ cells. Circuit-level simulations show that the proposed 3T-3MTJ cell structure can achieve a desirable tradeoff between storage density, access performance, and energy consumption compared to the prior 1T-1MTJ and 2T-2MTJ cell structures. Additionally, we propose a novel adaptive cache design based on the 3T-3MTJ cell structure, which can work in different modes to satisfy various memory access demands from different applications. Architecture level simulations validate the effectiveness of the proposed cache design. Linuo Xue, Bi Wu 0002, Yuanqing Cheng, Peiyuan Wang, Chando Park, Jimmy J. Kan, Seung-Hyuk Kang, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Building energy-efficient multi-level cell STT-RAM caches with data compressionabstractSpin-transfer torque magnetic random access memory (STT-RAM) technology has emerged as a potential replacement of SRAM in cache design, especially for building large-scale and energy-efficient last level caches. Compared with single-level cell (SLC), multi-level cell (MLC) STT-RAM is expected to double cache capacity and increase system performance. However, the two-step read/write access schemes incur considerable energy consumption and performance degradation. In this paper, we propose two techniques using data compression to optimize MLC STT-RAM cache design. The first technique tries to compress a cache line and fit it into only the soft-bit region of the cells, so that reading or writing this cache line takes only one step which is fast and energy-efficient. We introduce a second technique to increase the cache capacity by enabling the left hard-bit region to store another compressed cache line, which can improve the system performance for memory intensive workloads. The experimental results show that, compared with a conventional MLC STT-RAM last level cache design, our overhead minimized technique reduces the dynamic energy consumption by 38.2% on average with the same system performance, and our capacity augmented technique boosts the system performance by 6.1% with 19.2% dynamic energy saving on average, across the evaluated multi-programmed benchmarks. Liu Liu 0017, Ping Chi, Shuangchen Li, Yuanqing Cheng, Yuan Xie 0001 |
ASP-DAC | 4 |
| 2017 | Thermosiphon: A thermal aware NUCA architecture for write energy reduction of the STT-MRAM based LLCsabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. STT-MRAM (Spin Transfer Torque Magnetic Memory) is proposed as a promising solution for the low power cache design due to its high integration density and ultra-low leakage. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM and observe that the temperature can affect the write delay and energy significantly. Then, we explore the NUCA (Non-Uniform Cache Access) design of the CMPs (Chip-Multi-Processors)with STT-MRAM based LLC (Last Level Cache). A thermal aware data migration policy, called “Thermosiphon”, which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions based on the thermal distribution and adaptively migrate write intensive data considering the temperature gradient among different thermal regions. Compared to the conventional NUCA design, our proposed design can save 22.5% write energy with negligible hardware overhead. Bi Wu 0002, Yuanqing Cheng, Pengcheng Dai, Jianlei Yang 0001, Youguang Zhang, Dijun Liu, Ying Wang 0001, Weisheng Zhao 0001 |
ICCAD | 2 |
| 2017 | STT-RAM Buffer Design for Precision-Tunable General-Purpose Neural Network AcceleratorabstractMultilevel spin toque transfer RAM (STT-RAM) is a suitable storage device for energy-efficient neural network accelerators (NNAs), which relies on large-capacity on-chip memory to support brain-inspired large-scale learning models from conventional artificial neural networks to current popular deep convolutional neural networks. In this paper, we investigate the application of multilevel STT-RAM to general-purpose NNAs. First, the error-resilience feature of neural networks is leveraged to tolerate the read/write reliability issue in multilevel cell STT-RAM using approximate computing. The induced read/write failures at the expense of higher storage density can be effectively masked by a wide spectrum of NN applications with intrinsic forgiveness. Second, we present a precision-tunable STT-RAM buffer for the popular general-purpose NNA. The targeted STT-RAM memory design is able to transform between multiple working modes and adaptable to meet the varying quality constraint of approximate applications. Lastly, the reconfigurable STT-RAM buffer not only enables precision scaling in NNA but also provides adaptiveness to the demand for different learning models with distinct working-set sizes. Particularly, we demonstrate the concept of capacity/precision-tunable STT-RAM memory with the emerging reconfigurable deep NNA and elaborate on the data mapping and storage mode switching policy in STT-RAM memory to achieve the best energy efficiency of approximate computing. Lili Song, Ying Wang 0001, Yinhe Han 0001, Huawei Li 0001, Yuanqing Cheng, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Architecture design with STT-RAM: Opportunities and challengesabstractThe emerging spin-transfer torque magnetic random-access memory (STT-RAM) has attracted a lot of interest from both academia and industry in recent years. It has been considered as a promising replacement of SRAM and DRAM in the cache and memory system design thanks to many advantages, including non-volatility, low leakage power, SRAM comparable read performance and read energy consumption, higher density than SRAM, better scalability than conventional CMOS technologies, and good CMOS compatibility. However, the disadvantages of STT-RAM, such as higher write energy and longer write latency than SRAM, also bring design challenges. This paper introduces state-of-the-art architectural approaches to adopt STT-RAM in the cache and memory system design by taking advantage of the opportunities brought by STT-RAM as well as overcoming the challenges. Ping Chi, Shuangchen Li, Yuanqing Cheng, Seung-Hyuk Kang, Yuan Xie 0001 |
ASP-DAC | 3 |
| 2016 | ODESY: a novel 3T-3MTJ cell design with optimized area DEnsity, scalability and latencYabstractThe STT-RAM (Spin-Transfer Torque Magnetic RAM) technology is a promising candidate for cache memory because of its high density, low standy-power, and non-volatility. As technology scales, especially under 40nm technology node, the read disturbance becomes severe since the read current approaches closely to the switching current. In addition, the read latency and access performance degrade significantly as well. The conventional 1T-1MTJ and 2T-2MTJ cell designs cannot address these challenges efficiently. In this paper, we propose a novel 3T-3MTJ cell structure using the advanced perpendicular MTJ technology. This memory cell has higher storage density and better performance, and is particularly suitable for the deeply scaled technology node. A two-stage sensing scheme is also proposed to facilitate the read operation of the 3T-3MTJ cell design. Circuit-level and architecture-level simulations show that the proposed 3T-3MTJ cell structure can achieve a better tradeoff between storage density, access performance, energy consumption, and reliability compared to the prior 1T-1MTJ and 2T-2MTJ cell structures. Linuo Xue, Yuanqing Cheng, Jianlei Yang 0001, Peiyuan Wang, Yuan Xie 0001 |
ICCAD | 2 |
| 2016 | AES design improvement towards information safetyabstractWith the rapid development and globalization of semiconductor design and fabrication, integrated circuit (IC) is becoming more vulnerable to malicious modification called hardware Trojan. As Advanced Encryption Standard (AES) core has been widely used in security critical applications, it can easily become a target of Hardware Trojan. In this paper, 9 potential AES hardware Trojans, which cover Trojan types leaking key or plain text are demonstrated. Then three protections of varying strengths against the proposed Trojans, including combination, reorder, and reconfiguration logic insertion, are presented. The proposed protections are shown to effectively disfunction the 9 types of potential Trojans. Therefore, the information security of AES is significantly improved. Xiaoxiao Wang 0001, Xiaoying Zhao, Yuanqing Cheng, Donglin Su, Aixin Chen, Qihang Shi, Mark Tehranipoor |
ISCAS | 4 |
| 2016 | An efficient all-digital IR-Drop Alarmer for DVFS-based SoCabstractFor 40nm and below technologies, billions of transistors can be integrated into a single chip. Meanwhile, the operation frequency has reached over Giga Hertz. In this case, highly synchronized switching activities can induce significant current, which leads to IR-drop. Excessive IR-drop can cause timing failure, abnormal reset, or disruption of data processing. As a result, dynamic voltage and frequency scaling (DVFS) system implemented effective adaptation strategies are widely used by SoCs to mitigate IR-Drop noise and stabilize performance. As the basis of DVFS action, economic and accurate IR-drop monitors are in great need. This paper presents a novel and efficient IR-Drop Alarmer, which can cooperate with the DVFS system for fast IR-drop adaptation. The IR-drop alarming threshold of the proposed sensor is configurable between 45mV to 120mV. Considering a 1.1ns width IR-drop noise, the IR noise sampling window can be as small as 0.125ns, with alarming duration error rate less than 6.8% for 97% of the Monte Carlo samples considering process variations. Furthermore, the proposed alarmer is composed by all-digital standard gates without an y high frequency sampling clock, which is of low area overhead and power consumption. Liting Yu, Xiaoxiao Wang 0001, Yuanqing Cheng, Xiaoying Zhao, Pengyuan Jiao, Aixin Chen, Donglin Su, LeRoy Winemberg, Mehdi Sadi, Mark Tehranipoor |
ISCAS | 3 |
| 2016 | Quantitative evaluation of reliability and performance for STT-MRAMabstractDue to its non-volatility, high access speed, ultra low power consumption and unlimited writing/reading cycles, STT-MRAM (Spin Transfer Torque Magnetic Random Access Memory) has emerged as the most promising candidate for the next generation universal memory. However, the process of commercialization of STT-MRAM is hampered by its poor reliability. Generally, these reliability issues are caused by the PVT (Process Variations, Voltage, and Temperature) of both MTJ (Magnetic Tunneling Junction) and transistor. Mitigation and alleviating the impacts of the intrinsic properties and PVT on STT-MRAM is a challenging work. This paper discusses the errors occurring in STT-MRAM resulting from its poor reliability, and analyzes the causes of such errors. To obtain a quantitative assessment of PVT impact on STT-MRAM reliability, we investigate three aspects: writing/reading operation error rate, power consumption and access delay of a single cell. This study is carried out on Cadence platform for 45 nm technology node and the PMA (Perpendicular Magnetic Anisotropy) MTJ model used in the investigation comes from SP INLIB. These quantitative information would be helpful for designing reliability enhancing strategies of STT-MRAM. Liuyang Zhang, Aida Todri, Wang Kang 0001, Youguang Zhang, Lionel Torres, Yuanqing Cheng, Weisheng Zhao 0001 |
ISCAS | 6 |
| 2016 | Radiation-Induced Soft Error Analysis of STT-MRAM: A Device to Circuit ApproachabstractSpin-transfer torque magnetic random access memory (STT-MRAM) is a promising emerging memory technology due to its various advantageous features such as scalability, nonvolatility, density, endurance, and fast speed. However, the reliability of STT-MRAM is severely impacted by environmental disturbances because radiation strike on the access transistor could introduce potential write and read failures for 1T1MTJ cells. In this paper, a comprehensive approach is proposed to evaluate the radiation-induced soft errors spanning from device modeling to circuit level analysis. The simulation based on 3-D metal-oxide-semiconductor transistor modeling is first performed to capture the radiation-induced transient current pulse. Then a compact switching model of magnetic tunneling junction (MTJ) is developed to analyze the various mechanisms of STT-MRAM write failures. The probability of failure of 1T1MTJ is characterized and built as look-up-tables. This approach enables designers to consider the effect of different factors such as radiation strength, write current magnitude and duration time on soft error rate of STT-MRAM memory arrays. Meanwhile, comprehensive write and sense circuits are evaluated for bit error rate analysis under random radiation effects and transistors process variation, which is critical for performance optimization of practical STT-MRAM read and sense circuits. Jianlei Yang 0001, Peiyuan Wang, Yaojun Zhang, Yuanqing Cheng, Weisheng Zhao 0001, Yiran Chen 0001, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Temperature Impact Analysis and Access Reliability Enhancement for 1T1MTJ STT-RAMabstractSpin-transfer torque magnetic random access memory (STT-RAM) is a promising and emerging technology due to its many advantageous features such as scalability, nonvolatility, density, endurance, and fast access speed. However, the operation of STT-RAM is severely affected by environmental factors such as process variations and temperature. As the temperature rockets up in modern computing systems, it is highly desirable to understand thermal impact on STT-RAM operations and reliability. In this paper, a thermal-aware MTJ model, calibrated and validated by experimental measurements, is proposed as the basis for thoroughly thermal aware analysis of a 1T1MTJ STT-RAM cell structure. Using this model, we investigate temperature effect on memory cell access behavior in terms of access latency, energy, and reliability on a 45-nm technology node. Thermal impact on a more advanced 11-nm technology node is also evaluated in the paper. Additionally, we propose a thermal-aware design for STT-RAM sensing circuit using a body-biasing technique, which can enlarge read margin dramatically to enhance read reliability under temperature variations. Moreover, our proposed technique can suppress read disturbance effectively as well. Experimental results show that our proposed sensing circuit can enlarge read margin by 2.47× when reading “0” and 3.15× when reading “1,” and reduce read disturbance error rate by 55.6% on average. Bi Wu 0002, Yuanqing Cheng, Jianlei Yang 0001, Aida Todri, Weisheng Zhao 0001 |
IEEE Trans. Reliab. | 2 |
| 2016 | Alleviating Through-Silicon-Via Electromigration for 3-D Integrated Circuits Taking Advantage of Self-Healing EffectabstractThree-dimensional integration is considered to be a promising technology to tackle the global interconnect scaling problem for terascale integrated circuits (ICs). Three-dimensional ICs typically employ through-silicon-vias (TSVs) to vertically connect planar circuits. Due to its immature fabrication process, several defects, such as void, misalignment, and dust contamination, may be introduced. These defects can significantly increase current densities within TSVs and cause severe electromigration (EM) effects, which can degrade the reliability of 3-D ICs considerably. In this paper, we propose an effective framework to mitigate EM effect of the defective TSV. At first, we analyze various possible TSV defects and their impacts on EM reliability. Based on the observation that EM can be significantly alleviated by self-healing effect, we design an EM mitigation module to protect defective TSVs from EM. To guarantee EM mitigation efficiency, we propose two defective TSV protection schemes, i.e., neighbor sharing and global sharing. Experimental results show that the global-sharing scheme performs the best and can improve the EM mean time to failure by more than 70× on average with only 0.7% area overhead and less than 0.5% performance degradation compared with naked design without any EM protection. Yuanqing Cheng, Aida Todri, Jianlei Yang 0001, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | A Study of 3-D Power Delivery Networks With Multiple Clock DomainsabstractOngoing advancements in 3-D manufacturing are enabling 3-D ICs to contain several processing cores, hardware accelerators, and dedicated peripherals. Most of these functional units operate with independent clock frequencies for power management reasons or simply for being hard intellectual properties. Thus, as diverse and heterogeneous circuits can be implemented on a 3-D IC, it also leads to the use of multiple clock domains. While these domains allow many functional units to run in parallel to exploit 3-D potentials, they also introduce power delivery challenges. This paper proposes an efficient analysis for assessing the worst case power supply noise on 3-D power delivery networks (PDNs) with multiple clock domains. This paper discusses power and thermal integrity issues that arise from multiple clock domains that share the same 3-D global PDN. We first examine power supply noise distribution on each tier and investigate scenarios that lead to worst case noise. Thermal analyses are also performed and heat distribution among clock domains and tiers is examined. In addition, the impact of clock domain structure and frequency on the overall power supply noise and temperature distribution has been quantified. Experiments show that the multiclock domains can induce excessive noise and the through-silicon-vias can contribute to power supply noise and heat transfer among tiers. This paper presents a summary of guidelines for modeling, analyzing, and exploring a design of reliable 3-D PDNs with multiple clock domains. Aida Todri, Yuanqing Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | PSI Conscious Write Scheduling: Architectural Support for Reliable Power Delivery in 3-D Die-Stacked PCMabstractIn 3-D-stacked memory chips, the problem of power supply integrity (PSI) is aggravating due to the additional through-silicon-via resistance and the higher current density in 3-D power delivery network. In particular, for the 3-D phase-change memory (PCM) well known for its high-amplitude programming current, IR-drop violation poses a serious threat that enforces a strict guard band of requesting concurrence, and consequently reduces the write throughput. This paper presents the implication of an IR-drop phenomenon in a 3-D PCM cube, and investigates IR-drop's impacts on write management in the PCM. From the obtained SPICE simulation results, we find that the issued writes have to meet the IR-drop constraint to be reliably processed, and then propose a PSI conscious write scheduler to improve the write performance within the constraint of the IR-drops and the power budget in the 3-D PCM cube. First, a Bloom-filter-based method is proposed to avoid the invalid write decisions for the PCM. Second, to support fine-grained write management in the cutting-edge PCM, we develop an inexpensive approach, weighted token assignment (WTA), to filter out PSI-unsafe write decisions by employing a support vector machine-based learning model. Last, a write reordering policy is proposed to cooperate with WTA and optimize the total write throughput for better memory performance. In the simulated hybrid main memory composed of both dynamic random access memory and 3-D PCM, the proposed scheduler significantly improves the write throughput. Ying Wang 0001, Yinhe Han 0001, Huawei Li 0001, Lei Zhang 0008, Yuanqing Cheng, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | A case of precision-tunable STT-RAM memory design for approximate neural networkabstractMulti-level STT-RAM cell is able to boost the memory density at the expense of read/write reliability. However, the induced data integrity issue in STT-RAM memory can be effectively masked by a wide spectrum of applications with intrinsic forgiveness, which belong to the specific domain such as multimedia, synthesis and mining. In this work, we leverage the reconfigurable capability of MLC STT-RAM to provide variable-precision data storage for popular machine learning architectures. The targeted STT-RAM memory design is able to transform between multiple work modes and adaptable to meet the varying quality constraint of approximate applications. Particularly, we demonstrate the concept of precision-tunable STT-RAM memory with the emerging Convolution Neural Network accelerators and elaborate on the data mapping policy in STT-RAM memory to achieve the best energy-efficiency. Ying Wang 0001, Lili Song, Yinhe Han 0001, Yuanqing Cheng, Huawei Li 0001, Xiaowei Li 0001 |
ISCAS | 4 |
| 2015 | A body-biasing of readout circuit for STT-RAM with improved thermal reliabilityabstractAs the integration density rockets up for contemporary VLSI circuits, power consumption limits the scalability of technology advancement of CMOS. Spin transfer torque-magnetic random access memory (STT-MRAM), as one of the emerging non-CMOS technologies, has the promising prospect of low standby power, fast access speed and compatibility with the CMOS fabrication process. However, with the technology node scaling down, typical 1 Transistor-1 Magnetic Tunnel Junction (1T-1MTJ) STT-RAM cell suffers from severe reliability challenges, especially for read operation under temperature fluctuation. In this paper, we quantitatively analyze the temperature effect on read reliability of STT-RAM cell and propose a novel body-biasing feedback readout circuit design to improve the read sensing margin under different temperatures. The experiments based on 40nm CMOS technology and MTJ compact model validate the effectiveness of the proposed method. The improved sensing margin also permits a smaller sensing current for reading such that higher read energy efficiency can be achieved. Lun Yang, Yuanqing Cheng, Yuhao Wang 0002, Hao Yu 0001, Weisheng Zhao 0001, Aida Todri |
ISCAS | 2 |
| 2014 | Power supply noise-aware workload assignments for homogeneous 3D MPSoCs with thermal considerationabstractIn order to improve performance and reduce cost, multi-processor system on chip (MPSoC) is increasingly becoming attractive. At the same time, 3D integration emerges as a promising technology for high density integration. 3D homogeneous MPSoCs combine the benefits of both. However, high current demand and large on-chip switching activity variations introduce severe power supply noises (PSN) for 3D MPSoCs, which can increase critical path delay, and degrade chip performance and reliability. Meanwhile, thermal gradient should also be considered for 3D MPSoCs to avoid hot spots. In the paper, we investigate the PSN effects of different workloads and propose an effective PSN estimation method. Then, a heuristic workload assignment algorithm is proposed to suppress PSN under the given thermal constraint. The experimental results show that PSNs can be reduced significantly compared with thermal-balanced workload assignment scheme, and the system performance can be improved as well. Yuanqing Cheng, Aida Todri, Alberto Bosio, Luigi Dilillo, Patrick Girard 0001, Arnaud Virazel |
ASP-DAC | 1 |
| 2013 | TSV Minimization for Circuit - Partitioned 3D SoC Test Wrapper Design
Yuanqing Cheng, Lei Zhang 0008, Yinhe Han 0001, Xiaowei Li 0001 |
J. Comput. Sci. Technol. | 1 |
| 2013 | Thermal-Constrained Task Allocation for Interconnect Energy Reduction in 3-D Homogeneous MPSoCsabstract3-D technology that stacks silicon dies with through silicon vias (TSVs) is a promising solution to overcome the interconnect scaling problem in giga-scale integrated circuits (ICs). Thermal dissipation is a major challenge for 3-D integration and prior thermal-balanced task scheduling methods for 3-D multiprocessor system-on-chips (MPSoCs) typically balance power gradient across vertical stacks based on the assumption of strong thermal correlation among processing cores within a stack. On the other hand, 3-D MPSoCs typically employ network-on-chip (NoC) as the communication infrastructure which consumes a large portion of the energy budget. As TSVs consume much less energy than horizontal links in 3-D MPSoCs when transmitting the same amount data due to the reduced interconnect distance between vertical adjacent cores, it motivates to allocate heavily communicating tasks within the same vertical stack as much as possible, and thus traffic is restricted in the third dimension to reduce interconnect energy. However, aggregating active tasks within the same stack probably exacerbates the power density and result in hot spots. In this paper, we explore the tradeoff between thermal and interconnect energy when allocating tasks in 3-D Homogeneous MPSoCs, and propose an efficient heuristic. Experimental results show that the proposed technique can reduce interconnect energy by more than 25% on average with almost the same peak temperature when compared with prior thermal-balanced solutions. Yuanqing Cheng, Lei Zhang 0008, Yinhe Han 0001, Xiaowei Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Wrapper Chain Design for Testing TSVs Minimization in Circuit-Partitioned 3D SoCabstractThree dimensional (3D) System-on-Chips (SoCs) that typically employ through-silicon vias (TSVs) as vertical interconnects, emerge as a promising solution to continue Moore's law. Whereas, it also brings challenging problems, one of which is the test wrapper chain design and optimization, especially for circuit-partitioned 3D SoCs in which scan chains can cross among layers. Test time is the primary goal for wrapper chain design, both for 2D and 3D SoCs. The 3D SoC wrapper chain design problem can be converted into the well-studied2D one by projecting wrapper chain components of all layers to one virtual layer. Thereafter, we can leverage 2D optimization algorithms to determine the composition of wrapper chains and thus guarantee minimal testing time for 3D SoCs. One specific thing for circuit-partitioned 3D SoCs is that TSVs are needed to connect cross-layer wrapper structures to form the wrapper chains. As TSVs occupy planar chip area and will aggravate the routing congestion problem, it is necessary to reduce TSVs for test purpose as much as possible. In this work, we observe that by varying the connection orders of wrapper chain components, e.g., scan chains and I/O cells, the TSVs consumed vary significantly. Based on the above, we formulate this problem and propose novel heuristic to tackle it. Experimental results show that the proposed solution can save on average 33.2% amount of TSVs when compared to a prior intuitive method. Yuanqing Cheng, Lei Zhang 0008, Yinhe Han 0001, Xiaowei Li 0001 |
Asian Test Symposium | 1 |