Yanan Sun 0003

dblp:44/8711-3 · DBLP profile ↗
← Back
43ranked-venue papers
8as first author
32since 2021 · last 2026
0000-0001-8281-9121ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 8 first-author · 30 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DBP-CIM: Energy-Efficient 8T SRAM-Based Diagonal-Block Parallel Computing-in-Memory With Compact Data Layout for Arithmetic Operations
abstract
In this work, an energy-efficient bit-parallel static random-access memory (SRAM)-based computing-in-memory (SRAM-CIM) is proposed for general-purpose in-memory arithmetic operations to adapt diverse computing tasks. A compact diagonal-block parallel (DBP) mapping scheme and a novel arithmetic flow are proposed to address the hardware underutilization issue in conventional two-sided stationary bit-parallel CIM architectures. Specifically, the DBP mapping method is implemented to enhance the throughput by reorganizing the intermediate and final results into diagonal memory blocks, effectively reducing the vacant CIM cells caused by the dynamic bit width during computing. In addition, the proposed hardware-efficient arithmetic flows employ: 1) a pipelined ADD scheme to reduce the critical path latency in near-memory computing units; and 2) shift-based arithmetic operations that halve the hardware resources required for multiplication and division while reducing energy consumption. The post-layout simulations on 28-nm CMOS technology show that the proposed DBP-CIM achieves higher energy efficiency and throughput for general-purpose arithmetic operations, compared with state-of-the-art works. Furthermore, evaluations on the general-purpose benchmarks demonstrate that the DBP-CIM reduces energy consumption and computing cycles by up to 55.9% and 60.9%, compared to the conventional bit-parallel CIM.
Dengfeng Wang, Chengjun Chang, Weifeng He, Guanghui He 0002, Yanan Sun 0003
IEEE Trans. Very Large Scale Integr. Syst.5
2026 A Heterogeneous CIM Architecture With Splittable Nonvolatile Computing-in-SRAM Cell Pairs Enabling Efficient On-Chip Neural Network Inference
abstract
Heterogeneous computing-in-memory (CIM) offers a promising solution for efficient neural network (NN) accelerations by leveraging the characteristics of different types of memories. However, the previous heterogeneous CIM either requires additional data transfer between different types of isolated memories or solely relies on in situ embedded single-level (SL) or three-level (TL) nonvolatile memories (NVMs), making it difficult to trade off between storage density and robustness benefits. In this article, a new heterogeneous CIM architecture (NVS-SPT) with enhanced storage density and inference robustness is proposed to enable full on-chip acceleration of practical-scale NNs while scalable to larger models. A splittable nonvolatile computing-in-static random access memory (SRAM) cell pairs (nvS2RAM-CIM) is proposed with hybrid in situ embedded SL and TL resistive random access memory (ReRAM) groups, allowing flexible configuration as split or linked state to enhance storage density and restore yield. A layerwise hybrid-coding search (LHCS) algorithm with bitwise and tritwise data-aware mapping (BTM) method is proposed to determine the optimal weight coding patterns with high array utilizations. In addition, a merged hybrid-coding block (MHCB) generation scheme is employed to enable high computing parallelism by merging the dense computing patterns. The proposed NVS-SPT demonstrates up to$4.2\times $higher storage density compared with previous heterogeneous CIM with pure SL-ReRAMs and achieves up to 44.7% enhanced NN accuracy, compared with previous unified ternary coding. Furthermore, the proposed NVS-SPT exhibits up to$1.72\times $and$1.44\times $enhanced energy efficiency with$3.10\times $and$1.42\times $higher computing density, compared with previous heterogeneous CIM based on pure SL- or TL-ReRAMs, respectively.
Dengfeng Wang, Liukai Xu, Weifeng He, Guanghui He 0002, Xueqing Li 0002, Yanan Sun 0003
IEEE Trans. Very Large Scale Integr. Syst.7
2025 SAGA: A Memory-Efficient Accelerator for GANN Construction via Harnessing Vertex Similarity
abstract
Graph-traversal-based Approximate Nearest Neighbor (GANN) search and construction have become key retrieval techniques in various domains, such as recommendation systems and social networks. However, deploying GANN in real-world scenarios faces significant challenges, as high-dimensional vertices within the graph can lead to intensive memory demands. Although architectures like NDSearch have been proposed to accelerate GANN search, they are hard to deploy for GANN construction, as their pre-processing methods introduce massive overhead in dynamic graphs. In this paper, given the observation that neighboring vertices in a dynamic graph exhibit feature similarity, we propose SAGA, the first accelerator that alleviates memory bound in GANN construction. To capture this similarity, we directly leverage the first step of construction to gather vertices with the same starting point into a cluster to minimize the similarity detection overhead. Next, we decompose vertices into key and non-key ones, where their deltas fall in a narrow range, which is suitable to be quantized to lower bit widths. Building upon this approach, we design a specialized architecture, which efficiently implements the GANN construction by twolevel scheduling and a mixed-precision supported bit-serial unit. Through comprehensive evaluation, we demonstrate that SAGA can achieve an average speedup of $9.30 \times 4.87 \times 4.15 \times$ and $35.46 \times 7.60 \times 5.15 \times$ energy savings over CPU, GPU and NDSearch, respectively, while retaining task accuracy.
Xueyuan Liu 0001, Chunyu Qi, Yuanzheng Yao, Yanan Sun 0003, Xiaoyao Liang, Zhuoran Song
DAC5
2025 KVO-LLM: Boosting Long-Context Generation Throughput for Batched LLM Inference
abstract
With the widespread deployment of long-context large language models (LLMs), efficient and high-quality generation is becoming increasingly important. Modern LLMs employ batching and key-value (KV) cache to improve generation throughput and quality. However, as the context length and batch size rise drastically, the KV cache incurs extreme external memory access (EMA) issues. Recent LLM accelerators face substantial processing element (PE) under-utilization due to the low arithmetic intensity of attention with KV cache, while existing KV cache compression algorithms struggle with hardware inefficiency or significant accuracy degradation. To address these issues, an algorithm-architecture co-optimization, KVO-LLM, is proposed for long-context batched LLM generation. At the algorithm level, we propose a KV cache quantization-aware pruning method that first adopts salient-token-aware quantization and then prunes KV channels and tokens by attention guided pruning based on salient tokens identified during quantization. Achieving substantial savings on hardware overhead, our algorithm reduces the EMA of KV cache over 91% with significant accuracy advantages compared to previous KV cache compression algorithms. At the architecture level, we propose a multi-core jointly optimized accelerator that adopts operator fusion and cross-batch interleaving strategy, maximizing PE and DRAM bandwidth utilization. Compared to the state-of-the-art LLM accelerators, KVO-LLM improves generation throughput by up to $7.32 \times$, and attains $5.52 \sim 8.38 \times$ better energy efficiency.
Dongxu Lyu, Gang Wang 0063, Wenjie Li 0003, Jianfei Jiang 0001, Yanan Sun 0003, Guanghui He 0002
DAC8
2025 BitPattern: Enabling Efficient Bit-Serial Acceleration of Deep Neural Networks through Bit-Pattern Pruning
abstract
Bit-serial computation shows promise for accelerating deep neural networks (DNNs) by exploiting inherent bit sparsity. However, the original unstructured bit sparsity poses two major challenges for existing bit-serial accelerators (BSA): (1) workload imbalance from irregular bit distribution, and (2) inefficient memory access due to unpredictable non-zero bit locations. To address these issues, this paper proposes BitPattern, an algorithm/hardware co-design to efficiently accelerate bitserial computation through bit-pattern pruning. At the algorithm level, we employ bit-pattern pruning to identify optimal combinations of predefined patterns and apply compression encoding to minimize weight storage. We further devise a pattern-similaritybased merging method to balance the bit-serial workload. At the hardware level, we co-design a bit-serial accelerator with a dedicated bit-pattern decoder and PE to leverage the potential of structured bit-pattern sparsity. The evaluation on several deep learning benchmarks shows that BitPattern can achieve $1.72 \times$ memory reduction with negligible accuracy loss, and up to $2.11 \times$ speedup and $1.86 \times$ energy saving compared to state-of-the-art bit-serial accelerators.
Gang Wang 0063, Wenjie Li 0003, Dongxu Lyu, Yanan Sun 0003, Jianfei Jiang 0001, Guanghui He 0002
DAC6
2025 VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible Accelerator
abstract
Large Language Models (LLMs) excel in natural language processing tasks but pose significant computational and memory challenges for edge deployment due to their intensive resource demands. This work addresses the efficiency of LLM inference by algorithm-hardwaredataflow tri-optimizations. We propose a novel voting-based KV cache eviction algorithm, balancing hardware efficiency and algorithm accuracy by adaptively identifying unimportant kv vectors. From a dataflow perspective, we introduce a flexible-product dataflow and a runtime reconfigurable PE array for matrix-vector multiplication. The proposed approach effectively handles the diverse dimensional requirements and solves the challenges of incrementally varying sequence lengths. Additionally, an element-serial scheduling scheme is proposed for nonlinear operations, such as softmax and layer normalization (layernorm). Results demonstrate a substantial reduction in latency, accompanied by a significant decrease in hardware complexity, from $O(N)$ to $O(1)$. The proposed solution is realized in a custom-designed accelerator, VEDA, which outperforms existing hardware platforms. This research represents a significant advancement in LLM inference on resource-constrained edge devices, facilitating real-time processing, enhancing data privacy, and enabling model customization.
Zhican Wang, Hongxiang Fan, Haroon Waris, Gang Wang 0063, Jianfei Jiang 0001, Yanan Sun 0003, Guanghui He 0002
DAC7
2025 HPIM-NoC: A Priori-Knowledge-Based Optimization Framework for Heterogeneous PIM-Based NoCs
abstract
Network-on-Chip (NoC) accelerators with heterogeneous Processing-in-Memory (PIM) cores achieve superior performance than homogeneous ones for neural networks. Dedicated simulators and architecture search frameworks are pivotal for obtaining performance, power, and area (PPA) metrics, as well as guiding the design process. However, existing simulators are primarily designed for homogeneous NoC and lack support for simulating heterogeneous PIM-based NoC architectures. Besides, current search frameworks for heterogeneous NoC architectures only focus on workload allocation and mapping strategies, failing to explore heterogeneous PIM configurations in a larger design space. In this work, we propose HPIM-NoC, a joint simulation and search framework for heterogeneous PIM-based NoC architectures. HPIM-NoC not only supports the simulation of heterogeneous PIM cores, but also provides more accurate latency results by introducing NoC transmission delays and pipelines in co-simulation. HPIM-NoC implements a three-stage heterogeneous search process based on priori knowledge and employs a specific simulated annealing algorithm tailored for heterogeneous architecture search. The search process is accelerated by precomputing core PPA metrics and reducing NoC simulation frequency. In addition, the framework integrates a customized layout algorithm to optimize the placement of heterogeneous NoC, minimizing communication latency and overall area. Experimental results on various neural networks demonstrate that HPIM-NoC can quickly find near-optimal configurations within a limited time. The proposed acceleration method reduces the search time of HPIM-NoC by $2.12 \times$, $2.17 \times$, and $2.96 \times$, respectively. Compared to homogeneous architectures, the Fusions of Metrics (FoMs) of heterogeneous PIM-based NoC architectures found by HPIM-NoC are reduced by $\mathbf{1. 1 8 \%, ~} \mathbf{1 6. 9 4 \%}$, and $\mathbf{3 7. 4 1 \%}$ for ResNet-18 under three settings, respectively.
Shuai Yuan 0016, Angxin Cai, Qiushi Lin, Guoxing Wang, Yu Wang 0002, Zhenhua Zhu 0002, Yanan Sun 0003
DAC7
2025 KeeA*: Epistemic Exploratory A* Search via Knowledge Calibration
abstract
In recent years, neural network-guided heuristic search algorithms, such as Monte-Carlo tree search and A$^\*$ search, have achieved significant advancements across diverse practical applications. Due to the challenges stemming from high state-space complexity, sparse training datasets, and incomplete environmental modeling, heuristic estimations manifest uncontrolled inherent biases towards the actual expected evaluations, thereby compromising the decision-making quality of search algorithms. Sampling exploration enhanced A$^\*$ (SeeA$^\*$) was proposed to improve the efficiency of A$^\*$ search by constructing an dynamic candidate subset through random sampling, from which the expanded node was selected. However, uniform sampling strategy utilized by SeeA$^\*$ facilitates exploration exclusively through the injection of randomness, which completely neglects the heuristic knowledge relevant to open nodes. Moreover, the theoretical support of cluster sampling remains ambiguous. Despite the existence of potential biases, heuristic estimations still encapsulate certain valuable information. In this paper, epistemic exploratory A$^\*$ search (KeeA$^\*$) is proposed to integrate heuristic knowledge for calibrating the sampling process. We first theoretically demonstrate that SeeA$^\*$ with cluster sampling outperforms uniform sampling due to the distribution-aware selection with higher variance. Building on this insight, cluster scouting and path-aware sampling are introduced in KeeA$^\*$ to further exploit heuristic knowledge to increase the sampling mean and variance, respectively, thereby generating higher-quality extreme candidates and enhancing overall decision-making performance. Finally, empirical results on retrosynthetic planning and logic synthesis demonstrate superior performance of KeeA$^*$ compared to state-of-the-art heuristic search algorithms.
Dengwei Zhao, Shikui Tu, Yanan Sun 0003, Lei Xu 0001
NeurIPS3
2025 Hop-CIM: An all-digital two-level approximate SRAM-CIM macro for high energy-efficient HNN acceleration with data-aware early exit and column-wise partial-sum reuse
Shunqin Cai, Liukai Xu, Dengfeng Wang, Keqing Ouyang, Weizhong Wu, Zhi Li 0058, Yanan Sun 0003
Integr.10
2025 HyCTor: A Hybrid CNN-Transformer Network Accelerator With Flexible Weight/Output Stationary Dataflow and Multicore Extension
abstract
Hybrid convolutional neural network (CNN) and Transformer networks are emerging in computer vision, combining convolutional, linear, and attention layers to achieve high accuracies with moderate model sizes. Developing the accelerators for hybrid networks is pivotal to simultaneously optimize the static matrix multiplication (MM) in convolutional and linear layers, as well as dynamic MM in attention layers. However, the existing accelerators are primarily designed for either CNNs or Transformers, resulting in increased data movement to support dynamic MM and potential under-utilization of hardware for static MM. To enhance computational performance and energy efficiency for hybrid networks, we propose HyCTor, an accelerator featuring flexible output-stationary (OS) and weight-stationary (WS) dataflows, along with a multicore extension for higher throughput. The parallel array of HyCTor supports interlayer slicing and intralayer splicing to improve the utilization for static MM, and enables seamless switching between OS and WS dataflow to minimize the data movement in dynamic MM. By leveraging structured sparsity in OS dataflow and unstructured sparsity in WS dataflow, the computational efficiency is further boosted for each layer through flexible dataflow selection based on the sparsity ratio. Besides, a novel QuadLoop-mesh topology is proposed to address the complex data dependencies in hybrid networks and minimize data transmission distances in the multicore HyCTor. Experimental results on ResNet-18, ViT-B, and TransIAR-AF show that the proposed single-core HyCTor achieves$1.83\times $,$1.65\times $, and$2.41\times $speedup than state-of-the-art (SOTA) accelerators with 100% utilization rate in most layers, and$3.82\times $–$38.5\times $speedup than RTX4090 GPU. The energy efficiency of HyCTor is improved by$1.81\times $–$8.77\times $compared with SOTA accelerators. Moreover, the 4-core HyCTor achieves speedups of$3.32\times $,$2.58\times $, and$2.91\times $, while the 16-core HyCTor achieves speedups of$7.05\times $,$4.05\times $, and$9.64\times $compared to 1-core HyCTor on three networks.
Shuai Yuan 0016, Weifeng He, Zhenhua Zhu 0002, Fangxin Liu, Zhuoran Song, Guohao Dai 0001, Guanghui He 0002, Yanan Sun 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2025 Robust Monolithic 3D Carbon-Based Computing-in-SRAM With Variation-Aware Bit-Wise Data-Mapping for High-Performance and Integration Density
abstract
Bit-serial computing-in memory with SRAM cells (SRAM-CIM) enables a full set of integer and floating-point arithmetic operations and various data-intensive computations. Carbon nanotube field-effect transistors (CN-MOSFETs) with high scalability, energy-efficiency, and low process thermal budget are attractive to realize high-dense monolithic three-dimensional (M3D) SRAM-CIM. However, CN-MOSFETs possess unique process variations with asymmetric spatial correlations which can significantly influence the performance and reliability of carbon-based SRAM-CIM. In this paper, new M3D-4N4P SRAM-CIM cells with CN-MOSFETs are proposed with optimized profiles for achieving ultra-high integration density while preserving robustness of data-access and computation. Furthermore, the variation-aware bit-wise data-mapping method is proposed for enhancing the performance of carbon-based SRAM-CIM by leveraging the spatial correlations of CN-MOSFETs. By minimizing the area skew of vertically-stacked layers, the areas of proposed M3D-4N4P SRAM-CIM cells are reduced by up to 50.32% compared to the previous 6N2P SRAM-CIM cells assuming carbon nanotube transistor technology. The proposed M3D-4N4P SRAM-CIM array also achieves by up to$2.17\times $higher throughput on arithmetic operations and 18.34% lower computing latency with 25.36% reduced energy consumptions on MAC-based benchmarks, respectively, compared to the previous 2D-6N2P SRAM-CIM array.
Dengfeng Wang, Weifeng He, Qin Wang 0009, Hailong Jiao, Yanan Sun 0003
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 HEIRS: Hybrid Three-Dimension RRAM- and SRAM-CIM Architecture for Multi-task Transformer Acceleration
abstract
Large-scale transformer with millions of weights achieves great success in multiple natural language processing (NLP) tasks. To release the memory bottleneck of multi-task model deployment, transfer learning tunes part of weights with shared parameters among tasks. Moreover, computing-in-memory (CIM) emerges as an efficient solution for neural network acceleration. With higher storage density, RRAM-CIM can store the large-scale model without costly weight loading, compared with another mainstream SRAM-CIM. However, the RRAM rewrite for tuned and dynamic weight matrix-vector-multiplication (MVM) in transformers requires high-cost RRAM writing in RRAM-CIM. Current hybrid CIM can compensate for the weakness of RRAM-CIM by adding SRAM-CIM with independent MVM. However, the tuned weights in transfer learning cannot be implemented due to the demand for the cooperative addition of MVM results from both shared and tuned weights. In this paper, a hybrid three-dimension RRAM-CIM and SRAM-CIM architecture (HEIRS) is proposed for multi-task transformer acceleration, with monolithically 3D integration of high-density RRAM-CIM and high-performance SRAM-CIM. The 3D RRAM-CIM with ultra-high density stores the whole model with mitigated off-chip weight loading. The SRAM-CIM is employed for efficiently performing dynamic weight MVM without RRAM rewrite. Moreover, a novel hybrid-CIM paradigm is proposed with an input selective adder tree, to support cooperative addition in transfer learning. Experiments show that, compared with RRAM-CIM and SRAM-CIM, the proposed HEIRS improves the energy efficiency by up to 7.83x and 2.29x on BERT, respectively. Meanwhile, the latency is also reduced by up to 85.5% and the storage density is enhanced by 7.2x, compared to RRAM-CIM.
Liukai Xu, Shuai Yuan 0016, Dengfeng Wang, Xueqing Li 0002, Yanan Sun 0003
DAC6
2024 VSPIM: SRAM Processing-in-Memory DNN Acceleration via Vector-Scalar Operations
abstract
Processing-in-Memory (PIM) has been widely explored for accelerating data-intensive machine learning computation that mainly consists of general-matrix-multiplication (GEMM), by mitigating the burden of data movements and exploiting the ultra-high memory parallelism. The two mainstreams of PIM, the analog- and digital-type, have both been exploited in accelerating machine learning workloads by numerous outstanding prior works. Currently, the digital-PIM is increasingly favored due to the broader computing support and the avoidance of errors caused by intrinsic non-idealities, e.g., process variation. Nevertheless, it still lacks further optimization considering the characteristics of the GEMM computation, including better efficient data layout and scheduling, and the ability to handle the sparsity of activations at the bit-level. To boost the performance and efficiency of digital SRAM PIM, we propose the architecture called VSPIM that performs the computation in a bit-serial fashion, with unique support of vector-scalar computing pattern. The novelties of the VSPIM can be concluded as follows: 1) support bit-serial based scalar-vector computing via ingenious parallel bit-broadcasting; 2) refine the GEMM mapping strategy and computing pattern to enhance performance and efficiency; 3) powered by the introduced scalar-vector operation, the bit-sparsity of activation is leveraged to halt unnecessary computation to maximize efficiency and throughput. Our comprehensive evaluation shows that, compared to the state-of-the-art SRAM-based digital-PIM design (Neural Cache), VSPIM can significantly boost the performance and energy efficiency by up to$8.87\times$and$4.81\times$respectively, with negligible area overhead, upon multiple representative neural networks.
Chen Nie, Chenyu Tang, Jie Lin 0004, Chenyang Lv, Ting Cao 0007, Weifeng Zhang 0003, Li Jiang 0002, Xiaoyao Liang, Weikang Qian, Yanan Sun 0003, Zhezhi He
IEEE Trans. Computers11
2023 DeepTH: Chip Placement with Deep Reinforcement Learning Using a Three-Head Policy Network
abstract
Modern very-large-scale integrated (VLSI) circuit placement with huge state space is a critical task for achieving layouts with high performance. Recently, reinforcement learning (RL) algorithms have made a promising breakthrough to dramatically save design time than human effort. However, the previous RL-based works either require a large dataset of chip placements for pre-training or produce illegal final placement solutions. In this paper, DeepTH, a three-head policy gradient placer, is proposed to learn from scratch without the need of pre-training, and generate superior chip floorplans. Graph neural network is initially adopted to extract the features from nodes and nets of chips for estimating the policy and value. To efficiently improve the quality of floorplans, a reconstruction head is employed in the RL network to recover the visual representation of the current placement, by enriching the extracted features of placement embedding. Besides, the reconstruction error is used as a bonus during training to encourage exploration while alleviating the sparse reward problem. Furthermore, the expert knowledge of floorplanning preference is embedded into the decision process to narrow down the potential action space. Experiment results on the ISPD 2005 benchmark have shown that our method achieves 19.02% HPWL improvement than the analytic placer DREAMPlace and 19.89% improvement at least than the state-of-the-art RL algorithms.
Dengwei Zhao, Shuai Yuan 0016, Yanan Sun 0003, Shikui Tu, Lei Xu 0001
DATE3
2023 TL-nvSRAM-CIM: Ultra-High-Density Three-Level ReRAM-Assisted Computing-in-nvSRAM with DC-Power Free Restore and Ternary MAC Operations
abstract
Accommodating all the weights on-chip for large-scale NNs remains a great challenge for SRAM based computing-in-memory (SRAM-CIM) with limited on-chip capacity. Previous non-volatile SRAM-CIM (nvSRAM-CIM) addresses this issue by integrating high-density single-level ReRAMs on the top of high-efficiency SRAM-CIM for weight storage to eliminate the off-chip memory access. However, previous SL-nvSRAM-CIM suffers from poor scalability for an increased number of SL-ReRAMs and limited computing efficiency. To overcome these challenges, this work proposes an ultra-high-density three-level ReRAMs-assisted computing-in-nonvolatile-SRAM (TL-nvSRAM-CIM) scheme for large NN models. The clustered n-selector-n-ReRAM (cluster-nSnRs) is employed for reliable weight-restore with eliminated DC power. Furthermore, a ternary SRAM-CIM mechanism with differential computing scheme is proposed for energy-efficient ternary MAC operations while preserving high NN accuracy. The proposed TL-nvSRAM-CIM achieves 7.8x higher storage density, compared with the state-of-art works. Moreover, TL-nvSRAM-CIM shows up to 2.9x and 2.0x enhanced energy efficiency, respectively, compared to the baseline designs of SRAM-CIM and ReRAM-CIM, respectively.
Dengfeng Wang, Liukai Xu, Songyuan Liu, Zhi Li 0058, Weifeng He, Xueqing Li 0002, Yanan Sun 0003
ICCAD8
2023 CDAR-DRAM: Enabling Runtime DRAM Performance and Energy Optimization via In-Situ Charge Detection and Adaptive Data Restoration
abstract
With the increasing of dynamic random access memory’s (DRAM) capacity, the refresh operation rapidly becomes a major concern to the performance of the current computational system. Moreover, conservative timing parameters adopted for access operations make an increasing amount of negative impact on system performance and energy efficiency. In this article, we propose an in-situ charge detection and adaptive data restoration DRAM (CDAR-DRAM) architecture, which can dynamically adjust the refresh rate and relax the constraints on access timing by removing pessimistic timing margins for PVT variations. CDAR-DRAM employs a low-cost skewed-inverter-based detector to monitor the bitline voltage in runtime and estimate real-time timing parameters of cells. Based on the detector, an adaptive refresh and restore scheme (CDAR-ref) is presented, which progressively reduces the refresh rate and partially restores cells’ voltage just enough for cells with sufficient charge, thereby optimizing both refresh and restoration operations. Moreover, a supplementary adaptive access scheme (CDAR-acc) is presented, which detects the runtime charge level of recently accessed rows and reduces access latency aggressively, benefitting workloads in a single-core system and memory nonintensive workloads in a multicore system. CDAR’s flexibility allows the two schemes to be combined. The evaluation shows that in an eight-core system, the combined scheme improves performance and energy efficiency by 15.2% and 22.6%, respectively.
Yuxuan Qin, Chuxiong Lin, Weifeng He, Yanan Sun 0003, Zhigang Mao, Mingoo Seok
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 A DFT-Compatible In-Situ Timing Error Detection and Correction Structure Featuring Low Area and Test Overhead
abstract
In-situ timing error detection and correction (EDAC) structure is widely adopted in timing-error resilient circuits to reduce the conservative timing guardband induced by process, voltage, and temperature (PVT) variations. However, it introduces the latch-based datapath as well as extra detection and propagation logic, therefore challenges the design-for-testability (DFT) implementation. In this article, we propose a novel DFT-compatible EDAC structure with significant signal control simplification and test-pattern complexity reduction, featuring low area and test overhead. This structure leverages a new scannable EDAC cell (SEDC) which can be configured for timing EDAC in normal mode, or for shift operations as a flip-flop in scan mode. Specifically, the proposed detection logic can be controlled succinctly in scan shift operations and then observed via the global error propagation logic with simple control signal configurations during the test. Therefore, the sophisticated test pattern generation and critical path sensitization are removed. Based on the structure, a shift-based test method is presented to cover the EDAC structure with a low test pattern complexity and test time overheads. As compared with previous works, the proposed SEDC saves 30.5% area, 16.6% power, and 20.3% delay. Besides, our test method cooperating with the proposed EDAC structure reduces$149\times $and$23\times $static and at-speed test patterns, respectively, together with$232\times $static test cycles, and up to$25\times $at-speed test cycles on average, which proves the effectiveness for DFT.
Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 CREAM: Computing in ReRAM-Assisted Energy- and Area-Efficient SRAM for Reliable Neural Network Acceleration
abstract
SRAM-based computing-in-memory (CIM) has been widely explored to accelerate neural networks (NNs). However, it is challenging to store all weights of many modern NNs due to limited on-chip SRAM capacity. This bottleneck induces a large amount of off-chip DRAM accesses and impedes the improvement of performance and energy efficiency. This paper proposes a new approach of computing in resistive random-access memory (ReRAM)-assisted energy- and area-efficient SRAM (CREAM) for accelerating large-scale NNs while eliminating the DRAM access. The NN weights are all stored in high-density on-chip ReRAMs and restored to the proposed non-volatile SRAM (nvSRAM) CIM cells with array-level parallelism. Furthermore, to deal with the influence of ReRAM and CMOS variations, a novel layer-wise and bit-wise weight-configuration search algorithm is proposed by leveraging different sensitivity of each layer in NN models. A data-aware weight-mapping method is also presented to efficiently map NN models to ReRAMs in CREAM for high computation parallelism. The experiment results show$10.3\times $weight storage density over the standard 6T SRAM array. Evaluations of ResNet-18 and VGG-9 on CIFAR-10/CIFAR-100 datasets show up to$3.47\times $and$1.70\times $energy efficiency over two baseline designs of SRAM-CIM and ReRAM-CIM, respectively, in addition to 15.6% higher accuracy than ReRAM-CIM under device variations.
Yanan Sun 0003, Dengfeng Wang, Liukai Xu, Zhi Li 0058, Songyuan Liu, Weifeng He, Yongpan Liu, Huazhong Yang, Xueqing Li 0002
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 BC-MVLiM: A Binary-Compatible Multi-Valued Logic-in-Memory Based on Memristive Crossbars
abstract
Logic-in-memory with memristive crossbars is an attractive approach for realizing beyond von Neumann architectures. Multi-valued logic (MVL) containing more than two logic levels can enhance the computing speed with reduced number of logic operations. In this paper, a binary-compatible multi-valued logic-in-memory (BC-MVLiM) scheme is proposed with memristive dual-crossbars where both inputs and outputs are represented by the multi-level cells of memristors. Both of the binary and multiple-valued logic operations can be implemented in the proposed BC-MVLiM scheme depending on the radix of inputs. The proposed BC-MVLiM circuitry supports multiple row-wise and column-wise logic gates with multiple fan-ins and fan-outs for binary and ternary systems by leveraging both inter- and intra-crossbar operations. Experimental results show that the proposed BC-MVLiM-based multi-digit adder enhances the computation speed by up to 76.10% and 83.82%, for binary and ternary systems, respectively, as compared with the previously published memristive logic designs. By preventing the errors from propagating across multiple stages, the error rate of proposed BC-MVLiM is also reduced by up to 98.20% compared to the previous memristive logic designs in the presence of device variations.
Yanan Sun 0003, Zhi Li 0058, Weifeng He, Qin Wang 0009, Zhigang Mao
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 Write or not: programming scheme optimization for RRAM-based neuromorphic computing
abstract
One main fault-tolerant method for a neural network accelerator based on resistive random access memory crossbars is the programming-based method, which is also known as write-and-verify (W-V). In the basic W-V scheme, all devices in crossbars are programmed repeatedly until they are close enough to their targets, which costs huge overhead. To reduce the cost, we optimize the W-V scheme by proposing a probabilistic termination criterion on a single device and a systematic optimization method on multiple devices. Furthermore, we propose a joint algorithm that assists the novel W-V scheme by incremental retraining, which further reduces the W-V cost. Compared to the basic W-V scheme, our proposed method improves the accuracy by 0.23% for ResNet18 on CIFAR10 with only 9.7% W-V cost under variation with σ = 1.2.
Ziqi Meng, Yanan Sun 0003, Weikang Qian
DAC2
2022 CREAM: computing in ReRAM-assisted energy and area-efficient SRAM for neural network acceleration
abstract
Computing-in-memory has been widely explored to accelerate DNN. However, most existing CIM cannot store all NN weights due to limited SRAM capacity for edge AI devices, inducing a large amount off-chip DRAM access. In this paper, a new computing in ReRAM-assisted energy and area-efficient SRAM (CREAM) is proposed for implementing large-scale NNs while eliminating off-chip DRAM access. The weights of DNN are all stored in the high-dense on-chip ReRAM devices and restored to the proposed nvSRAM-CIM cells with array-level parallelism. A data-aware weight-mapping method is also proposed to enhance the CIM performance while fully exploiting the hardware utilization. Experiment results show that the proposed CREAM scheme enhances the storage density by up to 7.94x compared to the traditional SRAM arrays. The energy-efficiency of proposed CREAM is also enhanced by 2.14x and 1.99x, compared to the traditional SRAM-CIM with off-chip DRAM access and ReRAM-CIM circuits, respectively.
Liukai Xu, Songyuan Liu, Zhi Li 0058, Dengfeng Wang, Yanan Sun 0003, Xueqing Li 0002, Weifeng He
DAC6
2022 An 8T/Cell FeFET-Based Nonvolatile SRAM with Improved Density and Sub-fJ Backup and Restore Energy
abstract
In normally-off instant-on applications, power-gating of the embedded memory is an effective way for higher power efficiency by preventing long-standby-time leakage energy. Recent efforts of nonvolatile SRAM (nvSRAM) design with in-cell NVM element backup provide an efficient way for both normal-mode computing and off-mode backup and restore (B&R) operations. For these efforts, circuit innovations are required to achieve optimal balance between B&R energy and area overheads. In this paper, we report a novel 8T/cell FeFET-based nvSRAM design that outperforms prior FeFET-based designs with higher density, while still maintaining the advantage of only sub-fJ energy for each B&R operation, 363x lower than the existing RRAM-based nvSRAM design. Compared with prior FeFET-based designs, this design reduces the B&R transistor count per cell from 4 to only 2, which leads to a significant total area overhead reduction of 11%.
Nuo Xiu, Juejian Wu, Yanan Sun 0003, Huazhong Yang, Narayanan Vijaykrishnan, Sumitha George, Xueqing Li 0002
ISCAS5
2022 MSLM-RF: A Spatial Feature Enhanced Random Forest for On-Board Hyperspectral Image Classification
abstract
Hyperspectral imaging (HSI) greatly improves the capacity to identify and monitor ground objects due to the high spectral resolution. As the real-time remote sensing monitoring and warning tasks are getting more attention, new algorithms for low-power on-board classification are required to reduce the transmission time of satellite downlink. In this paper, we propose the Multi-Scale Local Maximum Random Forest (MSLM-RF) to significantly reduce the energy consumption while retaining high classification accuracy. The proposed MSLM-RF uses multi-scale maximum filters for spatial feature extraction and Random Forest for classification after spectral and spatial features fusion. The spatial features are efficiently extracted with low computational complexity by regarding the maximum light intensity values in different ranges of pixels as anchor points. MSLM-RF only consists of integer comparisons and a few additions, thereby eliminating the energy-hungry operations such as multiplication and exponentiation. According to experimental results on the HSI benchmark datasets, MSLM-RF delivers a better trade-off in accuracy and computational complexity than the state-of-the-art classification algorithms. Besides, MSLM-RF gets higher average classification accuracy and lower energy consumption than the previous on-board algorithms. The obtained results show the suitability of the proposed algorithm to accomplish practical real-time classification tasks on-board with low energy consumption.
Shuai Yuan 0016, Yanan Sun 0003, Weifeng He, Qianrong Gu, Zhigang Mao, Shikui Tu
IEEE Trans. Geosci. Remote. Sens.2
2021 CDAR-DRAM: An In-situ Charge Detection and Adaptive Data Restoration DRAM Architecture for Performance and Energy Efficiency Improvement
abstract
As the capacity of DRAM continues to grow, the refresh operation rapidly becomes the performance and power-efficiency bottleneck. Also, restore time, the time given for recharging cells post access, makes an increasingly large amount of negative impact on performance. To tackle these problems, in this paper, we propose an in-situ charge detection and adaptive data restoration DRAM (CDAR-DRAM) architecture, which can dynamically adjust the refresh rate and also relax the constraints on restore time. The proposed CDAR-DRAM employs a low-cost skewed-inverter-based detector, which can reduce the excessive timing margins that prior work added to guarantee the functionality of leaky DRAM cells under the worst-case temperature condition. Moreover, an adaptive DRAM refresh and restore scheme is proposed, which can switch automatically between two modes: (i) a refresh mode that supports adaptive refresh rate, and (ii) a restore mode that relaxes the constraints on restore time dynamically for cells having sufficient charge. With the transistor-and architecture-level simulations, we evaluate the CDAR-DRAM in an 8-core system across different workloads. Compared with the prior art, the proposed architecture achieves a 9.4% improvement in system performance and a 14.3% reduction in energy consumption, without requiring the time-consuming profiling process which many prior works employed.
Chuxiong Lin, Weifeng He, Yanan Sun 0003, Zhigang Mao, Mingoo Seok
DAC3
2021 Digital Offset for RRAM-based Neuromorphic Computing: A Novel Solution to Conquer Cycle-to-cycle Variation
abstract
Resistance variation in memristor device hinders the practical use of resistive random access memory (RRAM) crossbars as neural network (NN) accelerators. Previous fault-tolerant methods cannot effectively handle cycle-to-cycle variation (CCV). Many of them also use a pair of positive-weight and negative-weight crossbars to store a weight matrix, which implicitly enhances the fault tolerance but doubles the hardware cost. This paper proposes a novel solution that dramatically reduces the NN accuracy loss under CCV, while still using a single crossbar to store a weight matrix. The key idea is to introduce digital offsets into the crossbar, which further enables two techniques to conquer CCV. The first is a variation-aware weight optimization method that determines the optimal target weights to be written into the crossbar; the second is a post-writing tuning method that optimally sets the digital offsets to recover the accuracy loss due to variation. Simulation results show that the accuracy maintains the ideal value for LeNet with MNIST and only drops by 2.77% over the ideal value for ResNet-18 with CIFAR-10 under a large resistance variation. Moreover, compared to state-of-the-art fault-tolerant methods, our method achieves a better NN accuracy with at least 50% fewer crossbars.
Ziqi Meng, Weikang Qian, Yilong Zhao 0004, Yanan Sun 0003, Li Jiang 0002
DATE4
2021 An Area-Efficient Scannable In Situ Timing Error Detection Technique Featuring Low Test Overhead for Resilient Circuits
abstract
Timing error detection is a key technique for resilient circuits to explore the timing margins, yet it hinders the scan shift operations and increases the excessive test overhead. In this paper, we propose an area-efficient scannable in situ timing error detection technique consisting of a lightweight scannable error-detection cell and propagation logics, featuring low design-for-test effort and test overhead. The proposed error-detection cell fully reuses its main and shadow latches to construct the latch-based error-detection structure in normal mode, or the flip-flop-based datapath in scan mode. Therefore, it not only offers the time-borrowing ability to lower the correction overheads, but also supports the scan shift operations and detection logic tests. Besides, the dependency of error signal generation on the critical path sensitization is eliminated by configuring input and clock signals of error propagation logics, and thereby the detection and propagation logic can be tested easily. Benefiting from the technique, a set of test methods is presented with lower test pattern scales and test cycle overheads. As compared with previous works, the proposed cell saves at least 30.5% area overhead. Besides, experimental results across several benchmark circuits show that 116x of test patterns, 232x of static test cycles, and 26x of at-speed test cycles are saved on average, proving the effectiveness of the proposed technique for the design-for-test requirement.
Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
ICCAD3
2021 Investigation of Dynamic Leakage-Suppression Logic Techniques Crossing Different Technology Nodes from 180 nm Bulk CMOS to 7 nm FinFET Plus Process
abstract
Leakage power reduction techniques are crucial for energy-efficient circuits. This paper investigates the leakage suppression capability, performance, and reliability of dynamic leakage suppression logic (DLSL) and feedforward leakage self-suppression logic (FLSL) techniques, crossing different technology nodes from TSMC 180 nm bulk CMOS to 7 nm FinFET Plus process. Compared with CMOS benchmarks, experimental results show that DLSL-based benchmarks demonstrate a leakage power reduction for four orders of magnitude in 180 nm and 130 nm technologies, while only two orders of magnitude in other technologies. Moreover, FLSL offers a 4-28× performance improvement over DLSL at a cost of 2× leakage power.
Zihan Lian, Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
ISCAS5
2021 Design of Ternary Logic-in-Memory Based on Memristive Dual-Crossbars
abstract
Implementing logic within memristive crossbar is an attractive approach to overcome the memory wall in conventional von Neumann architectures. Ternary logic with three logic levels can reduce the number of logic operations and enhance the computing speed compared to the binary logic. In this paper, a ternary logic-in-memory scheme is proposed based on the memristive dual-crossbar structure where the inputs and outputs are represented by the multi-level cells of memristors. Two inter-crossbar ternary logic gates and one intra-crossbar binary logic gate for both row and column-wise operations are supported in the proposed scheme to effectively reduce the operation latency. Experimental results show that the operation steps of the proposed multi-trit ternary adder are reduced by up to 83.82%, as compared with previously published binary memristive logic designs. The computation energy consumed by the proposed ternary adder is also reduced by up to 35.87% as compared to previously published binary IMPLY logic design.
Yanan Sun 0003, Weifeng He, Qin Wang 0009
ISCAS2
2021 An Energy-Efficient Logic Cell Library Design Methodology with Fine Granularity of Driving Strength for Near- and Sub-Threshold Digital Circuits
abstract
Commercial multi-threshold standard logic cell libraries are designed for nominal super-threshold circuits. If blindly used at near- and sub-threshold voltages, such libraries exhibit excessively coarse granularity in driving strength, leading to sub-optimal logic synthesis and placement-and- routing results. To tackle this problem, a holistic methodology for designing a near- and sub-threshold standard cell library that has fine driving strength granularity is presented in this paper. Meanwhile, the proposed methodology leverages inverse narrow width effect, reverse short channel effect and forward body biasing to modulate the driving strength at low area overheads. Based on the proposed methodology, we develop a 65nm multi-threshold-voltage, multi-channel-length library and benchmark it against the commercial library across several common circuits. The results show a 26.6% reduction in power-delay product, a 28.1% reduction in energy-delay product, and a 27.0% reduction in leakage power at 5.8% area overhead on average, confirming the efficiency of the methodology in near- and sub-threshold digital circuits design.
Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
ISCAS3
2021 An Ultra-Low Leakage Bitcell Structure with the Feedforward Self-Suppression Scheme for Near-Threshold SRAM
abstract
Leakage power consumption has become a critical issue for low power Static Random-Access Memory (SRAM) design in the near-threshold regime. In this paper, an ultra-low leakage fourteen-transistor SRAM bitcell structure with the feedforward self-suppression scheme is presented. To reduce the leakage power significantly as well as maintain the data stability in hold state, a cross-coupled dynamic leakage- suppression inverter-based structure is adopted. Furthermore, the bypass scheme is employed to enable the speed modulation for bitcell read and write operations. As compared with state- of-the-art designs, a 65 nm 8kb SRAM array with the proposed bitcell structure achieves 75× leakage power, 45% write power as well as 65% read power reduction at 0.4V. Further comparisons with different processes verify up to 38k times leakage power reduction in 130nm planar process and 139× in 7nm plus FinFET process.
Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok
ISCAS4
2021 Unary Coding and Variation-Aware Optimal Mapping Scheme for Reliable ReRAM-Based Neuromorphic Computing
abstract
Neural network (NN) computing contains a large number of multiply-and-accumulate (MAC) operations. The performance of NN accelerator is limited with the traditional von Neumann architecture due to the tremendous off-chip memory accesses. Resistive random-access memory (ReRAM)-based crossbars can naturally perform matrix–vector multiplication (MVM) operations and are well suitable for NN accelerators. In the existing ReRAM-based NN accelerators, the synaptic weights represented by the conductances of ReRAMs are mainly based on the binary coding. However, the imperfect fabrication process combined with stochastic filament-based switching leads to resistance variations of ReRAMs, which can significantly alter the weights in binary synapses and degrade the NN accuracy. Moreover, the NN accuracy further deteriorates with multilevel cells (MLCs) used for reducing hardware overhead. In this article, a novel unary coding of synaptic weights is proposed to overcome the resistance variations of MLCs and achieve reliable ReRAM-based neuromorphic computing. A variation-aware optimal mapping scheme is also proposed in compliance with the unary coding to guarantee high accuracy by leveraging a unique feature of unary coding—the existence of multiple ways to represent the same value. The optimal mapping obtains very small errors for weights with resistance variations of MLCs. Our simulation results show that under resistance variations, the proposed method achieves less than 0.08% and 3.43% accuracy loss on CIFAR10 and ImageNet, respectively, compared to the ideal accuracy. With each synaptic weight represented by four 2-b MLCs, the proposed method improves the accuracy over the traditional binary coding scheme by 83.39% and 87.6% for CIFAR10 and ImageNet, respectively.
Yanan Sun 0003, Zhi Li 0058, Yilong Zhao 0004, Jiachen Jiang, Weikang Qian, Zhezhi He, Li Jiang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 ITT-RNA: Imperfection Tolerable Training for RRAM-Crossbar-Based Deep Neural-Network Accelerator
abstract
Deep neural networks (DNNs) have gained a strong momentum among various applications. The enormous matrix-multiplication exhibited in the above DNNs is computation and memory intensive. Resistive random-access memory crossbar (RRAM-crossbar) consisting of memristor cells can naturally carry out the matrix-vector multiplication. RRAM-crossbar-based accelerator, therefore, has two orders of magnitude of higher energy-efficiency than conventional accelerators. The imperfect fabrication process of RRAM-crossbars, however, causes various defects and process variations. These fabrication imperfections not only result in significant yield loss but also degrade the accuracy of DNNs executed on the RRAM-crossbars. In this article, we first propose an accelerator-friendly neural-network training method, by leveraging the inherent self-healing capability of the neural network, to prevent the large-weight synapses from being mapped to the imperfect memristors. Next, we propose a dynamic adjustment mechanism to extend the above method for DNNs, such as multilayer perceptrons (MLPs), wherein the imperfect-memristor induced errors can accumulate and magnify through multiple layers. Such off-device training method is a pure software solution, and it is unable to provide enough accuracy for convolutional neural networks (CNNs). Several works propose error-tolerable hardware design by allowing the retraining of CNNs on the RRAM-crossbar. Although this hardware-based on-device training method is effective, the frequent write operation on RRAM-crossbar hurt the endurance of RRAM-crossbars. Consequently, we propose a software and hardware co-design methodology to effectively preserve the classification accuracy of CNN with few on-device training iterations. The experimental results show that the proposed method can guarantee ≤1.1% loss of accuracy for resistance variations in MLP and CNN. Moreover, the proposed method can guarantee ≤1% loss of accuracy even when stuck-at-faults (SAFs) rate = 20%.
Zhuoran Song, Yanan Sun 0003, Lerong Chen, Tianjian Li, Naifeng Jing, Xiaoyao Liang, Li Jiang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Go Unary: A Novel Synapse Coding and Mapping Scheme for Reliable ReRAM-based Neuromorphic Computing
abstract
Neural network (NN) computing contains a large number of multiply-and-accumulate (MAC) operations, which is the speed bottleneck in traditional von Neumann architecture. Resistive random access memory (ReRAM)-based crossbar is well suited for matrix-vector multiplication. Existing ReRAM-based NNs are mainly based on the binary coding for synaptic weights. However, the imperfect fabrication process combined with stochastic filament-based switching leads to resistance variations, which can significantly affect the weights in binary synapses and degrade the accuracy of NNs. Further, as multi-level cells (MLCs) are being developed for reducing hardware overhead, the NN accuracy deteriorates more due to the resistance variations in the binary coding. In this paper, a novel unary coding of synaptic weights is presented to overcome the resistance variations of MLCs and achieve reliable ReRAM-based neuromorphic computing. The priority mapping is also proposed in compliance with the unary coding to guarantee high accuracy by mapping those bits with lower resistance states to ReRAMs with smaller resistance variations. Our experimental results show that the proposed method provides less than 0.45% and 5.48% accuracy loss on LeNet (on MNIST dataset) and VGG16 (on CIFAR-10 dataset), respectively, with acceptable hardware cost.
Yanan Sun 0003, Weikang Qian, Ziqi Meng, Li Jiang 0002
DATE2
2020 ESNreram: An Energy-Efficient Sparse Neural Network Based on Resistive Random-Access Memory
abstract
The sparsity in the deep neural networks (DNNs) can be leveraged by methods such as pruning and quantization to assist the energy-efficient deployment of large-scale deep neural networks onto hardware platforms, such as GPU and ASIC, for better performance and power efficiency. However, for the metal-oxide resistive random access memory (ReRAM) architecture, the study of energy-efficient methods still shrink the model size or constrain the precision of DNN by leveraging the DNN sparsity. Due to the circuit features of ReRAM, reading bit-0 naturally consumes less energy than reading bit-1. In this paper, we exploit the fine-grained tuning method on the bit-level to reduce energy consumption of ReRAM. Specifically, we present the gradient-search and the weight-group update algorithm, which can significantly unbalance the proportion of bit-1 and bit-0 inside the weights of DNN with negligible NN accuracy loss. Experiments demonstrate that the percentage of bit-0, in some typical convolutional neural networks (CNNs), increases to 33.8%, with less than 0.5% degradation in NN accuracy. The energy reduction can be up to 65%.
Zhuoran Song, Yilong Zhao 0004, Yanan Sun 0003, Xiaoyao Liang, Li Jiang 0002
ACM Great Lakes Symposium on VLSI3
2018 Electronic-Photonic Integrated Circuit Design and Crosstalk Modeling for a High Density Multi-Lane MZM Array
abstract
To meet the rapidly increasing bandwidth density demand in high-performance computers and data centers, the number of parallel lanes in a short-reach optical interconnect link is expected to increase by several folds. This paper presents a multi-lane traveling-wave Mach-Zehnder modulator (MZM) array design using silicon photonics technology. We treat the MZM as an electronic-photonic integrated circuit and adopt a newly developed multi-physics cross-layer design methodology. The design spans electrical/optical device modeling, electromagnetic simulation, and circuit analysis. The crosstalk issue within a closely-packed MZM array is investigated. The simulated 3-dB EO bandwidth of a single MZM is 17.5 GHz. The complete link considering crosstalk in the MZM array achieves BER-12at 25 Gbps. A prototype chip with 4 MZMs is designed using IMEC 50G silicon photonics technology, saving significant chip area compared to conventional design.
Pengfei Ji, Yanan Sun 0003, Weifeng He
ISCAS4
2018 Variation-Aware Global Placement for Improving Timing-Yield of Carbon-Nanotube Field Effect Transistor Circuit
abstract
As the conventional silicon-based CMOS technology marches toward the sub-10nm region, the problem of high power density becomes increasingly serious. Under this circumstance, the carbon-nanotube field effect transistors (CNFETs) emerge as a promising alternative to the conventional silicon-based CMOS devices. However, they experience a much larger variation than the silicon-based CMOS devices, which results in a large circuit delay variation and hence, a significant timing yield loss. One of the main variation sources is the carbon-nanotube (CNT) density variation. However, it shows a special property not existing for silicon-based CMOS devices, namely the asymmetric spatial correlation. In this work, we propose novel global placement algorithms to reduce the timing yield loss caused by the CNT density variation. To effectively reduce the statistical circuit delay, we first develop a statistical delay measure for a segment of gates. Based on this measure, we further develop a segment-based strategy and a path-based placement strategy to reduce the delays of the statistically critical paths. Experimental results demonstrated that both of our approaches effectively improve the timing yield.
Chen Wang 0072, Yanan Sun 0003, Shiyan Hu 0001, Li Jiang 0002, Weikang Qian
ACM Trans. Design Autom. Electr. Syst.2
2017 A 0.33 V 2.5 μW cross-point data-aware write structure, read-half-select disturb-free sub-threshold SRAM in 130 nm CMOS
Wei Jin 0004, Weifeng He, Jianfei Jiang 0001, Haichao Huang, Xuejun Zhao, Yanan Sun 0003, Naifeng Jing
Integr.6
2016 Enabling in-situ logic-in-memory capability using resistive-RAM crossbar memory
abstract
Recently, logic-in-memory (LIM) is gaining growing interest because it eliminates the unnecessary data movement between the memory and logic components that embarrasses both performance and power dissipation in modern microprocessors. However, most of the existing LIM just puts the logic and memory closer rather than a true integration due to the incompatibility of logic and memory circuit structures. In this paper, we propose a in-situ LIM design by leveraging the emerging ReRAM memory in a crossbar structure. It performs in-situ logic processing using the same memory cells based on the resistive states of ReRAMs with non-destructive operations, and therefore can exploit the large internal bandwidth available in the array without data readout. The logic exploration exposes that the proposed LIM can support different logical functions and get in-situ results without moving data in and out of the memory. We believe that the proposed design provides a promising solution for a true logic processing capability within memory.
Naifeng Jing, Taozhong Li, Zhongyuan Zhao 0004, Wei Jin 0004, Yanan Sun 0003, Weifeng He, Zhigang Mao
FPT5
2015 Carbon-based sleep switch dynamic logic circuits with variable strength keeper for lower-leakage currents and higher-speed
abstract
A new variable strength keeper technique is presented in this paper for achieving higher-speed and lower-leakage currents in wide fan-in dynamic logic gates with carbon nanotube transistors. The strength of the keeper is dynamically adjusted depending on the logical state of the dynamic node during evaluation phase in a domino logic circuit. While providing similar noise immunity, the evaluation delay and power-delay product of the proposed domino logic circuits are reduced by up to 13.33% and 13.84%, respectively, as compared to the standard domino logic circuits in a 16nm carbon nanotube transistor technology. Furthermore, the proposed technique provides up to 77.98% savings in average leakage power consumption as compared to the standard domino logic circuits in idle mode.
Yanan Sun 0003, Volkan Kursun
ISCAS1
2015 A Novel Robust and Low-Leakage SRAM Cell With Nine Carbon Nanotube Transistors
abstract
A novel static random-access memory (SRAM) cell with nine carbon nanotube MOSFETs (9-CN-MOSFETs) is proposed in this paper. With the new 9-CN-MOSFET SRAM cell, the read data stability is enhanced by 99.09%, while providing similar read speed as compared with the conventional six-transistor (6T) SRAM cell in a 16-nm carbon nanotube transistor technology. The worst-case write voltage margin is increased by 4.57× and 3.90× with the proposed 9-CN-MOSFET SRAM cell as compared with the conventional 6T SRAM cell and a previously published eight-transistor (8T) SRAM cell, respectively. A 1 Kibit SRAM array with the new memory cells consumes 34.18% and 12.27% lower leakage power as compared with the memory arrays with 6T and 8T SRAM cells, respectively, in idle mode. The overall electrical quality is enhanced by up to 13.63× with the proposed 9-CN-MOSFET memory circuit as compared with the other memory cells that are evaluated in this paper.
Yanan Sun 0003, Hailong Jiao, Volkan Kursun
IEEE Trans. Very Large Scale Integr. Syst.1
2013 Low-power and compact NP dynamic CMOS adder with 16nm carbon nanotube transistors
abstract
Low-power, compact, and high-performance NP dynamic CMOS circuits implemented with a 16nm carbon nanotube transistor technology are presented in this paper. The performances of two-stage pipeline 32-bit carry lookahead adders are evaluated with two circuit techniques: the carbon nanotube MOSFET (CN-MOSFET) domino logic and the CN-MOSFET NP dynamic CMOS. While providing similar propagation delay, the total area of CN-MOSFET NP dynamic CMOS circuit is reduced by 13.61% as compared to the CN-MOSFET domino adder. Miniaturization of the CN-MOSFET NP dynamic CMOS adder reduces the dynamic switching power consumption by 30.42% as compared to the CN-MOSFET domino circuit. Furthermore, the CN-MOSFET NP dynamic CMOS circuit provides 49.32% savings in leakage power consumption as compared to the CN-MOSFET domino adder.
Yanan Sun 0003, Volkan Kursun
ISCAS1
2011 Leakage current and bottom gate voltage considerations in developing maximum performance 16nm N-channel carbon nanotube transistors
abstract
The influence of substrate voltage on carbon nanotube MOSFET (CN-MOSFET) performance is investigated in this paper. The optimum device profiles with different transistor sizes are identified for achieving the highest on-state to off-state current ratio (Ion/Ioff)Tradeoffs between subthreshold leakage current and device performance are evaluated with different substrate bias voltages. Technology development guidelines are provided for achieving low-leakage, high-speed, area efficient, and manufacturable integrated circuits with carbon nanotube transistors.
Yanan Sun 0003, Volkan Kursun
ISCAS1
2011 Uniform carbon nanotube diameter and nanoarray pitch for VLSI of 16nm P-channel MOSFETs
abstract
Uniformities of carbon nanotube diameters and nanoarray pitch values of all the transistors across a chip are required to enable low cost very large scale integration (VLSI) with the carbon nanotube technology. Nanotube diameter and nanoarray pitch are concurrently optimized and unified in this paper with two different substrate bias voltages considering a wide range of p-channel transistor sizes. A performance and density metric is evaluated to identify the optimum p-type device profiles suitable for very large scale integration with a 16nm carbon nanotube transistor technology.
Yanan Sun 0003, Volkan Kursun
VLSI-SoC1