VLDB 2026 Research / reviewers in the wild / expert
Ke Chen 0018
dblp:47/6529-18
· DBLP profile ↗
33ranked-venue papers
8as first author
27since 2021 · last 2026
0000-0003-0981-3166ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 8 first-author · 26 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BIHDC: A Retrainable Fully-Binary Hyperdimensional Computing Accelerator for Edge FPGAsabstractHyperdimensional Computing (HDC) is a lightweight machine learning paradigm characterized by low complexity, efficient learning, and strong interpretability, making it well suited for edge intelligence. However, its accuracy on 2D image tasks still lags behind deep neural networks (DNNs), and existing HDC hardware often relies on in-memory computing or high-end FPGAs, which fail to meet the strict low-power and small-area requirements of edge devices and generally lack complete retraining capabilities. To address these limitations, we propose BIHDC, a lightweight HDC framework that introduces a learnable preprocessing scheme to enhance feature extraction and implements an edge FPGA-based fully binary accelerator supporting end-to-end processing, including preprocessing, encoding, training, retraining, and testing. Experimental results show that BIHDC improves classification accuracy by up to 5% compared to baseline HDC with binarized encoding and by 1.5% over baseline HDC with original encoding across four image datasets. Implemented on the Xilinx Zynq7000 platform, BIHDC reduces resource utilization and power consumption by 90% and 70%, respectively, providing an efficient and scalable solution for resource-constrained edge applications. Changzhen Han, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Weiqiang Liu 0001 |
ASP-DAC | 2 |
| 2026 | An 8T-SRAM Near-Memory Architecture for Multiplierless Approximate DCT
Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Lixia Han, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2026 | A Low Complexity BPF Polar Decoder Design with Sensitivity-Aware LUT Compression Framework
Bozhi Xiu, Yanjing Zhang, Chenggang Yan 0002, Ke Chen 0018, Bi Wu 0002, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2026 | XHDC: An Adaptive Hyperdimensional Computing Accelerator with Incremental Learning
Ruofan Yan, Changzhen Han, Ke Chen 0018, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2026 | Instant-CIM: An Instant Neural Radiance Field Computing-In-Memory Architecture for Low-Power and Real-Time AR/VR RenderingabstractNovel View Synthesis is a foundational technique for creating immersive Augmented and Virtual Reality (AR/VR) experiences, aiming to generate photorealistic images of a scene from arbitrary camera viewpoints using only a limited set of source images, with Neural Radiance Fields (NeRF) emerging as the state-of-the-art solution. However, real-time NeRF rendering on low-power devices remains challenging due to its memory-intensive hash encoding and compute-intensive Multilayer Perception (MLP). In this work, we propose Instant-CIM, the fully on-chip Computing-in-Memory (CIM) architecture for efficient NeRF rendering. At the algorithm level, Instant-CIM proposes a spatially-adaptive framework that dynamically selects the number of active hash encoding levels per spatial region based on a composite importance score derived from density and gradient. The approach replaces uniform level allocation with a threshold-based strategy that activates finer encoding levels only in regions with high representation complexity. At the hardware level, Instant-CIM proposes an in-situ hash engine that implements in-memory hash query and interpolation through 3D scene grid decomposition and Z-order based mapping schemes. Meanwhile, Instant-CIM proposes a sparse MLP engine that leverages differential-based input complemented by a precision-adjustable skipping mechanism to fully exploit spatial similarities. Comprehensive evaluation across synthetic datasets demonstrates that Instant-CIM achieves 3.0×~4.9× improvement in rendering speed and 8.6×~33× enhancement in energy efficiency compared to state-of-the-art NeRF architecture. Lixia Han, Hui Chen 0015, Xueming Fu, Ke Chen 0018, Peng Huang 0004, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Computers | 6 |
| 2026 | E2CAP: An Energy-Efficient FPGA Accelerator for Deep Reinforcement Learning With Experience Compression and Configurable PE ArrayabstractDeep reinforcement learning (DRL) has emerged as a powerful tool for solving complex decision-making tasks in domains such as robotics, autonomous systems, and gaming. However, accelerating DRL on hardware platforms faces three major challenges: 1) large memory footprint of experience replay data; 2) inefficient weight access due to weight transposition; and 3) computational imbalance across processing elements during deep neural network (DNN) inference and training. To address these issues, we propose E2CAP, an energy-efficient field-programmable gate array (FPGA) accelerator tailored for DRL workloads. E2CAP integrates a compression strategy that significantly reduces the data volume of the experience pool in the DRL model, enabling on-chip deployment of the experience pool. In addition, E2CAP features a configurable array processing core (CAP) based on configurable processing elements (CPEs) that support both intra-PE and inter-PE configurability. This design effectively addresses the issues of weight trans position and computational imbalance, improving computational efficiency and hardware resources utilization. We implement full on-chip DRL inference and training using the DQN algorithm on E2CAP. Experimental results demonstrate a 98.19% compression ratio for experience replay data. Compared to the state-of-the-art CPU and GPU, E2CAP achieves$33.01\times $and$17.72\times $higher energy efficiency, respectively, while outperforming existing FPGA-based DRL accelerators in multidimensional performance. Fen Ge, Xinjun Zhou, Ke Chen 0018, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2026 | Neuromorphic Hyperdimensional Computing for Efficiently Processing Event-Based DataabstractThe neuromorphic sensor’s event-based data output offers significant benefits, including minimal data redundancy and exceptional time resolution, which guarantee low power consumption and heightened sensitivity during the data acquisition process. Spiking neural network (SNN), with its inherent event-driven characteristic, is well-suited for processing event-based data, and its spike-based computing mechanism enhances the efficiency of data processing. Recent studies are exploring the integration of brain-inspired hyperdimensional computing (HDC) with SNN to leverage HDC’s advantages, including the low inference and training complexity, aiming to further reduce hardware overhead associated with SNN deployment. However, existing works have not effectively harnessed the information output by SNN during hyperdimensional encoding, leading to considerable area and energy overhead. In this article, an efficient neuromorphic HDC method is proposed, featuring a simplified hyperdimensional encoding approach that considers the temporal dynamics of SNN. In addition, a lightweight accelerator design matching the proposed method is also given. Experimental results show that the proposed accelerator achieves over 50% area reduction and reduces energy consumption by 20%–90%. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Gong Zhang 0002, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | PreDAC: An Efficient Framework of Pre-Refining Enhanced Design Space Exploration for Approximate ComputingabstractApproximate computing has emerged as a promising solution in energy-efficiency applications. Recently, attention has shifted from approximate components to Design Space Exploration (DSE) algorithms. However, traditional DSE algorithms face challenges in efficiently obtaining optimal solutions within large and complex design spaces. This paper introduces a prerefining enhanced design space exploration framework that provides customized design space and cost-performance formula for applications. Experimental results demonstrate that integrating this pre-refining step into various DSE algorithms leads to substantial performance gains, including up to $87 \times$ speedup and a 23% improvement in hardware overhead. Moreover, the innovative cost-performance-based DSE algorithm attains a $7.7 \times$ acceleration and further optimizes hardware metrics by an additional 8.8% compared to advanced frameworks employing the same pre-refinement. Ziying Cui, Ke Chen 0018, Bi Wu 0002, Yu Gong 0002, Chenggang Yan 0002, Weiqiang Liu 0001 |
DAC | 2 |
| 2025 | High-Performance Co-Processing Architecture Using SOT-MRAM-Based In-memory Computing SchemeabstractIn recent years, the rapid advancement of processor performance has highlighted the increasing inadequacy of memory bandwidth. To mitigate this challenge, the In-memory Computing (IMC) architecture has been proposed. Among various IMC implementations, Spin-Orbit Torque Random Access Memory (SOT-MRAM) stands out as an ideal generic co-processor due to its high density and low leakage current. However, how to efficiently integrate with existing instruction set architectures (ISA) while satisfying the characteristics of the SOT-MRAM IMC hardware poses a challenge. In this work, an instruction-driven in-memory co-processor architecture is proposed. By combining the proposed hardware-software collaborative memory redirection scheme, the SOT-MRAM-based in-memory computing can be merged into the existing computing architecture without disturbing ISA. Further, to satisfy the inter-row operation computation characteristics of SOT-MRAM IMC, a data pre-scheduling scheme oriented to efficient computation is proposed. Experimental results demonstrate a 6.2x speedup and 64% power reduction compared to a CPU-only architecture, and a 28% improvement in acceleration ratio over the state-of-the-art IMC architecture. Bi Wu 0002, Ke Chen 0018, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2025 | High-Radix Generalized Hyperbolic CORDIC and Its Hardware ImplementationabstractIn this paper, we propose a high-radix generalized hyperbolic coordinate rotation digital computer (HGH-CORDIC). This algorithm not only computes logarithmic and exponential functions with any fixed base but also significantly reduces the number of iterations required compared to traditional CORDIC methods. Initially, we present the general iteration formulas for HGH-CORDIC. Subsequently, we discuss its pivotal convergence properties and selection criteria, exemplifying these with commonly used cases. Through extensive software simulations, we validate the theoretical foundations of our approach. Finally, we explore efficient hardware implementation strategies. Our analysis indicates that, relative to state-of-the-art radix-2 GH-CORDIC, the proposed HGH-CORDIC can decrease the number of iterations by more than$50\%$while maintaining comparable accuracy. Synthesized under the 28nm CMOS technology, the reports show that the reference circuit can save about$40\%$area and power consumption averagely for$2^{x}$and$log_{2}x$calculations compared with the latest CORDIC method. Hui Chen 0015, Lianghua Quan, Ke Chen 0018, Weiqiang Liu 0001 |
IEEE Trans. Computers | 3 |
| 2025 | LAHDC: Logic-Aggregation-Based Query for Embedded Hyperdimensional Computing AcceleratorabstractWith low complexity and robustness, hyperdimensional computing (HDC) has become a promising paradigm for edge-side applications. HDC employs hypervectors (generally with 2–10 K dimensions) to represent input samples, and performs logical operations in hyperdimensional space to complete perceptual tasks. Compared to deep neural network (DNN), HDC is more suitable for lightweight edge-side applications (i.e., speech, activity recognition), due to its low complexity and less computational scheduling. However, existing HDC’s querying process relies on trained class hypervectors, resulting in on-chip storage and transmission overhead which limits the application of ASIC-based or FPGA-based HDC accelerators in embedded systems. In this article, a logic-aggregation-based query method called LAHDC is proposed to eliminate such overhead. In addition, an ultratiny HDC accelerator design matching LAHDC is also proposed, as well as an automated tool to search for optimal structure and generate hardware design code. Experimental results show that, compared to existing ASIC-based HDC accelerators, the proposed design reduce the area/energy by more than 95%/80%. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Gong Zhang 0002, Weiqiang Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | SS-MRAM: A Segment-Based Search Scheme With Configurable Matching for High-Accuracy Hyperdimensional Computing in CAM ApplicationsabstractWith the rise of hyperdimensional computing (HDC), content-addressable memories (CAMs) have emerged as an ideal choice for processing high-dimensional data. However, despite the advantages of high parallelism and low latency offered by CAM technology, it fails to address the significant loss of inference accuracy caused by closely matching hamming distances (HD). Emerging analog-based imprecise in-memory computing technologies frequently provide a minimum detectable HD that is insufficient for meeting the requirements of high similarity tasks. This limitation provides opportunities for using digital methods to realize fully exact matching memory computing technology based on magnetic random access memory (MRAM). In this work, a segment search scheme based on STT-MRAM devices and address-matching technology is proposed, which achieves zero loss in inference accuracy. An adaptive amplification structure is initially implemented by integrating the discharge method of latch structures, accompanied by the design of a 14T-2MTJ cell circuit. Through the optimization of the matching step, HD are calculated for each segment of configurable-dimensional vectors, ultimately facilitating the classification ranking of query hypervectors based on a comprehensive array architecture. Experimental results indicate that there is zero loss in inference accuracy when each segment of the dimension is configured to be below 16 bits. At 16 bits, the search power consumption in the worst-case matching scenario is measured at 1.73 fJ/bit, while the loss in inference accuracy does not exceed 0.4%. Bi Wu 0002, Shuo Ran, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | A Time Efficient Comprehensive Model of Approximate Multipliers for Design Space ExplorationabstractMultipliers play an essential role in various data processing applications and have garnered significant attention in approximate computing (AxC) for their energy-efficient features. However, formulating a precise error model for approximate data processing algorithms in conjunction with hardware metrics presents a challenge, leading to substantial time consumption in the design space exploration. This paper introduces an analytical model for approximate multipliers while considering input patterns. This model furnishes accurate error metrics, along with high-precision hardware metrics for various approximate multiplier configurations, impervious to variations in input data distribution. The proposed error model reduces the runtime by an average factor of 120.85 and, in some instances, by as much as 2,500 times, when contrasted with simulation-based methods. The design space exploration is performed on a 3×3 convolution circuit, revealing a comparable Pareto-optimal set and substantial reductions of up to 79.46% in the Power-Delay-Product (PDP) and 71.98% in area compared to the accurate counterpart. Additionally, the result of the Gaussian Blur application experiment demonstrates a 68.59% reduction in PDP and a 56.21% reduction in area, all while maintaining a PSNR of 30 dB. Ziying Cui, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Yu Gong 0002, Weiqiang Liu 0001 |
ARITH | 2 |
| 2024 | Most Significant One-Driven Shifting Dynamic Efficient Multipliers for Large Language ModelsabstractLarge Language Models (LLMs) have demonstrated exceptional performance but demand significantly more computational power and memory compared to Deep Neural Networks (DNNs). This necessitates the development of more energy-efficient hardware designs. This paper introduces a novel weight approximation strategy for quantized LLMs, resulting in the creation of a highly efficient approximate multiplier based on Most Significant One (MSO) shifting. When compared to energy-efficient approximate logarithmic multipliers and precision-demanding approximate non-logarithmic multipliers, the proposed design strikes an optimal balance between accuracy and hardware cost. It maintains a superior level of accuracy while incurring hardware costs comparable to logarithmic multipliers and, in some cases, even outperforming them. In particular, when compared to exact multiplier, the proposed design achieves significant reductions, including up to a 28.31% reduction in area, a 57.84% decrease in power consumption, and an 11.86% reduction in delay. The experiments demonstrate that the proposed multiplier in DNNs can save approximately 60% of energy without compromising task accuracy. Similarly, experiments on the Transformer accelerator indicate substantial energy savings for LLMs using the proposed design. Ke Chen 0018, Bi Wu 0002, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2024 | Fully Learnable Hyperdimensional Computing Framework With Ultratiny Accelerator for Edge-Side ApplicationsabstractBrain-inspired hyperdimensional computing (HDC) is a new computational paradigm that encodes input sample into a hypervector (generally with dimensions of$2K-10K$), and performs simple arithmetic and logic operations in the hyperdimensional space to complete perceptual tasks like human brain. Due to its simplicity, interpretability, and robustness, HDC has gradually become a competitor and substitute for deep neural network (DNN) in many tasks. However, there exists an accuracy gap between existing heuristic HDC algorithms and DNN in computer vision tasks, as existing encoding methods have difficulty in filtering out large amount of background and noise in the images, and effectively extracting the spatial structure features of images. In addition, the existing hardware for HDC deployment mainly focuses on in-memory computing (IMC), application specific integrated circuit (ASIC), or high-capacity field programmable gate array (high-capacity FPGA), which cannot meet the flexibility, small area, and low power requirements of edge-side applications. In this paper, a fully learnable HDC framework with learnable preprocessing, encoding and querying, is proposed to boost the accuracy in computer vision tasks, as well as an ultra-tiny accelerator based on edge-side FPGA which matches the proposed framework. Experiments show that on multiple commonly-used image datasets, the proposed HDC framework has an average computation reduction of 80% compared to other most advanced strategies, while achieves a 1.2% accuracy increase. Evaluation on edge-side FPGA shows that compared to other FPGA based state-of-the-art designs, the proposed accelerator saves more than$10\boldsymbol{\times}$hardware resource and power consumption. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Gong Zhang 0002, Weiqiang Liu 0001 |
IEEE Trans. Computers | 3 |
| 2024 | Edge-Side Fine-Grained Sparse CNN Accelerator With Efficient Dynamic Pruning SchemeabstractWith the rapid development of the Internet of Things (IoT), it has become a common concern of academia and industry to provide real-time high performance services for edge-side applications and to bestow intelligence on massive edge-side devices. Due to the limitations of storage space, volume and power consumption of edge side devices, it is difficult for existing convolutional neural networks with large number of parameters and large amount of computation to match them. Network pruning can effectively alleviate the excessive parameters and computation issues in CNNs. However, fine-grained pruning is not hardware friendly, while other structured pruning schemes will result in a much higher loss of accuracy under the same compression ratio. In this paper, an model compression strategy is given including the proposed efficient fine-grained pruning scheme, a dynamic pruning & training method, and a weight importance judgment method. Depending on this strategy, sparse VGG16 (ResNet50) model can be obtained by training from scratch, and achieves a total of$16\times $compression ratio with 1/32 indexing overhead. Further, a light-weight, high-performance sparse CNN accelerator with modified systolic array is proposed. Implementing VGG16 and ResNet50 on the proposed accelerator, the experimental results show that compared with the most advanced design, the proposed accelerator can achieve 8.13 Frames Per Second (FPS) with$2.17\times $better power efficiency and at most$4.14\times $better calculation density. Bi Wu 0002, Tianyang Yu, Ke Chen 0018, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | A Potential Enabler for High-Performance In-Memory Multi-Bit Arithmetic Schemes With Unipolar Switching SOT-MRAMabstractDue to the physical separation of data processing and storage, the conventional Von Neumann architecture exists excessive data migration overhead to curtail the progress of data-intensive applications. In this way, the Computing-in-Memory (CiM) architecture is proposed. Due to the boolean property of the memory cell, the current CiM mainly focuses on single-bit logic design. For the multi-bit arithmetic design, a prevalent patchwork approach is employed using single-bit logic, leaving the design with insufficient parallelism. This paper proposes a high-performance in-memory multi-bit addition (M-Add) and multiplication (M-Mul) scheme based on unipolar switching SOT-MRAM. For the M-Add scheme, transmission logic-based circuit design is proposed to realize single-step inter-column XOR operations, which is logically fits perfectly the g operator of parallel prefix algorithm. Further, the oBK algorithm is presented to maximize the g operator occupancy. For the M-Mul scheme, mapping the Booth decoder to the control signal of the proposed modified flip-flop queue, only two steps are required to realize the decoding of three encoded signals in parallel. The simulation results indicate the proposed design reduces the latency of N-bit Add (N-bit Mul) by an average of 82.6% (31.5%) compared to state-of-the-art CiM designs. Further, a CNN application based on proposed operations achieves 1.23 TOPS/w on the CIFAR-10 dataset, with an average of 47.57% increase over other CiM designs. Bi Wu 0002, Tianyang Yu, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Toward Efficient Retraining: A Large-Scale Approximate Neural Network Framework With Cross-Layer OptimizationabstractLeveraging approximate multipliers in approximate neural networks (ApproxNNs) can effectively reduce hardware area and power consumption, making them suitable for edge-side applications. However, the propagation of layer-by-layer errors limits the application of approximate multipliers to large-scale ApproxNNs and complex tasks. Currently, retraining techniques that consider approximate multiplication errors are commonly used to compensate for the accuracy loss. However, due to the irregularity of the errors introduced by approximate multiplier, it is difficult for the existing generic acceleration hardware (e.g., GPU) to efficiently simulate its function and accelerate retraining, which thereby leads to a huge retraining overhead in ApproxNNs’ application. In this article, we propose an ApproxNN framework that introduces errors with regular and controlled positions for high-efficiency retraining of large-scale ApproxNNs. An approximate multiplier design that matches this framework is also presented to verify the effectiveness of the proposed ApproxNN framework. Experiment results demonstrate that the proposed ApproxNN framework is able to achieve up to 46$\times$speedup in retraining, and the proposed approximate multiplier reduces area/power-delay product (PDP) by 31%/63% compared to the exact multiplier. Compared with the floating-point neural network (NN) model, an accuracy decrease of only 1.13% is achieved when applied to ResNet50 on ImageNet dataset with only 15-epochs retraining, which surpasses other state-of-the-art designs. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Exact and Approximate Squarers for Error-Tolerant ApplicationsabstractApproximate computing is considered an innovative paradigm with wide applications to high performance and low power systems. These applications have relaxed requirements for accuracy, so they can tolerate errors in results and achieve high performance. In approximate computing, multipliers have been widely studied, but squarers (as similar schemes) have not received much attention. In this paper, an accurate squarer is designed based on a Radix-8 Booth-folding square algorithm to reduce the number of partial products and the depth of the partial product array. Several approximate squarers (R8AS1, R8AS2 and R8AS3) are proposed based on the exact squarer to reduce power and delay. Two approximate partial product generators are also designed to simplify the Radix-8 Booth square encoder in R8AS1 and R8AS2. In addition, approximate compressors with compensation are used in the partial product compression stage to reduce additional area and power consumption in R8AS3. Synthesis results for power, area, and delay at 28 nm CMOS technology are presented. Compared with designs in the technical literature with the same accuracy, the proposed 16-bit designs reduce the PDP by 37%; in general, the PDP is decreased by up to 51%. Finally, the proposed approximate squarers are implemented in a square-law detector as a communication application and achieve an SNR close to 30 dB. Also, the three proposed approximate squarers are applied to the k-means clustering algorithm for machine learning to accomplish high performance in classification. Ke Chen 0018, Chenyu Xu, Haroon Waris, Weiqiang Liu 0001, Paolo Montuschi, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2023 | MLiM: High-Performance Magnetic Logic in-Memory Scheme With Unipolar Switching SOT-MRAMabstractConventional computing architectures based on the von Neumann structure are suffering from the severe ‘memory wall’ issue due to the isolation and speed mismatch between memory and processor. As a promising solution, the concept of logic in-memory (LiM) has been proposed to effectively reduce the overhead of data migration and has been extensively studied in various memory technologies such as SRAM, DRAM, MRAM, ReRAM, etc. Among them, SOT-MRAM combines the advantages of non-volatility, low static power consumption, ultra-fast read/write speed, and high density, has emerged as one of the most promising candidates for low-power LiM implementations. In this paper, four in-memory logic operations, AND, OR, MAJ and full-addition (FA), are proposed based on the Unipolar Switching (US) SOT-MRAM devices. Incorporating the emerging switching behavior of SOT-MRAM, these operations can be performed with the basic memory access operations (read/write) with negligible modifying peripheral circuits. Meanwhile, by optimizing the operation steps, the performance degradation caused by the instability of SOT-MRAM device can be minimized in the proposed LiM architecture. Detailed simulation results show that the proposed design can reduce the latency (energy) of AND, OR operations at least by 71.2%, 74.4% (30.0%, 35.4%) compared with the existing SRAM and STT-MRAM designs. For MAJ and FA operations, the performance is improved by at least 34.7% and 44.8% compared to the existing design. The robustness of our design is demonstrated by the 100% pass of the 1000 samples Monte Carlo simulations for the sufficient switching current margin and the effectiveness of basic operations. Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | Approximate Softmax Functions for Energy-Efficient Deep Neural NetworksabstractApproximate computing has emerged as a new paradigm that provides power-efficient and high-performance arithmetic designs by relaxing the stringent requirement of accuracy. Nonlinear functions (such as softmax, rectified linear unit (ReLU), Tanh, and Sigmoid) are extensively used in deep neural networks (DNNs). However, they incur significant power dissipation due to the high circuit complexity. As DNNs are error-tolerant, the design of approximation-linear functions is possible and desired. In this article, the design of an approximate softmax function (AxSF) is proposed. AxSF is based on a double hybrid structure (DHS). AxSF divides the input of the softmax function into two parts for different processing methods. The most significant bits (MSBs) are processed with lookup tables (LUTs) and an exact restoring array divider (EXDr). Taylor’s expansion and a logarithmic divider are used for the less significant bits (LSBs). An improved DHS (IDHS) is also proposed to reduce the hardware complexity. In IDHS, a novel Booth multiplier is utilized for the hybrid scheme to improve the partial product generation and compression, while the truncated implementation is applied to the divider unit. The proposed DHS and IDHS are compared with existing softmax designs. The results show that the proposed approximate softmax design reduces hardware by 48% and delay by 54% while retaining a high accuracy. Ke Chen 0018, Haroon Waris, Weiqiang Liu 0001, Fabrizio Lombardi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | PAxC: A Probabilistic-oriented Approximate Computing Methodology for ANNsabstractIn spite of the rapidly increasing number of approximate designs in circuit logic stack for Artificial Neural Networks (ANNs) learning. A principled and systematic approximate hardware incorporating domain knowledge is still lacking. As the layer of ANN becomes deeper, the errors introduced by approximate hardware will be accumulated quickly, which can result in unexpected results. In this paper, we propose a probabilistic-oriented approximate computing (PAxC) methodology based on the notion of approximate probability to overcome the conceptual and computational difficulties inherent to probabilistic ANN learning. The PAxC makes use of minimum likelihood error in both circuit and application level to maintain the aggressive approximate datapaths to boost the benefits from the tradeoff between accuracy and energy. Compared with a baseline design, the proposed method significantly reduces the power-delay product (PDP) with a negligible accuracy loss. Simulation and a case study of image processing validate the effectiveness of the proposed methodology. Chenghua Wang, Ke Chen 0018, Weiqiang Liu 0001 |
DATE | 3 |
| 2022 | An Energy-efficient and High-precision Approximate MAC with Distributed Arithmetic CircuitsabstractIn this paper, an approximate distributed arithmetic (DA) based parallel MAC is proposed. First, by adopting three kinds of approximation methods, the novel structure significantly reduces hardware complexity. Then, the result is compensated according to the analysis of the probability to enhance the precision. The hardware and error metric evaluation demonstrates that the proposed MAC achieves 25% power-delay product reduction while maintaining better precision. Finally, the Gaussian Blur application is employed to verify the proposed DA-based MAC with 6dB average PSNR improvement compared with recent state-of-the-art work. Ziying Cui, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Weiqiang Liu 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | Data Stream Oriented Fine-grained Sparse CNN Accelerator with Efficient Unstructured Pruning StrategyabstractNetwork pruning can effectively alleviate the excessive parameters and computation issues in CNNs. However, unstructured pruning is not hardware friendly, while structured pruning will result in a significant loss of accuracy. In this paper, an unstructured fine-grained pruning strategy is proposed and achieves a 16X compression ratio with a top-1 accuracy loss of 1.4% for VGG-16. Combined with the proposed hardware-oriented hyperparameter selection method, compression rates of up to 64X can be obtained while fully meeting the edge-side accuracy requirements. Further, a light-weight, high-performance sparse CNN accelerator with modified systolic array is proposed for pruned VGG-16. The experimental results show that compared with the most advanced design, the proposed accelerator can achieve 21 Frames Per Second (FPS) with 3X better power efficiency and 2.19X better calculation density. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Energy-efficient Oriented Approximate Quantization Scheme for Fine-Grained Sparse Neural Network AccelerationabstractFor edge-side applications with severe power constraints, using pruning and quantization to compress models while maintaining network accuracy has become a widely deployed form of Convolutional Neural Networks (CNNs). For the same model accuracy, fine-grained non-regular pruning can bring higher model compression rate than coarse-grained regular pruning, but also introduces a larger indexing overhead. Besides, this overhead increases dramatically as the pruning granularity decreases. In this work, an approximate quantization scheme for fine-grained pruning is proposed. By reusing part of the quantized data bits, the proposed scheme can merge quantized data and indexes approximately, reducing the indexing overhead as well. Meanwhile, since the approximate compensation of index bits, the proposed scheme achieves an effective model accuracy improvement compared to the case of direct quantization to low bit-width. Experimental results show that, for 2:4 fine-grained pruning and 8-bit quantization scenario, the proposed method can save 20% of memory space and transmission cost. Compared with the direct quantization to 6-bit approach, the proposed scheme improves the accuracy by nearly 0.5% in the simulation of ImageNet dataset on ResNet50, despite occupying the same storage space. When deploying Yolov2-tiny at 16 × compression ratio, the energy efficiency of the CNN accelerator with the proposed approximate quantization is 1.33-3.82× that of other state-of-the-art designs. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
ICCD | 3 |
| 2022 | Editorial Special Issue on Circuits and Systems for Emerging Computing ParadigmsabstractAS Dennard’s law is coming to an end, on-chip power consumption reduction and throughput improvement due to technology scaling pose serious challenges; workloads of today’s applications (such as AI, big data, and the IoT) have also reached extremely high levels of complex computation. Power dissipation has become the fundamental barrier to scale computing performance across all technology platforms. Computation at nanoscales requires innovative approaches. Shanshan Liu 0001, Bi Wu 0002, Ke Chen 0018, Weiqiang Liu 0001, Máire O'Neill, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | GBC: An Energy-Efficient LSTM Accelerator With Gating Units Level Balanced Compression StrategyabstractRecurrent Neural Networks (RNNs) have emerged as one of the most popular neural networks for processing time-series problems, widely used in machine translation, automatic speech recognition, and other natural language processing applications. However, conventional RNNs suffered from vanishing and exploding gradients, resulting in poor network performance in applications with long-term input information. As a variant of RNN, Long Short-Term Memory (LSTM) had been proposed to tackle this issue. Nevertheless, at the same time, LSTM introduces gating units and many additional parameters, which makes it challenging to be implemented directly on resource-limited platforms, such as Field Programmable Gate Arrays (FPGAs). This work first investigated the overall maximum achievable compression rates of different gating units and their correlations. Then, Gating Units Level Balanced Compression (GBC) strategy is proposed. After Top-$k$pruning, the proposed GBC strategy can attain a compression rate of$36.6\times $for LSTM. Further, the theoretical analysis indicates that for the existing gating units level LSTM compression variants, the GBC strategy still has further potential for compression. A complementary compression of the GBC strategy is performed on the existing coupled-gate LSTM to verify the analysis. Experimental results show that GBC achieves an additional$32\times $(overall$42.7\times $) compression rate with negligible accuracy loss. Finally, hardware experiments conducted on Xilinx ADM-PCIE-7V3 FPGAs also demonstrate that the accelerator designed in this paper achieves an improvement of 7.4%-191.5% in energy efficiency compared to the state-of-the-art designs. Bi Wu 0002, Zhengkuan Wang, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2019 | Efficient Implementations of Reduced Precision Redundancy (RPR) Multiply and Accumulate (MAC)abstractMultiply and Accumulate (MAC) is one of the most common operations in modern computing systems. It is for example used in matrix multiplication and in new computational environments such as those executed on neural networks for deep machine learning. MAC is also used in critical systems that must operate reliably such as object recognition for vehicles. Therefore, MAC implementations must be able to cope with errors that may be caused for example by radiation. A common scheme to deal with soft errors in arithmetic circuits is the use of Reduced Precision Redundancy (RPR). RPR instead of replicating the entire circuit, uses reduced precision copies which significantly reduce the overhead while still being able to correct the largest errors. This paper considers the implementation of RPR Multiply and Accumulate circuits. First, it is shown that the properties of signed integer multiplication (two´s complement format) can be used to make RPR more efficient. Then its principles are extended to the MAC operation by proposing RPR implementations that improve the error correction capabilities with a limited impact on circuit overhead. The proposed schemes have been implemented and tested. The results show that they can significantly reduce the Mean Square Error (MSE) at the output when the circuit is affected by a soft error and the implementation overhead of the proposed schemes is extremely low. Ke Chen 0018, Linbin Chen, Pedro Reviriego, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2018 | Design and Application of an Approximate 2-D Convolver with Error CompensationabstractThis paper proposes an error compensation scheme of two-dimensional (2D) convolver in which both approximate circuit- and algorithm-level techniques are utilized in the design. Truncation and voltage scaling are used as circuit techniques, while bit-width reduction is utilized at the algorithm level. These different techniques are related to the configuration of the convolver by which its operation can be configured to meet different and often contrasting figures of merit. An extensive evaluation of different error metrics is performed. An error analysis is also presented to substantiate the simulation results; an error compensation scheme is introduced to remedy a loss of accuracy in computation. Convolution for image processing is treated in detail to show the effectiveness of the proposed approach. The design, the analysis and the simulation results show that the approximate techniques utilized in the inexact convolver can operate in synergy. Ke Chen 0018, Jie Han 0001, Paolo Montuschi, Weiqiang Liu 0001, Fabrizio Lombardi |
ISCAS | 1 |
| 2017 | Two Approximate Voting Schemes for Reliable ComputingabstractThis paper relies on the principles of inexact computing to alleviate the issues arising in static masking by voting for reliable computing in the nanoscales. Two schemes that utilize in different manners approximate voting, are proposed. The first scheme is referred to as inexact double modular redundancy (IDMR). IDMR does not resort to triplication, thus saving overhead due to modular replication. This scheme is crudely adaptive in its operation, i.e., it allows a threshold to determine the validity of the module outputs. IDMR operates by initially establishing the difference between the values of the outputs of the two modules; only if the difference is below a preset threshold, then the voter calculates the average value of the two module outputs. The second scheme (ITDMR) combines IDMR with TMR (triple modular redundancy) by using novel conditions in the comparison of the outputs of the three modules. Within an inexact framework, the majority is established using different criteria; in ITDMR, adaptive operation is carried further than IDMR to include approximate voting in a pairwise fashion. So, the validity of the three inputs is established and when only two of the three inputs satisfy the threshold condition, the IDMR operation is utilized. An extensive analysis that includes the voting circuits as well as a probabilistic framework is included. The proposed IDMR and ITDMR schemes improve the power dissipation and tolerance to variations compared to a traditional TMR. To further validate the applicability of the proposed schemes, inexact voting has been used in two applications (image processing and FIR filtering); the simulation results show that performance is substantially improved over TMR. Ke Chen 0018, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Computers | 1 |
| 2015 | An approximate voting scheme for reliable computing
Ke Chen 0018, Fabrizio Lombardi, Jie Han 0001 |
DATE | 1 |
| 2015 | On the Nonvolatile Performance of Flip-Flop/SRAM Cells With a Single MTJabstractIn this brief, three nonvolatile flip-flop (FF)/SRAM cells that utilize a single magnetic tunneling junction (MTJ) as nonvolatile resistive element are proposed. These cells have the same core (i.e., 6T) but they employ different numbers of MOSFETs to implement the so-called instantly ON, normally OFF mode of operation. The additional transistors are utilized for the restore operation to ensure that the data stored in the nonvolatile circuitry can be written back into the FF core once the power is made available. These three cells (7T, 9T, and 11T) are extensively analyzed in terms of their operations in 32 nm technology, such as operational delays (for the write, read, and restore operations), the static noise margin (SNM), critical charge and process variations (in both the MOSFETs and the resistive element). Simulation results show that an increase in the number of MOSFETs in the cells causes improvements in critical charge and tolerance to process variations at the expense of an increase in power dissipation. The SNM and the delay of the restore operation, however, do not necessarily increase with the number of MOSFETs in the cell, but rather on the control of access to the storage nodes from the single MTJ. Ke Chen 0018, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | On the Restore Operation in MTJ-Based Nonvolatile SRAM CellsabstractThis brief investigates the Restore mechanism of a nonvolatile static random access memory (NVSRAM) cell that utilizes two magnetic tunneling junctions (MTJs) as nonvolatile resistive elements and a 6T SRAM core. Two cells are proposed by employing different mechanisms for the Restore operation once the power is reestablished. The proposed cells use the bitline and supply as mechanisms to initiate the Restore operation, so connecting the two MTJs to different nodes of the NVSRAM circuitry. The cells are extensively analyzed in terms of their operations with respect to different figures of merit, such as operational delays (for the Write, Read, and Restore operations), the static noise margin, power consumption, critical charge, and process variations (in both the MOSFETs and the resistive elements). Simulation results show that the cell with the MTJs connected to the supply offers the best performance in terms of power for the Read/Restore operations; it also achieves the best Read delay, but the worst Restore delay. Ke Chen 0018, Jie Han 0001, Fabrizio Lombardi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |