VLDB 2026 Research / reviewers in the wild / expert
Bi Wu 0002
dblp:99/10047-2
· DBLP profile ↗
33ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0001-9972-0478ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 7 first-author · 24 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BIHDC: A Retrainable Fully-Binary Hyperdimensional Computing Accelerator for Edge FPGAsabstractHyperdimensional Computing (HDC) is a lightweight machine learning paradigm characterized by low complexity, efficient learning, and strong interpretability, making it well suited for edge intelligence. However, its accuracy on 2D image tasks still lags behind deep neural networks (DNNs), and existing HDC hardware often relies on in-memory computing or high-end FPGAs, which fail to meet the strict low-power and small-area requirements of edge devices and generally lack complete retraining capabilities. To address these limitations, we propose BIHDC, a lightweight HDC framework that introduces a learnable preprocessing scheme to enhance feature extraction and implements an edge FPGA-based fully binary accelerator supporting end-to-end processing, including preprocessing, encoding, training, retraining, and testing. Experimental results show that BIHDC improves classification accuracy by up to 5% compared to baseline HDC with binarized encoding and by 1.5% over baseline HDC with original encoding across four image datasets. Implemented on the Xilinx Zynq7000 platform, BIHDC reduces resource utilization and power consumption by 90% and 70%, respectively, providing an efficient and scalable solution for resource-constrained edge applications. Changzhen Han, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Weiqiang Liu 0001 |
ASP-DAC | 3 |
| 2026 | An 8T-SRAM Near-Memory Architecture for Multiplierless Approximate DCT
Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Lixia Han, Chenghua Wang, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2026 | A Low Complexity BPF Polar Decoder Design with Sensitivity-Aware LUT Compression Framework
Bozhi Xiu, Yanjing Zhang, Chenggang Yan 0002, Ke Chen 0018, Bi Wu 0002, Weiqiang Liu 0001 |
ISCAS | 5 |
| 2026 | A quality-configurable approximate cache design based on NAND-like SOT MRAM with high energy efficiency
Zhengyi Hou, Luyao Shi, Bi Wang 0002, Bi Wu 0002, Lirida A. B. Naviner, Zhaohao Wang |
Integr. | 5 |
| 2026 | Neuromorphic Hyperdimensional Computing for Efficiently Processing Event-Based DataabstractThe neuromorphic sensor’s event-based data output offers significant benefits, including minimal data redundancy and exceptional time resolution, which guarantee low power consumption and heightened sensitivity during the data acquisition process. Spiking neural network (SNN), with its inherent event-driven characteristic, is well-suited for processing event-based data, and its spike-based computing mechanism enhances the efficiency of data processing. Recent studies are exploring the integration of brain-inspired hyperdimensional computing (HDC) with SNN to leverage HDC’s advantages, including the low inference and training complexity, aiming to further reduce hardware overhead associated with SNN deployment. However, existing works have not effectively harnessed the information output by SNN during hyperdimensional encoding, leading to considerable area and energy overhead. In this article, an efficient neuromorphic HDC method is proposed, featuring a simplified hyperdimensional encoding approach that considers the temporal dynamics of SNN. In addition, a lightweight accelerator design matching the proposed method is also given. Experimental results show that the proposed accelerator achieves over 50% area reduction and reduces energy consumption by 20%–90%. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Gong Zhang 0002, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | PreDAC: An Efficient Framework of Pre-Refining Enhanced Design Space Exploration for Approximate ComputingabstractApproximate computing has emerged as a promising solution in energy-efficiency applications. Recently, attention has shifted from approximate components to Design Space Exploration (DSE) algorithms. However, traditional DSE algorithms face challenges in efficiently obtaining optimal solutions within large and complex design spaces. This paper introduces a prerefining enhanced design space exploration framework that provides customized design space and cost-performance formula for applications. Experimental results demonstrate that integrating this pre-refining step into various DSE algorithms leads to substantial performance gains, including up to $87 \times$ speedup and a 23% improvement in hardware overhead. Moreover, the innovative cost-performance-based DSE algorithm attains a $7.7 \times$ acceleration and further optimizes hardware metrics by an additional 8.8% compared to advanced frameworks employing the same pre-refinement. Ziying Cui, Ke Chen 0018, Bi Wu 0002, Yu Gong 0002, Chenggang Yan 0002, Weiqiang Liu 0001 |
DAC | 3 |
| 2025 | MIRACLE: Multimodal Information Retrieval via a Combined In-Memory Processing and Content Addressable Memory ApproachabstractThe rapid advancement of information technology has brought multimodal information retrieval into the research spotlight. Neural networks, particularly Transformers, have emerged as the dominant solution for extracting multimodal feature vectors. While neural network acceleration has been extensively explored, the subsequent retrieval stage in multimodal scenarios remains under-optimized. Conventional retrieval approaches, such as cosine similarity sorting on von Neumann architectures, suffer from significant data migration and computational inefficiencies. Hashing methods enhance storage and computation efficiency but encounter challenges in energy-efficient implementation and mitigating accuracy losses due to modal heterogeneity. This paper presents a hybrid architecture that integrates in-memory processing (PIM) and content-addressable memory (CAM) to address these challenges. Transformer-extracted features are processed via in-memory random hashing leveraging device-intrinsic properties, with CAM facilitating parallel search space reduction. A final cosine similarity reranking stage refines the results while balancing accuracy with energy efficiency. Experimental evaluations validate that the proposed method, when compared to the baseline traditional CPU-based cosine similarity retrieval, 1) achieves almost identical level of accuracy, dramatically outperforming other pure CAMbased Hamming distance retrieval approaches; and 2) reduces latency by $9.45 \times$ and energy consumption by $30.20 \times$. Xuehui Liu, Tianyang Yu, Shuo Ran, Bi Wu 0002, Xiaotao Jia, Weiqiang Liu 0001, Gang Qu 0001, Weisheng Zhao 0001 |
DAC | 6 |
| 2025 | High-Performance Co-Processing Architecture Using SOT-MRAM-Based In-memory Computing SchemeabstractIn recent years, the rapid advancement of processor performance has highlighted the increasing inadequacy of memory bandwidth. To mitigate this challenge, the In-memory Computing (IMC) architecture has been proposed. Among various IMC implementations, Spin-Orbit Torque Random Access Memory (SOT-MRAM) stands out as an ideal generic co-processor due to its high density and low leakage current. However, how to efficiently integrate with existing instruction set architectures (ISA) while satisfying the characteristics of the SOT-MRAM IMC hardware poses a challenge. In this work, an instruction-driven in-memory co-processor architecture is proposed. By combining the proposed hardware-software collaborative memory redirection scheme, the SOT-MRAM-based in-memory computing can be merged into the existing computing architecture without disturbing ISA. Further, to satisfy the inter-row operation computation characteristics of SOT-MRAM IMC, a data pre-scheduling scheme oriented to efficient computation is proposed. Experimental results demonstrate a 6.2x speedup and 64% power reduction compared to a CPU-only architecture, and a 28% improvement in acceleration ratio over the state-of-the-art IMC architecture. Bi Wu 0002, Ke Chen 0018, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2025 | Learning-Based Realtime Synthetic Aperture Radar Imaging for Embedded System on SatelliteabstractSpaceborne Synthetic Aperture Radar (SAR), due to its ability of all-weather and all-day sensing, is extensively utilized across various fields. Realtime imaging on satellite is highly advantageous as it significantly reduces the communication cost and delay between satellite and ground station, making it exceptionally suitable for time-sensitive applications like maritime search and rescue. However, the limited computing power of embedded system on satellite, together with the substantial computational demands of traditional imaging algorithms, pose challenges for realtime imaging. To address these issues, a lightweight learning-based model with adjustable complexity is proposed for realtime imaging. Furthermore, we collect a large amount of real-world echo data from satellite and construct the first large-scale dataset, for training and evaluating the learning-based SAR imaging task. Experimental results and evaluation in the embedded system show that, the proposed learning-based model achieves up to 19× speedup than traditional algorithm. Tianyang Yu, Bi Wu 0002, Weiqiang Liu 0001 |
ISCAS | 2 |
| 2025 | LAHDC: Logic-Aggregation-Based Query for Embedded Hyperdimensional Computing AcceleratorabstractWith low complexity and robustness, hyperdimensional computing (HDC) has become a promising paradigm for edge-side applications. HDC employs hypervectors (generally with 2–10 K dimensions) to represent input samples, and performs logical operations in hyperdimensional space to complete perceptual tasks. Compared to deep neural network (DNN), HDC is more suitable for lightweight edge-side applications (i.e., speech, activity recognition), due to its low complexity and less computational scheduling. However, existing HDC’s querying process relies on trained class hypervectors, resulting in on-chip storage and transmission overhead which limits the application of ASIC-based or FPGA-based HDC accelerators in embedded systems. In this article, a logic-aggregation-based query method called LAHDC is proposed to eliminate such overhead. In addition, an ultratiny HDC accelerator design matching LAHDC is also proposed, as well as an automated tool to search for optimal structure and generate hardware design code. Experimental results show that, compared to existing ASIC-based HDC accelerators, the proposed design reduce the area/energy by more than 95%/80%. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Gong Zhang 0002, Weiqiang Liu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | SS-MRAM: A Segment-Based Search Scheme With Configurable Matching for High-Accuracy Hyperdimensional Computing in CAM ApplicationsabstractWith the rise of hyperdimensional computing (HDC), content-addressable memories (CAMs) have emerged as an ideal choice for processing high-dimensional data. However, despite the advantages of high parallelism and low latency offered by CAM technology, it fails to address the significant loss of inference accuracy caused by closely matching hamming distances (HD). Emerging analog-based imprecise in-memory computing technologies frequently provide a minimum detectable HD that is insufficient for meeting the requirements of high similarity tasks. This limitation provides opportunities for using digital methods to realize fully exact matching memory computing technology based on magnetic random access memory (MRAM). In this work, a segment search scheme based on STT-MRAM devices and address-matching technology is proposed, which achieves zero loss in inference accuracy. An adaptive amplification structure is initially implemented by integrating the discharge method of latch structures, accompanied by the design of a 14T-2MTJ cell circuit. Through the optimization of the matching step, HD are calculated for each segment of configurable-dimensional vectors, ultimately facilitating the classification ranking of query hypervectors based on a comprehensive array architecture. Experimental results indicate that there is zero loss in inference accuracy when each segment of the dimension is configured to be below 16 bits. At 16 bits, the search power consumption in the worst-case matching scenario is measured at 1.73 fJ/bit, while the loss in inference accuracy does not exceed 0.4%. Bi Wu 0002, Shuo Ran, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | A Time Efficient Comprehensive Model of Approximate Multipliers for Design Space ExplorationabstractMultipliers play an essential role in various data processing applications and have garnered significant attention in approximate computing (AxC) for their energy-efficient features. However, formulating a precise error model for approximate data processing algorithms in conjunction with hardware metrics presents a challenge, leading to substantial time consumption in the design space exploration. This paper introduces an analytical model for approximate multipliers while considering input patterns. This model furnishes accurate error metrics, along with high-precision hardware metrics for various approximate multiplier configurations, impervious to variations in input data distribution. The proposed error model reduces the runtime by an average factor of 120.85 and, in some instances, by as much as 2,500 times, when contrasted with simulation-based methods. The design space exploration is performed on a 3×3 convolution circuit, revealing a comparable Pareto-optimal set and substantial reductions of up to 79.46% in the Power-Delay-Product (PDP) and 71.98% in area compared to the accurate counterpart. Additionally, the result of the Gaussian Blur application experiment demonstrates a 68.59% reduction in PDP and a 56.21% reduction in area, all while maintaining a PSNR of 30 dB. Ziying Cui, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Yu Gong 0002, Weiqiang Liu 0001 |
ARITH | 3 |
| 2024 | A Combined Content Addressable Memory and In-Memory Processing Approach for k-Clique Counting Accelerationabstractk-Clique counting problem plays an important role in graph mining which has seen a growing number of applications. However, current k-Clique counting accelerators cannot meet the performance requirement mainly because they struggle with high data transfer issue incurred by the intensive set intersection operations and the inability of load balancing. In this paper, we propose to solve this problem with a hybrid framework of content addressable memory (CAM) and in-memory processing (PIM). Specifically, we first utilize CAM for binary induced subgraph generation in order to reduce the search space, then we use PIM to implement in-place parallel k-Clique counting through iterative Boolean logic "AND" like operation. To take full advantage of this combined CAM and PIM framework, we develop dynamic task scheduling strategies that can achieve near optimal load balancing among the PIM arrays. Experimental results demonstrate that, compared with state-of-the-art CPU and GPU platforms, our approach achieves speedups of 167.5× and 28.8×, respectively. Meanwhile, the energy efficiency is improved by 788.3× over the GPU baseline. Xidi Ma, Tianyang Yu, Bi Wu 0002, Gang Qu 0001, Weisheng Zhao 0001 |
DAC | 5 |
| 2024 | Most Significant One-Driven Shifting Dynamic Efficient Multipliers for Large Language ModelsabstractLarge Language Models (LLMs) have demonstrated exceptional performance but demand significantly more computational power and memory compared to Deep Neural Networks (DNNs). This necessitates the development of more energy-efficient hardware designs. This paper introduces a novel weight approximation strategy for quantized LLMs, resulting in the creation of a highly efficient approximate multiplier based on Most Significant One (MSO) shifting. When compared to energy-efficient approximate logarithmic multipliers and precision-demanding approximate non-logarithmic multipliers, the proposed design strikes an optimal balance between accuracy and hardware cost. It maintains a superior level of accuracy while incurring hardware costs comparable to logarithmic multipliers and, in some cases, even outperforming them. In particular, when compared to exact multiplier, the proposed design achieves significant reductions, including up to a 28.31% reduction in area, a 57.84% decrease in power consumption, and an 11.86% reduction in delay. The experiments demonstrate that the proposed multiplier in DNNs can save approximately 60% of energy without compromising task accuracy. Similarly, experiments on the Transformer accelerator indicate substantial energy savings for LLMs using the proposed design. Ke Chen 0018, Bi Wu 0002, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2024 | Fully Learnable Hyperdimensional Computing Framework With Ultratiny Accelerator for Edge-Side ApplicationsabstractBrain-inspired hyperdimensional computing (HDC) is a new computational paradigm that encodes input sample into a hypervector (generally with dimensions of$2K-10K$), and performs simple arithmetic and logic operations in the hyperdimensional space to complete perceptual tasks like human brain. Due to its simplicity, interpretability, and robustness, HDC has gradually become a competitor and substitute for deep neural network (DNN) in many tasks. However, there exists an accuracy gap between existing heuristic HDC algorithms and DNN in computer vision tasks, as existing encoding methods have difficulty in filtering out large amount of background and noise in the images, and effectively extracting the spatial structure features of images. In addition, the existing hardware for HDC deployment mainly focuses on in-memory computing (IMC), application specific integrated circuit (ASIC), or high-capacity field programmable gate array (high-capacity FPGA), which cannot meet the flexibility, small area, and low power requirements of edge-side applications. In this paper, a fully learnable HDC framework with learnable preprocessing, encoding and querying, is proposed to boost the accuracy in computer vision tasks, as well as an ultra-tiny accelerator based on edge-side FPGA which matches the proposed framework. Experiments show that on multiple commonly-used image datasets, the proposed HDC framework has an average computation reduction of 80% compared to other most advanced strategies, while achieves a 1.2% accuracy increase. Evaluation on edge-side FPGA shows that compared to other FPGA based state-of-the-art designs, the proposed accelerator saves more than$10\boldsymbol{\times}$hardware resource and power consumption. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Gong Zhang 0002, Weiqiang Liu 0001 |
IEEE Trans. Computers | 2 |
| 2024 | Edge-Side Fine-Grained Sparse CNN Accelerator With Efficient Dynamic Pruning SchemeabstractWith the rapid development of the Internet of Things (IoT), it has become a common concern of academia and industry to provide real-time high performance services for edge-side applications and to bestow intelligence on massive edge-side devices. Due to the limitations of storage space, volume and power consumption of edge side devices, it is difficult for existing convolutional neural networks with large number of parameters and large amount of computation to match them. Network pruning can effectively alleviate the excessive parameters and computation issues in CNNs. However, fine-grained pruning is not hardware friendly, while other structured pruning schemes will result in a much higher loss of accuracy under the same compression ratio. In this paper, an model compression strategy is given including the proposed efficient fine-grained pruning scheme, a dynamic pruning & training method, and a weight importance judgment method. Depending on this strategy, sparse VGG16 (ResNet50) model can be obtained by training from scratch, and achieves a total of$16\times $compression ratio with 1/32 indexing overhead. Further, a light-weight, high-performance sparse CNN accelerator with modified systolic array is proposed. Implementing VGG16 and ResNet50 on the proposed accelerator, the experimental results show that compared with the most advanced design, the proposed accelerator can achieve 8.13 Frames Per Second (FPS) with$2.17\times $better power efficiency and at most$4.14\times $better calculation density. Bi Wu 0002, Tianyang Yu, Ke Chen 0018, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | A Potential Enabler for High-Performance In-Memory Multi-Bit Arithmetic Schemes With Unipolar Switching SOT-MRAMabstractDue to the physical separation of data processing and storage, the conventional Von Neumann architecture exists excessive data migration overhead to curtail the progress of data-intensive applications. In this way, the Computing-in-Memory (CiM) architecture is proposed. Due to the boolean property of the memory cell, the current CiM mainly focuses on single-bit logic design. For the multi-bit arithmetic design, a prevalent patchwork approach is employed using single-bit logic, leaving the design with insufficient parallelism. This paper proposes a high-performance in-memory multi-bit addition (M-Add) and multiplication (M-Mul) scheme based on unipolar switching SOT-MRAM. For the M-Add scheme, transmission logic-based circuit design is proposed to realize single-step inter-column XOR operations, which is logically fits perfectly the g operator of parallel prefix algorithm. Further, the oBK algorithm is presented to maximize the g operator occupancy. For the M-Mul scheme, mapping the Booth decoder to the control signal of the proposed modified flip-flop queue, only two steps are required to realize the decoding of three encoded signals in parallel. The simulation results indicate the proposed design reduces the latency of N-bit Add (N-bit Mul) by an average of 82.6% (31.5%) compared to state-of-the-art CiM designs. Further, a CNN application based on proposed operations achieves 1.23 TOPS/w on the CIFAR-10 dataset, with an average of 47.57% increase over other CiM designs. Bi Wu 0002, Tianyang Yu, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Toward Efficient Retraining: A Large-Scale Approximate Neural Network Framework With Cross-Layer OptimizationabstractLeveraging approximate multipliers in approximate neural networks (ApproxNNs) can effectively reduce hardware area and power consumption, making them suitable for edge-side applications. However, the propagation of layer-by-layer errors limits the application of approximate multipliers to large-scale ApproxNNs and complex tasks. Currently, retraining techniques that consider approximate multiplication errors are commonly used to compensate for the accuracy loss. However, due to the irregularity of the errors introduced by approximate multiplier, it is difficult for the existing generic acceleration hardware (e.g., GPU) to efficiently simulate its function and accelerate retraining, which thereby leads to a huge retraining overhead in ApproxNNs’ application. In this article, we propose an ApproxNN framework that introduces errors with regular and controlled positions for high-efficiency retraining of large-scale ApproxNNs. An approximate multiplier design that matches this framework is also presented to verify the effectiveness of the proposed ApproxNN framework. Experiment results demonstrate that the proposed ApproxNN framework is able to achieve up to 46$\times$speedup in retraining, and the proposed approximate multiplier reduces area/power-delay product (PDP) by 31%/63% compared to the exact multiplier. Compared with the floating-point neural network (NN) model, an accuracy decrease of only 1.13% is achieved when applied to ResNet50 on ImageNet dataset with only 15-epochs retraining, which surpasses other state-of-the-art designs. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | MLiM: High-Performance Magnetic Logic in-Memory Scheme With Unipolar Switching SOT-MRAMabstractConventional computing architectures based on the von Neumann structure are suffering from the severe ‘memory wall’ issue due to the isolation and speed mismatch between memory and processor. As a promising solution, the concept of logic in-memory (LiM) has been proposed to effectively reduce the overhead of data migration and has been extensively studied in various memory technologies such as SRAM, DRAM, MRAM, ReRAM, etc. Among them, SOT-MRAM combines the advantages of non-volatility, low static power consumption, ultra-fast read/write speed, and high density, has emerged as one of the most promising candidates for low-power LiM implementations. In this paper, four in-memory logic operations, AND, OR, MAJ and full-addition (FA), are proposed based on the Unipolar Switching (US) SOT-MRAM devices. Incorporating the emerging switching behavior of SOT-MRAM, these operations can be performed with the basic memory access operations (read/write) with negligible modifying peripheral circuits. Meanwhile, by optimizing the operation steps, the performance degradation caused by the instability of SOT-MRAM device can be minimized in the proposed LiM architecture. Detailed simulation results show that the proposed design can reduce the latency (energy) of AND, OR operations at least by 71.2%, 74.4% (30.0%, 35.4%) compared with the existing SRAM and STT-MRAM designs. For MAJ and FA operations, the performance is improved by at least 34.7% and 44.8% compared to the existing design. The robustness of our design is demonstrated by the 100% pass of the 1000 samples Monte Carlo simulations for the sufficient switching current margin and the effectiveness of basic operations. Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | An Energy-efficient and High-precision Approximate MAC with Distributed Arithmetic CircuitsabstractIn this paper, an approximate distributed arithmetic (DA) based parallel MAC is proposed. First, by adopting three kinds of approximation methods, the novel structure significantly reduces hardware complexity. Then, the result is compensated according to the analysis of the probability to enhance the precision. The hardware and error metric evaluation demonstrates that the proposed MAC achieves 25% power-delay product reduction while maintaining better precision. Finally, the Gaussian Blur application is employed to verify the proposed DA-based MAC with 6dB average PSNR improvement compared with recent state-of-the-art work. Ziying Cui, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Weiqiang Liu 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2022 | Data Stream Oriented Fine-grained Sparse CNN Accelerator with Efficient Unstructured Pruning StrategyabstractNetwork pruning can effectively alleviate the excessive parameters and computation issues in CNNs. However, unstructured pruning is not hardware friendly, while structured pruning will result in a significant loss of accuracy. In this paper, an unstructured fine-grained pruning strategy is proposed and achieves a 16X compression ratio with a top-1 accuracy loss of 1.4% for VGG-16. Combined with the proposed hardware-oriented hyperparameter selection method, compression rates of up to 64X can be obtained while fully meeting the edge-side accuracy requirements. Further, a light-weight, high-performance sparse CNN accelerator with modified systolic array is proposed for pruned VGG-16. The experimental results show that compared with the most advanced design, the proposed accelerator can achieve 21 Frames Per Second (FPS) with 3X better power efficiency and 2.19X better calculation density. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2022 | Energy-efficient Oriented Approximate Quantization Scheme for Fine-Grained Sparse Neural Network AccelerationabstractFor edge-side applications with severe power constraints, using pruning and quantization to compress models while maintaining network accuracy has become a widely deployed form of Convolutional Neural Networks (CNNs). For the same model accuracy, fine-grained non-regular pruning can bring higher model compression rate than coarse-grained regular pruning, but also introduces a larger indexing overhead. Besides, this overhead increases dramatically as the pruning granularity decreases. In this work, an approximate quantization scheme for fine-grained pruning is proposed. By reusing part of the quantized data bits, the proposed scheme can merge quantized data and indexes approximately, reducing the indexing overhead as well. Meanwhile, since the approximate compensation of index bits, the proposed scheme achieves an effective model accuracy improvement compared to the case of direct quantization to low bit-width. Experimental results show that, for 2:4 fine-grained pruning and 8-bit quantization scenario, the proposed method can save 20% of memory space and transmission cost. Compared with the direct quantization to 6-bit approach, the proposed scheme improves the accuracy by nearly 0.5% in the simulation of ImageNet dataset on ResNet50, despite occupying the same storage space. When deploying Yolov2-tiny at 16 × compression ratio, the energy efficiency of the CNN accelerator with the proposed approximate quantization is 1.33-3.82× that of other state-of-the-art designs. Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
ICCD | 2 |
| 2022 | Editorial Special Issue on Circuits and Systems for Emerging Computing ParadigmsabstractAS Dennard’s law is coming to an end, on-chip power consumption reduction and throughput improvement due to technology scaling pose serious challenges; workloads of today’s applications (such as AI, big data, and the IoT) have also reached extremely high levels of complex computation. Power dissipation has become the fundamental barrier to scale computing performance across all technology platforms. Computation at nanoscales requires innovative approaches. Shanshan Liu 0001, Bi Wu 0002, Ke Chen 0018, Weiqiang Liu 0001, Máire O'Neill, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | GBC: An Energy-Efficient LSTM Accelerator With Gating Units Level Balanced Compression StrategyabstractRecurrent Neural Networks (RNNs) have emerged as one of the most popular neural networks for processing time-series problems, widely used in machine translation, automatic speech recognition, and other natural language processing applications. However, conventional RNNs suffered from vanishing and exploding gradients, resulting in poor network performance in applications with long-term input information. As a variant of RNN, Long Short-Term Memory (LSTM) had been proposed to tackle this issue. Nevertheless, at the same time, LSTM introduces gating units and many additional parameters, which makes it challenging to be implemented directly on resource-limited platforms, such as Field Programmable Gate Arrays (FPGAs). This work first investigated the overall maximum achievable compression rates of different gating units and their correlations. Then, Gating Units Level Balanced Compression (GBC) strategy is proposed. After Top-$k$pruning, the proposed GBC strategy can attain a compression rate of$36.6\times $for LSTM. Further, the theoretical analysis indicates that for the existing gating units level LSTM compression variants, the GBC strategy still has further potential for compression. A complementary compression of the GBC strategy is performed on the existing coupled-gate LSTM to verify the analysis. Experimental results show that GBC achieves an additional$32\times $(overall$42.7\times $) compression rate with negligible accuracy loss. Finally, hardware experiments conducted on Xilinx ADM-PCIE-7V3 FPGAs also demonstrate that the accelerator designed in this paper achieves an improvement of 7.4%-191.5% in energy efficiency compared to the state-of-the-art designs. Bi Wu 0002, Zhengkuan Wang, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | An Energy Efficient Accelerator for Bidirectional Recurrent Neural Networks (BiRNNs) Using Hybrid-Iterative Compression With Error SensitivityabstractRecurrent Neural Networks (RNNs) have been widely used in many sequential applications, such as machine translation, speech recognition and sentiment analysis. Long Term Short Term Memory (LSTM) and Gated Recurrent Unit (GRU) are widely used variants of RNN due to their effectiveness in overcoming gradient vanishing and exploding problems; however, compared to conventional RNN, their massive storage and computation requirements hinder their application. In addition, the recurrent structure of RNNs makes them prone to accumulate errors, resulting in a severe loss of accuracy. In this work, we propose a hybrid-iterative compression (HIC) algorithm for LSTM/GRU. By exploiting the error sensitivity of RNN, the gating units are divided into error-sensitive and error-insensitive groups, that are compressed using different algorithms. By using this approach, a 37.1×/32.3× compression ratio is achieved with negligible accuracy loss for LSTM/GRU. Further, an energy efficient accelerator for bidirectional RNNs is proposed. In this accelerator, the data flow of the matrix operation unit based on the block structure matrix (MOU-S) is improved through rearranging weights; the utilization of BRAM is improved through a fine-grained parallelism configuration of matrix-vector multiplications (MVMs). Meanwhile, the timing matching strategy alleviates the load-imbalance problem between MOU-S and the matrix operation unit based on top- k pruning (MOU-P). When running at 200MHz on Xilinx ADM-PCIE-7V3 FPGA, the proposed design achieves an improvement in energy efficiency in a range of 5%-237% for LSTM networks, and an improvement of 58% for GRU networks compared with state-of-the-art designs. Guocai Nan, Zhengkuan Wang, Chenghua Wang, Bi Wu 0002, Zhican Wang, Weiqiang Liu 0001, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2020 | A Comparative Cross-layer Study on Racetrack Memories: Domain Wall vs SkyrmionabstractRacetrack memory (RM), a new storage scheme in which information flows along a nanotrack, has been considered as a potential candidate for future high-density storage device instead of hard disk drive (HDD). The first RM technology, which was proposed in 2008 by IBM, relies on a train of opposite magnetic domains separated by domain walls (DWs), named DW-RM. After 10 years of intensive research, a variety of fundamental advancements has been achieved; unfortunately, no product has been available until now. With increasing effort and resources dedicated to the development of DW-RM, it is likely that new materials and mechanisms will soon be discovered for practical applications. However, new concepts might also be on the horizon. Recently, an alternative information carrier, magnetic skyrmion, which was experimentally discovered in 2009, has been regarded as a promising replacement of DW for RM, named skyrmion-based RM (SK-RM). Intensive effort has been involved and amazing advances have been made in observing, writing, manipulating, and deleting individual skyrmions. So, what is the relationship between DW and skyrmion? What are the key differences between DW and skyrmion, or between DW-RM and SK-RM? What benefits could SK-RM bring and what challenges need to be addressed before application? In this review article, we intend to answer these questions through a comparative cross-layer study between DW-RM and SK-RM. This work will provide guidelines, especially for circuit and architecture researchers on RM. Wang Kang 0001, Bi Wu 0002, Xing Chen 0012, Daoqian Zhu, Zhaohao Wang, Xichao Zhang, Youguang Zhang, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2020 | Write Back Energy Optimization for STT-MRAM-based Last-level Cache with Data Pattern CharacterizationabstractTraditional memory technologies face severe challenges in meeting the ever-increasing power and memory bandwidth requirements for high-performance computing and big-data analyses. Several emerging memory technologies are promising as the replacements of SRAM or DRAM. Among them, STT-MRAM can be used to replace SRAM as the last-level cache (LLC). However, it suffers from high write energy and latency. In this article, we investigate data patterns written from SRAM-based upper-level cache to STT-MRAM-based LLC to explore the write energy reduction potential. Depending on the data layout within a cache line, redundant bits can be identified and eliminated from write back operations to save STT-MRAM write energy. We also propose a dynamic profiling method to accommodate different application characteristics. The extensive simulation results show that write energy can be saved by 37.05% ∼ 38.89% for static profiling and 19.76% ∼ 34.29% for dynamic profiling. Keren Liu, Bi Wu 0002, Weisheng Zhao 0001, Yuanqing Cheng, Ying Wang 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal ConsiderationabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. Spin transfer torque magnetic memory (STT-MRAM) is proposed as a promising solution for the low power cache design due to its high integration density and ultralow leakage power. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM, and observe that the temperature can affect the write delay and energy significantly. Then, we explore the nonuniform cache access (NUCA) design of the chip-multiprocessors with STT-MRAM-based last level cache (LLC). A thermal aware data migration policy, called “Thermosiphon,” which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions dynamically based on the thermal distribution monitored by thermal sensors available on-chip, and adaptively migrates write intensive data among different thermal regions considering the thermal gradient. Compared to the conventional NUCA design, our proposed design can save 41.2% write energy at most and 13.01% on average with negligible hardware overhead. Bi Wu 0002, Pengcheng Dai, Yuanqing Cheng, Ying Wang 0001, Jianlei Yang 0001, Zhaohao Wang, Dijun Liu, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | An Adaptive Thermal-Aware ECC Scheme for Reliable STT-MRAM LLC DesignabstractConsidering the insatiable demand for high-performance computing, on-chip cache capacity increases rapidly. Spin-transfer-torque magnetoresistive random-access memory (STT-MRAM) is a promising cache candidate due to ultralow standby power, high-access speed, and integration density. Unfortunately, when the feature size of magnetic tunnel junction (MTJ) scales down to 1 Xnm, read current approaches write current closely, which may result in read disturbance threatening the reliability of STT-MRAM. Furthermore, the elevating on-chip temperature reduces the thermal stability of STT-MRAM remarkably and aggravates the read disturbance. Error correction code (ECC) is an effective technique to enhance memory reliability. In this paper, we take advantage of the thermal dependence of STT-MRAM and propose a thermally adaptive ECC design, called “Chameleon,” that can adjust the ECC protection strength dynamically to reduce the ECC storage overhead and improve the cache access performance and energy efficiency. Experimental results show that compared to the conservative nonadaptive ECC scheme, our design can improve both cache performance and energy consumption effectively. Bi Wu 0002, Yuanqing Cheng, Ying Wang 0001, Dijun Liu, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | An Adaptive 3T-3MTJ Memory Cell Design for STT-MRAM-Based LLCsabstractThe STT-MRAM technology is a promising candidate for future on-chip cache memory because of its high density, low standby power, and nonvolatility. As the technology node scales, especially under 40-nm technology node, STT-MRAM cell design becomes a key issue to approach low power consumption, high access performance, and desirable reliability. The conventional 1T-1 magnetic tunnel junction (MTJ) and 2T-2MTJ cell designs cannot address these challenges efficiently. In this paper, we propose a novel 3T-3MTJ cell structure using the advanced perpendicular MTJ (p-MTJ) technology. It can store 2 bits with three MTJs. The differential sensing technique can be used to read out the most significant bit as fast as the 2T-2MTJ design. The sensing latency of 2 bits within the same cell is almost the same as the sensing latency of the 1T-1MTJ cell design. Therefore, the 3T-3MTJ cell can have the advantages of both 2T-2MTJ and 1T-1MTJ cells. Circuit-level simulations show that the proposed 3T-3MTJ cell structure can achieve a desirable tradeoff between storage density, access performance, and energy consumption compared to the prior 1T-1MTJ and 2T-2MTJ cell structures. Additionally, we propose a novel adaptive cache design based on the 3T-3MTJ cell structure, which can work in different modes to satisfy various memory access demands from different applications. Architecture level simulations validate the effectiveness of the proposed cache design. Linuo Xue, Bi Wu 0002, Yuanqing Cheng, Peiyuan Wang, Chando Park, Jimmy J. Kan, Seung-Hyuk Kang, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Thermosiphon: A thermal aware NUCA architecture for write energy reduction of the STT-MRAM based LLCsabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. STT-MRAM (Spin Transfer Torque Magnetic Memory) is proposed as a promising solution for the low power cache design due to its high integration density and ultra-low leakage. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM and observe that the temperature can affect the write delay and energy significantly. Then, we explore the NUCA (Non-Uniform Cache Access) design of the CMPs (Chip-Multi-Processors)with STT-MRAM based LLC (Last Level Cache). A thermal aware data migration policy, called “Thermosiphon”, which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions based on the thermal distribution and adaptively migrate write intensive data considering the temperature gradient among different thermal regions. Compared to the conventional NUCA design, our proposed design can save 22.5% write energy with negligible hardware overhead. Bi Wu 0002, Yuanqing Cheng, Pengcheng Dai, Jianlei Yang 0001, Youguang Zhang, Dijun Liu, Ying Wang 0001, Weisheng Zhao 0001 |
ICCAD | 1 |
| 2016 | PDS: pseudo-differential sensing scheme for STT-MRAMabstractSTT-MRAM has been considered as one of the most promising nonvolatile memory candidates in the next-generation of computer architecture. However, the read reliability and dynamic write power concerns greatly hinder its practical application. In this paper, we propose a synergistic solution, namely pseudo-differential sensing (PDS), to jointly address these two concerns. Three techniques, including cell cluster, asymmetric sensing amplifier (ASA) and self-error-detection-correction (SEDC), are proposed to implement the PDS concept. Our experimental results show that the PDS scheme with the 3T3MTJ cell cluster can reduce the area (~21.7%) and write power (~25.6%) of the differential sensing (DS) scheme while improve the read reliability (read margin, ~35.6%) of the typical sensing (TS) scheme for a 16 Mbit cache. Furthermore, the PDS scheme with the 1T3MTJ cell cluster can outperform both the TS and DS schemes in terms of area (~40.0%, ~66.1%), read latency (~16.6%, ~32.1%), read power (~16.7%, ~37.1%), write latency (~5.4%, 16.3%) and write power (~18.6%, ~43.4%). Wang Kang 0001, Tingting Pang, Bi Wu 0002, Weifeng Lv, Youguang Zhang, Guangyu Sun 0003, Weisheng Zhao 0001 |
DAC | 3 |
| 2016 | Temperature Impact Analysis and Access Reliability Enhancement for 1T1MTJ STT-RAMabstractSpin-transfer torque magnetic random access memory (STT-RAM) is a promising and emerging technology due to its many advantageous features such as scalability, nonvolatility, density, endurance, and fast access speed. However, the operation of STT-RAM is severely affected by environmental factors such as process variations and temperature. As the temperature rockets up in modern computing systems, it is highly desirable to understand thermal impact on STT-RAM operations and reliability. In this paper, a thermal-aware MTJ model, calibrated and validated by experimental measurements, is proposed as the basis for thoroughly thermal aware analysis of a 1T1MTJ STT-RAM cell structure. Using this model, we investigate temperature effect on memory cell access behavior in terms of access latency, energy, and reliability on a 45-nm technology node. Thermal impact on a more advanced 11-nm technology node is also evaluated in the paper. Additionally, we propose a thermal-aware design for STT-RAM sensing circuit using a body-biasing technique, which can enlarge read margin dramatically to enhance read reliability under temperature variations. Moreover, our proposed technique can suppress read disturbance effectively as well. Experimental results show that our proposed sensing circuit can enlarge read margin by 2.47× when reading “0” and 3.15× when reading “1,” and reduce read disturbance error rate by 55.6% on average. Bi Wu 0002, Yuanqing Cheng, Jianlei Yang 0001, Aida Todri, Weisheng Zhao 0001 |
IEEE Trans. Reliab. | 1 |