Chenggang Yan 0002

dblp:315/0827-2 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0001-5957-3207ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 2 first-author · 17 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BIHDC: A Retrainable Fully-Binary Hyperdimensional Computing Accelerator for Edge FPGAs
abstract
Hyperdimensional Computing (HDC) is a lightweight machine learning paradigm characterized by low complexity, efficient learning, and strong interpretability, making it well suited for edge intelligence. However, its accuracy on 2D image tasks still lags behind deep neural networks (DNNs), and existing HDC hardware often relies on in-memory computing or high-end FPGAs, which fail to meet the strict low-power and small-area requirements of edge devices and generally lack complete retraining capabilities. To address these limitations, we propose BIHDC, a lightweight HDC framework that introduces a learnable preprocessing scheme to enhance feature extraction and implements an edge FPGA-based fully binary accelerator supporting end-to-end processing, including preprocessing, encoding, training, retraining, and testing. Experimental results show that BIHDC improves classification accuracy by up to 5% compared to baseline HDC with binarized encoding and by 1.5% over baseline HDC with original encoding across four image datasets. Implemented on the Xilinx Zynq7000 platform, BIHDC reduces resource utilization and power consumption by 90% and 70%, respectively, providing an efficient and scalable solution for resource-constrained edge applications.
Changzhen Han, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Weiqiang Liu 0001
ASP-DAC4
2026 An 8T-SRAM Near-Memory Architecture for Multiplierless Approximate DCT
Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Lixia Han, Chenghua Wang, Weiqiang Liu 0001
ISCAS4
2026 A Low Complexity BPF Polar Decoder Design with Sensitivity-Aware LUT Compression Framework
Bozhi Xiu, Yanjing Zhang, Chenggang Yan 0002, Ke Chen 0018, Bi Wu 0002, Weiqiang Liu 0001
ISCAS3
2026 Neuromorphic Hyperdimensional Computing for Efficiently Processing Event-Based Data
abstract
The neuromorphic sensor’s event-based data output offers significant benefits, including minimal data redundancy and exceptional time resolution, which guarantee low power consumption and heightened sensitivity during the data acquisition process. Spiking neural network (SNN), with its inherent event-driven characteristic, is well-suited for processing event-based data, and its spike-based computing mechanism enhances the efficiency of data processing. Recent studies are exploring the integration of brain-inspired hyperdimensional computing (HDC) with SNN to leverage HDC’s advantages, including the low inference and training complexity, aiming to further reduce hardware overhead associated with SNN deployment. However, existing works have not effectively harnessed the information output by SNN during hyperdimensional encoding, leading to considerable area and energy overhead. In this article, an efficient neuromorphic HDC method is proposed, featuring a simplified hyperdimensional encoding approach that considers the temporal dynamics of SNN. In addition, a lightweight accelerator design matching the proposed method is also given. Experimental results show that the proposed accelerator achieves over 50% area reduction and reduces energy consumption by 20%–90%.
Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Gong Zhang 0002, Weiqiang Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2025 PreDAC: An Efficient Framework of Pre-Refining Enhanced Design Space Exploration for Approximate Computing
abstract
Approximate computing has emerged as a promising solution in energy-efficiency applications. Recently, attention has shifted from approximate components to Design Space Exploration (DSE) algorithms. However, traditional DSE algorithms face challenges in efficiently obtaining optimal solutions within large and complex design spaces. This paper introduces a prerefining enhanced design space exploration framework that provides customized design space and cost-performance formula for applications. Experimental results demonstrate that integrating this pre-refining step into various DSE algorithms leads to substantial performance gains, including up to $87 \times$ speedup and a 23% improvement in hardware overhead. Moreover, the innovative cost-performance-based DSE algorithm attains a $7.7 \times$ acceleration and further optimizes hardware metrics by an additional 8.8% compared to advanced frameworks employing the same pre-refinement.
Ziying Cui, Ke Chen 0018, Bi Wu 0002, Yu Gong 0002, Chenggang Yan 0002, Weiqiang Liu 0001
DAC5
2025 SS-MRAM: A Segment-Based Search Scheme With Configurable Matching for High-Accuracy Hyperdimensional Computing in CAM Applications
abstract
With the rise of hyperdimensional computing (HDC), content-addressable memories (CAMs) have emerged as an ideal choice for processing high-dimensional data. However, despite the advantages of high parallelism and low latency offered by CAM technology, it fails to address the significant loss of inference accuracy caused by closely matching hamming distances (HD). Emerging analog-based imprecise in-memory computing technologies frequently provide a minimum detectable HD that is insufficient for meeting the requirements of high similarity tasks. This limitation provides opportunities for using digital methods to realize fully exact matching memory computing technology based on magnetic random access memory (MRAM). In this work, a segment search scheme based on STT-MRAM devices and address-matching technology is proposed, which achieves zero loss in inference accuracy. An adaptive amplification structure is initially implemented by integrating the discharge method of latch structures, accompanied by the design of a 14T-2MTJ cell circuit. Through the optimization of the matching step, HD are calculated for each segment of configurable-dimensional vectors, ultimately facilitating the classification ranking of query hypervectors based on a comprehensive array architecture. Experimental results indicate that there is zero loss in inference accuracy when each segment of the dimension is configured to be below 16 bits. At 16 bits, the search power consumption in the worst-case matching scenario is measured at 1.73 fJ/bit, while the loss in inference accuracy does not exceed 0.4%.
Bi Wu 0002, Shuo Ran, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2024 A Time Efficient Comprehensive Model of Approximate Multipliers for Design Space Exploration
abstract
Multipliers play an essential role in various data processing applications and have garnered significant attention in approximate computing (AxC) for their energy-efficient features. However, formulating a precise error model for approximate data processing algorithms in conjunction with hardware metrics presents a challenge, leading to substantial time consumption in the design space exploration. This paper introduces an analytical model for approximate multipliers while considering input patterns. This model furnishes accurate error metrics, along with high-precision hardware metrics for various approximate multiplier configurations, impervious to variations in input data distribution. The proposed error model reduces the runtime by an average factor of 120.85 and, in some instances, by as much as 2,500 times, when contrasted with simulation-based methods. The design space exploration is performed on a 3×3 convolution circuit, revealing a comparable Pareto-optimal set and substantial reductions of up to 79.46% in the Power-Delay-Product (PDP) and 71.98% in area compared to the accurate counterpart. Additionally, the result of the Gaussian Blur application experiment demonstrates a 68.59% reduction in PDP and a 56.21% reduction in area, all while maintaining a PSNR of 30 dB.
Ziying Cui, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Yu Gong 0002, Weiqiang Liu 0001
ARITH4
2024 A Potential Enabler for High-Performance In-Memory Multi-Bit Arithmetic Schemes With Unipolar Switching SOT-MRAM
abstract
Due to the physical separation of data processing and storage, the conventional Von Neumann architecture exists excessive data migration overhead to curtail the progress of data-intensive applications. In this way, the Computing-in-Memory (CiM) architecture is proposed. Due to the boolean property of the memory cell, the current CiM mainly focuses on single-bit logic design. For the multi-bit arithmetic design, a prevalent patchwork approach is employed using single-bit logic, leaving the design with insufficient parallelism. This paper proposes a high-performance in-memory multi-bit addition (M-Add) and multiplication (M-Mul) scheme based on unipolar switching SOT-MRAM. For the M-Add scheme, transmission logic-based circuit design is proposed to realize single-step inter-column XOR operations, which is logically fits perfectly the g operator of parallel prefix algorithm. Further, the oBK algorithm is presented to maximize the g operator occupancy. For the M-Mul scheme, mapping the Booth decoder to the control signal of the proposed modified flip-flop queue, only two steps are required to realize the decoding of three encoded signals in parallel. The simulation results indicate the proposed design reduces the latency of N-bit Add (N-bit Mul) by an average of 82.6% (31.5%) compared to state-of-the-art CiM designs. Further, a CNN application based on proposed operations achieves 1.23 TOPS/w on the CIFAR-10 dataset, with an average of 47.57% increase over other CiM designs.
Bi Wu 0002, Tianyang Yu, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 Toward Efficient Retraining: A Large-Scale Approximate Neural Network Framework With Cross-Layer Optimization
abstract
Leveraging approximate multipliers in approximate neural networks (ApproxNNs) can effectively reduce hardware area and power consumption, making them suitable for edge-side applications. However, the propagation of layer-by-layer errors limits the application of approximate multipliers to large-scale ApproxNNs and complex tasks. Currently, retraining techniques that consider approximate multiplication errors are commonly used to compensate for the accuracy loss. However, due to the irregularity of the errors introduced by approximate multiplier, it is difficult for the existing generic acceleration hardware (e.g., GPU) to efficiently simulate its function and accelerate retraining, which thereby leads to a huge retraining overhead in ApproxNNs’ application. In this article, we propose an ApproxNN framework that introduces errors with regular and controlled positions for high-efficiency retraining of large-scale ApproxNNs. An approximate multiplier design that matches this framework is also presented to verify the effectiveness of the proposed ApproxNN framework. Experiment results demonstrate that the proposed ApproxNN framework is able to achieve up to 46$\times$speedup in retraining, and the proposed approximate multiplier reduces area/power-delay product (PDP) by 31%/63% compared to the exact multiplier. Compared with the floating-point neural network (NN) model, an accuracy decrease of only 1.13% is achieved when applied to ResNet50 on ImageNet dataset with only 15-epochs retraining, which surpasses other state-of-the-art designs.
Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2023 High Performance and Hardware-Efficient Approximate BPF Decoder for Polar codes
abstract
Belief propagation flip (BPF) decoding is a modified algorithm of BP decoding for polar codes, which has the error correction capability comparable to successive cancellation list (SCL) decoding while retaining the high throughput performance of BP decoding. However, the high complexity of BPF decoding algorithm limits its efficiency and maximum working frequency in hardware implementation. In this paper, a comprehensive BPF (CBPF) scheme is proposed by considering multi-factors affecting the selection of bits to be flipped. Additionally, two type processing elements (PEs) units in the decoders are proposed to reduce the latency. To further enhance the maximum working frequency and reduce the hardware efficiency, a parallel log-likelihood ratio (LLR) sorter using approximate computation is proposed. The proposed CBPF decoder with 1024 code length and 1/2 code rate is implemented on a 28nm CMOS technology, which achieves throughput of 20.48Gb/s at$\text{Eb}/\mathrm{N}0=4.0\ \text{dB}$with area occupied only$0.537mm^{2}$. Simulation results show that the proposed decoder has higher hardware efficiency and fairly good error correction performance compared to the state-of-the-art works.
Yuxuan Cui, Chenggang Yan 0002, Weiqiang Liu 0001
ISCAS2
2023 MLiM: High-Performance Magnetic Logic in-Memory Scheme With Unipolar Switching SOT-MRAM
abstract
Conventional computing architectures based on the von Neumann structure are suffering from the severe ‘memory wall’ issue due to the isolation and speed mismatch between memory and processor. As a promising solution, the concept of logic in-memory (LiM) has been proposed to effectively reduce the overhead of data migration and has been extensively studied in various memory technologies such as SRAM, DRAM, MRAM, ReRAM, etc. Among them, SOT-MRAM combines the advantages of non-volatility, low static power consumption, ultra-fast read/write speed, and high density, has emerged as one of the most promising candidates for low-power LiM implementations. In this paper, four in-memory logic operations, AND, OR, MAJ and full-addition (FA), are proposed based on the Unipolar Switching (US) SOT-MRAM devices. Incorporating the emerging switching behavior of SOT-MRAM, these operations can be performed with the basic memory access operations (read/write) with negligible modifying peripheral circuits. Meanwhile, by optimizing the operation steps, the performance degradation caused by the instability of SOT-MRAM device can be minimized in the proposed LiM architecture. Detailed simulation results show that the proposed design can reduce the latency (energy) of AND, OR operations at least by 71.2%, 74.4% (30.0%, 35.4%) compared with the existing SRAM and STT-MRAM designs. For MAJ and FA operations, the performance is improved by at least 34.7% and 44.8% compared to the existing design. The robustness of our design is demonstrated by the 100% pass of the 1000 samples Monte Carlo simulations for the sufficient switching current margin and the effectiveness of basic operations.
Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2023 Design of High Hardware Efficiency Approximate Floating-Point FFT Processor
abstract
The Fast Fourier Transformation (FFT), as a high-efficiency algorithm of the Discrete Fourier Transform (DFT), is widely used in Digital Signal Processing (DSP), wireless communication systems, spectrum analysis, and image processing. Approximate computing has shown effectiveness and feasibility to enhance the hardware efficiency of these applications. However, most approximate units in previous works are designed case by case, which has low efficiency and is difficult to find the optimal design. In this paper, a top-down design strategy for approximate floating-point (FP) FFT is proposed, which includes a mantissa bit-width adjustment algorithm and a step-by-step multiplier approximation algorithm. With the mantissa bit-width adjustment algorithm, the approximate 64 FP FFT achieved 50% area reduction and 70% power-delay product (PDP) reduction compared to the exact design with a 60dB Signal Noise Ratio (SNR) requirement, which is also at least 52% and 33% better than the previous approximate FP FFT. After using the step-by-step multiplier approximation algorithm, the approximate mantissa multiplier with an 8-bit fractional part reduced the area and PDP by 81.15% and 93.70%, respectively. The feasibility of the proposed approximate FFT design is verified in the channel estimation module of a wireless communication system, spectrum analysis, and image processing system.
Chenggang Yan 0002, Jipeng Ge, Chenghua Wang, Weiqiang Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 An Energy-efficient and High-precision Approximate MAC with Distributed Arithmetic Circuits
abstract
In this paper, an approximate distributed arithmetic (DA) based parallel MAC is proposed. First, by adopting three kinds of approximation methods, the novel structure significantly reduces hardware complexity. Then, the result is compensated according to the analysis of the probability to enhance the precision. The hardware and error metric evaluation demonstrates that the proposed MAC achieves 25% power-delay product reduction while maintaining better precision. Finally, the Gaussian Blur application is employed to verify the proposed DA-based MAC with 6dB average PSNR improvement compared with recent state-of-the-art work.
Ziying Cui, Ke Chen 0018, Bi Wu 0002, Chenggang Yan 0002, Weiqiang Liu 0001
ACM Great Lakes Symposium on VLSI4
2022 Data Stream Oriented Fine-grained Sparse CNN Accelerator with Efficient Unstructured Pruning Strategy
abstract
Network pruning can effectively alleviate the excessive parameters and computation issues in CNNs. However, unstructured pruning is not hardware friendly, while structured pruning will result in a significant loss of accuracy. In this paper, an unstructured fine-grained pruning strategy is proposed and achieves a 16X compression ratio with a top-1 accuracy loss of 1.4% for VGG-16. Combined with the proposed hardware-oriented hyperparameter selection method, compression rates of up to 64X can be obtained while fully meeting the edge-side accuracy requirements. Further, a light-weight, high-performance sparse CNN accelerator with modified systolic array is proposed for pruned VGG-16. The experimental results show that compared with the most advanced design, the proposed accelerator can achieve 21 Frames Per Second (FPS) with 3X better power efficiency and 2.19X better calculation density.
Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001
ACM Great Lakes Symposium on VLSI4
2022 Energy-efficient Oriented Approximate Quantization Scheme for Fine-Grained Sparse Neural Network Acceleration
abstract
For edge-side applications with severe power constraints, using pruning and quantization to compress models while maintaining network accuracy has become a widely deployed form of Convolutional Neural Networks (CNNs). For the same model accuracy, fine-grained non-regular pruning can bring higher model compression rate than coarse-grained regular pruning, but also introduces a larger indexing overhead. Besides, this overhead increases dramatically as the pruning granularity decreases. In this work, an approximate quantization scheme for fine-grained pruning is proposed. By reusing part of the quantized data bits, the proposed scheme can merge quantized data and indexes approximately, reducing the indexing overhead as well. Meanwhile, since the approximate compensation of index bits, the proposed scheme achieves an effective model accuracy improvement compared to the case of direct quantization to low bit-width. Experimental results show that, for 2:4 fine-grained pruning and 8-bit quantization scenario, the proposed method can save 20% of memory space and transmission cost. Compared with the direct quantization to 6-bit approach, the proposed scheme improves the accuracy by nearly 0.5% in the simulation of ImageNet dataset on ResNet50, despite occupying the same storage space. When deploying Yolov2-tiny at 16 × compression ratio, the energy efficiency of the CNN accelerator with the proposed approximate quantization is 1.33-3.82× that of other state-of-the-art designs.
Tianyang Yu, Bi Wu 0002, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001
ICCD4
2022 GBC: An Energy-Efficient LSTM Accelerator With Gating Units Level Balanced Compression Strategy
abstract
Recurrent Neural Networks (RNNs) have emerged as one of the most popular neural networks for processing time-series problems, widely used in machine translation, automatic speech recognition, and other natural language processing applications. However, conventional RNNs suffered from vanishing and exploding gradients, resulting in poor network performance in applications with long-term input information. As a variant of RNN, Long Short-Term Memory (LSTM) had been proposed to tackle this issue. Nevertheless, at the same time, LSTM introduces gating units and many additional parameters, which makes it challenging to be implemented directly on resource-limited platforms, such as Field Programmable Gate Arrays (FPGAs). This work first investigated the overall maximum achievable compression rates of different gating units and their correlations. Then, Gating Units Level Balanced Compression (GBC) strategy is proposed. After Top-$k$pruning, the proposed GBC strategy can attain a compression rate of$36.6\times $for LSTM. Further, the theoretical analysis indicates that for the existing gating units level LSTM compression variants, the GBC strategy still has further potential for compression. A complementary compression of the GBC strategy is performed on the existing coupled-gate LSTM to verify the analysis. Experimental results show that GBC achieves an additional$32\times $(overall$42.7\times $) compression rate with negligible accuracy loss. Finally, hardware experiments conducted on Xilinx ADM-PCIE-7V3 FPGAs also demonstrate that the accelerator designed in this paper achieves an improvement of 7.4%-191.5% in energy efficiency compared to the state-of-the-art designs.
Bi Wu 0002, Zhengkuan Wang, Ke Chen 0018, Chenggang Yan 0002, Weiqiang Liu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 A Dynamic Highly Reliable SRAM-Based PUF Retaining Memory Function
abstract
In this paper, a highly reliable SRAM based Physical Unclonable Function (PUF), which retains the memory function is proposed. The mismatch of NMOS is extracted during discharge process and amplified by the cross-coupled inverter to generate a response. At the beginning of the discharge process, the NMOSs are biased at sub-threshold region, which can improve the reliability and stability. The proposed PUF is designed in a 40nm CMOS process and each bit cell only consumes 4.98 μm2(3112F2). Post simulation shows that the bit error rate (BER) deterioration is 0.96% per 0.1V, 0.36% per 10° C with temperature variations from -40° C to 80° C and supply voltage variations from 0.9V to 1.3V. It achieves 1.8% native instability through the simulation. Meanwhile, the proposed PUF can retain memory function after a response is generated.
Chenghua Wang, Chenggang Yan 0002, Yijun Cui, Chongyan Gu, Máire O'Neill, Weiqiang Liu 0001
ISCAS3
2021 An Efficient High SFDR PDDS Using High-Pass-Shaped Phase Dithering
abstract
This brief proposes a high spurious-free dynamic range (SFDR) pulse output direct digital frequency synthesizer (PDDS) with low complexity and low power consumption. Independent and uniformly distributed (IUD) high-pass-shaped dither is added to the phase accumulator output, resulting in a wideband spurious-free and low close-in noise floor. The power efficiency and speed are increased by reducing the sampling frequency of the${m}$-sequence generator and the high-pass filter (HPF). The SFDR improves by 29 dB as a result of the HPF, which is confirmed with a field-programmable gate array (FPGA) implementation. The application specified integrated circuit (ASIC) occupies$1168~{\mu \text {m}^{2}}$on nangate 45-nm CMOS process and consumes 80.2 and$398~{\mu }\text{W}$from a 1.2-V supply with the dither generator running at${({1}/{4})f_{\text {clk}}}$and${f_{\text {clk}}}$(${f_{\text {clk}}}=2$GHz), respectively.
Chenggang Yan 0002, Jie Sun 0020, Weiqiang Liu 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2020 Active Noise Shaping SAR ADC Based on ISDM with the 5MHz Bandwidth
abstract
This paper reports a hybrid noise shaping (NS) successive-approximation registers (SAR) ADC based on incremental sigma-delta modulator (ISDM). The combination of ISDM and NS-SAR ADC achieves the high signal-to-noise ratio (SNR) with a few extra timing and hardware penalty. The reuse of integrator for both ISDM and NS releases the demands of high-resolution multi-input comparator replaced by only a 2-input one, which also helps suppress the noise from comparator and clock to the top plate of capacitor digital-to-analog converter (CDAC) relatively. With the embedded ISDM, the gain of amplifier used in the finite impulse response (FIR) filter for NS is relaxed remarkably, and wider bandwidth (BW) as well as lower oversampling rate (OSR) is realized compared to those traditional NS-SAR ADCs. The proposed hybrid ADC is designed in 40nm CMOS process with a 1.1V supply. It achieves the Signal-to-Noise Distortion Ratio (SNDR) and the Spurious Free Dynamic Range (SFDR) of 82.4dB and 97.1dBc respectively at 80MS/s, with a signal bandwidth (BW) of 5MHz and a total power consumption of 883uW.
Xin Li 0099, Chenggang Yan 0002, Jianhui Wu 0001
ISCAS3