Li Li 0003

dblp:53/2189-3 · DBLP profile ↗
← Back
53ranked-venue papers
1as first author
32since 2021 · last 2026
0000-0002-1047-6067ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 53 · 1 first-author · 32 since 2021
YearPublicationVenuePosition
2026 Lightweight Multi-View EEG Seizure Detection with a Two-Stage Low-Power BFP FFT-NN Accelerator
Chi Ben, Youbin Luo, Zhenglin Gu, Kexin Tian, Heng Zhang 0025, Li Li 0003
ISCAS7
2026 Perception-Core: Reconfigurable Energy-Efficient Domain-Specific Architecture with Multimodal Fusion for AIoT
Xinyu Wang 0027, Yuanhua Deng, Haixiang Ren, Hui Chen 0015, Li Li 0003
ISCAS8
2026 FRRAP: A Fast-Response Reconfigurable AI Processor and Its Application in Four-Class Seizure Monitoring
Heng Zhang 0025, Qiang Tao, Chi Ben, Youbin Luo, Guoqiang He, Sirui Zhu, Jingyi Ma, Li Li 0003
ISCAS9
2026 Thermal-Aware 3D-IC Floorplan Based On TSV-Coordination Simulated Annealing
Xingjie Zou, Linfeng Wu, Heng Zhang 0025, Xinyu Wang 0027, Li Li 0003
ISCAS7
2026 An Adaptive Congestion-aware Approximate Communication (ACAC) scheme and implementation for Network-on-Chip systems
Shize Zhou, Wenjie Fan 0004, Yongqi Xue, Shiping Li, Songfeng Deng, Jinlun Ji, Tong Cheng, Xinyu Wang 0027, Li Li 0003
Integr.10
2026 A 1.1 μJ/Inference Binary Spiking Neural Network Accelerator for DVS Gesture Recognition
abstract
Dynamic vision sensors (DVS) are bioinspired sensors that can generate sparse data streams with low latency, low power consumption and high dynamic range. Spiking neural networks (SNNs), which are inspired by biological brains, are event-based models, and therefore they can be used to process the binary data streams produced by such sensors naturally. However, SNN accelerators usually require more memory and longer time for inference, due to the extra time dimension in SNNs. In this paper, an energy efficient binary spiking neural network (BSNN) accelerator for DVS gesture recognition is proposed with algorithm and hardware codesign. We integrate the binary neural network (BNN) training method into SNN to directly train a BSNN model, which significantly reduces memory consumption. A temporal pooling (TP) layer is further proposed to reduce the time steps in SNNs while maintaining competitive accuracy. The proposed BSNN accelerator can achieve high parallelism with high resource utilization, and the sparsity of input spikes is utilized to further reduce power consumption. The proposed BSNN model achieves an accuracy of 95.49% on IBM DVS Gesture dataset. The implementation results show that the BSNN accelerator can achieve 38.2k inference per second with$1.1~\mu $J/inference energy consumption and 216.9 TOPS/W energy efficiency.
Congyi Sun, Xusen Zeng, Qiang Tao, Heng Zhang 0025, Qinyu Chen, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 Fast Modular Reduction Algorithm and Reconfigurable Domain-Specific Architecture Design Based on Generalized Mersenne Primes
abstract
Modular arithmetic enjoys a broad spectrum of applications. In recent years, the advancement of post-quantum cryptography (PQC) has imposed growing demands on the flexibility and scalability of domain-specific accelerators. This article presents a fast modular reduction algorithm based on generalized Mersenne (GM) primes, which employs approximate scaling and iterative compression approaches to rapidly converge the quotient value while reducing the precomputation complexity to rely solely on the modulus itself. Building upon this algorithm, we have designed a reconfigurable modular reduction array using multiplier units with smaller word length. Operating at 1GHz, the proposed array achieves an area reduction of 45.14% compared to Barrett-based structures, and 13.93% compared to Montgomery-based structures. The array has been integrated into a complete number-theoretic transform (NTT) acceleration architecture. The resulting reconfigurable GM/general modular reduction domain-specific architecture elevates the chip frequency to$0.42\sim 1$GHz under the same process technology. Under identical test conditions, it improves area efficiency by 10.6% and reduces energy consumption by 17.0% compared to the state-of-the-art ASIC design. When compared to the latest field-programmable gate array (FPGA) implementations, it achieves a reduction in area-time product (ATP) by 7.2%~18.4% for Kyber and by 13.4% for Dilithium. These results strongly demonstrate the notable advantages of the hardware-friendly GM algorithm.
Xinyu Wang 0027, Guoqiang He, Congyi Sun, Kai Chen 0034, Li Li 0003
IEEE Trans. Very Large Scale Integr. Syst.8
2026 QSNNA: An Energy-Efficient Quaternary Spiking Neural Network Accelerator for Seizure Detection
Heng Zhang 0025, Linfeng Wu, Linxiang Wang, Youbin Luo, Haochuan Pan, Xinyu Wang 0027, Guoqiang He, Qinyu Chen, Li Li 0003
IEEE Trans. Very Large Scale Integr. Syst.10
2025 Compact Interleaved Thermal Control for Improving Throughput and Reliability of Networks-on-Chip
abstract
Due to the scaling of sub-micron technology and the growing complexity of applications, escalating power density and traffic workloads heavily burden the network-on-chip (NoC) in multi-core systems and exacerbate thermal reliability issues. While recent thermal management techniques offer innovative solutions, they often employ the same management strategy for all tiles in NoC and activate it synchronously, which inevitably causes system oscillation and temperature cycling. In this paper, we propose a novel compact interleaved thermal control method that staggers the control phases of neighboring nodes to create negative feedback for each tile. We further explore the optimal control phase assignment by formulating it as a graph coloring problem to achieve the best performance. Experimental results demonstrate that the proposed method lowers the maximal spatial and temporal temperature variations up to 83.1% and 71.2% and improves the system throughput by 35.85% averagely compared with the state-of-the-art work. Besides, the proposed method can achieve a significant average enhancement of 317.50% and 234.83% in the minimal and average thermal-related mean time to failure (MTTF). Moreover, the method is scalable without extra power or area costs and is compatible with existing thermal management techniques.
Tong Cheng, Li Li 0003
ASP-DAC4
2025 FPGA-Par: An Efficient Algorithm for Elegant Partitioning in Multi-FPGA Systems
abstract
This work introduces FPGA-Par, an efficient graph partitioning algorithm for multi-FPGA systems. FPGA-Par utilizes an iterative balanced partitioning and supernodes transferring algorithm that transforms imbalanced partitioning problems into balanced ones, addressing the inefficiencies associated with traditional single-node movement methods, as well as eliminating the need for brute-force exploration of imbalance factors. This approach achieves partitions with flexible size ratios that satisfy resource constraints, resulting in improved partition quality. Experimental results demonstrate that FPGA-Par reduces time overhead by 55% compared to state-of-the-art imbalanced partitioning algorithms. Furthermore, it improves partition quality metrics, with Cut and Cutmaxdecreasing by 38% and 25%, respectively, and achieves a 1.25x increase in system frequency after mapping and routing on a multi-FPGA system.
Hengyue Gao, Chenyang Dai, Jinlun Ji, Qiyue Zhao, Jiangtao Yuan, Li Li 0003
ISCAS8
2025 NAME: NoC-based Accelerators Mapping Exploration for High Performance DNN Inference
abstract
The rapid advancement of deep learning, with increasingly large deep neural networks (DNNs), has led to the use of multi-core parallel processing in accelerators, utilizing Network-on-Chip (NoC) for interconnection. However, while multi-core architectures improve computational performance, they also increase data movement overhead. This paper analyzes NoC traffic patterns during DNN processing and explores mapping optimization to reduce computational and communicational overheads. We propose NoC-based Accelerators Mapping Exploration (NAME), an automated mapper for generating high-performance task mappings for NoC-based accelerators. NAME balances computational and communication overheads by dividing DNNs into scalable groups and interleaving NoC link usage, reducing traffic bottlenecks. Experiments show that NAME reduces execution time by up to 82% and 52%, respectively, compared to fixed mapping and the state-of-the-art framework AOME (Autonomous Optimal Mapping Exploration).
Jinlun Ji, Hengyue Gao, Heng Zhang 0025, Yuqi Lu, Yulong Song, Wenjie Fan 0004, Li Li 0003
ISCAS8
2025 HengNet: An Ultra-lightweight Model with Two-level Reuse Algorithm for Seizure Detection and Prediction
abstract
Traditional models based on electroencephalographic (EEG) signals for seizure monitoring encounter difficulties in simultaneously optimizing accuracy, response latency, and computational load. These challenges hinder their deployment in edge computing environments, where real-time local inference is critical. To address these issues, we introduce a novel network architecture, designated as HengNet. This architecture integrates a Two-level Reuse Algorithm (TRA), which strategically reutilizes outputs from intermediate layers, considerably reducing the average computational load per inference—vital for scenarios requiring frequent inferences. When tested on the CHB-MIT dataset, this patient-specific model attains classification accuracies of 95.67% and 99.60% for seizure prediction and detection, respectively. Notably, it maintains an average computational load of merely 0.05 million multiply-accumulate operations (MACs) per inference and has a compact model size of 6.87 K parameters. These results represent a significant advancement compared with existing methods. Operating at a rate of 32 inferences per second, the computational load of the model for seizure prediction has been reduced by more than 19.4 times, and for seizure detection, by more than 6.4 times.
Heng Zhang 0025, Linxiang Wang, Wenjie Fan 0004, Zhenglin Gu, Youbin Luo, Xingjie Zou, Chang Gao 0002, Qinyu Chen, Li Li 0003
ISCAS10
2025 An Energy Efficient Residual Spiking Neural Network Accelerator With Ternary Spikes
abstract
Spiking neural networks (SNNs) use discrete binary spikes to transfer information between neurons, which is different from artificial neural networks (ANNs). Although event-based characteristics bring potential computation power and efficiency to SNNs, the long processing time window of discrete spikes leads to high latency. In this brief, a spike version of the residual network using ternary spikes is proposed. A shorter time window is required to achieve competitive performance because the ability to transfer information of the ternary spikes is strengthened. An SNN accelerator based on the proposed residual network with ternary spikes is designed and implemented with 28 nm CMOS technology, and the core area is 0.63 mm2. The proposed SNN accelerator achieves the classification accuracy of 92.07% on CIFAR-10 dataset with SResNet20 and only 6 time steps. The accelerator achieves 0.39 mJ energy consumption per frame with a throughput of 165.7 FPS when running at 500 MHz.
Congyi Sun, Wenqing Song, Qinyu Chen, Chenyang Dai, Li Li 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 FAS-NoC: A Real-Time Fused Approximation Scheme Coordinating Communication and Computation for NoC-Based NN Accelerators
abstract
Network-on-Chip (NoC) is a scalable on-chip communication architecture widely used in neural network accelerators. However, data-intensive applications like machine learning place significant demands on the NoC’s communication and computation, and often have a degree of resilience to data noise, which allows to use approximation techniques to reduce execution time and energy consumption for both computation and communication, under the constraints of acceptable quality loss. Traditional approximate NoCs do not consider the data distribution characteristics of the neural networks, resulting in a lower approximate rate. Moreover, these schemes do not take the synergistic optimization of computation and communication, which limits reductions in execution time. In this paper, we propose a Fused Approximation Scheme of NoC (FAS-NoC) that incorporates the characteristics of data distribution in neural networks. FAS-NoC includes an approximate compression and recovery scheme based on data hierarchy, and uses a congestion-aware scheme to adjust the approximate rate of the node. Additionally, by leveraging the characteristics of recovered data after approximate communication, the scheme optimizes the design of computing units within the computing array. FAS-NoC collaboratively optimizes approximate communication and computation, organically integrating the two aspects. Compared with the state-of-the-art approximate framework ACDC (ACDC_ABDTR and ACDC_APPROX), the execution time of FAS-NoC is reduced by$48.94\%$and$47.72\%$, respectively. The experimental results show that the additional area overhead of the FAS-NoC only accounts for$0.81\%$of the original node, the additional power consumption overhead only accounts for$0.75\%$of the original node.
Chuanzhu Liu, Wenjie Fan 0004, Heng Zhang 0025, Chenyang Dai, Congyi Sun, Xinyu Wang 0027, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.8
2024 TTNNM: Thermal- and Traffic-Aware Neural Network Mapping on 3D-NoC-based Accelerator
abstract
3D Network on Chips (3D-NoCs) have ample on-chip wiring resources and high bandwidth, yet face numerous hotspots and higher temperature gradients due to increased integration and power density. This could lead to device failure, impacting system stability. Our paper introduces a thermal- and traffic-aware mapping method for 3D-NoC-based neural network accelerators. Firstly, based on the average load of different neural network layer, we determine their mapping sequences and suitable dies. Secondly, to minimize delay and alleviate hotspot temperatures, we allocate groups to appropriate nodes. Compared with previous works, TTNNM reduces the average temperature by 3.0°C, 2.2°C, 2.4°C, temperature variance by 58.4%, 64.8%, 73.0%, maximum temperature by 9.3°C, 7.9°C, 12.0°C, and packet latency by 31.7%, 17.2%, 25.1%.
Wenjie Fan 0004, Heng Zhang 0025, Jinlun Ji, Tong Cheng, Shiping Li, Li Li 0003
ACM Great Lakes Symposium on VLSI7
2024 Automatic Generation and Optimization Framework of NoC-Based Neural Network Accelerator Through Reinforcement Learning
abstract
Choices of dataflows, which are known as intra-core neural network (NN) computation loop nest scheduling and inter-core hardware mapping strategies, play a critical role in the performance and energy efficiency of NoC-based neural network accelerators. Confronted with an enormous dataflow exploration space, this paper proposes an automatic framework for generating and optimizing the full-layer-mappings based on two reinforcement learning algorithms including A2C and PPO. Combining soft and hard constraints, this work transforms the mapping configuration into a sequential decision problem and aims to explore the performance and energy efficient hardware mapping for NoC systems. We evaluate the performance of the proposed framework on 10 experimental neural networks. The results show that compared with the direct-X mapping, the direct-Y mapping, GA-base mapping, and NN-aware mapping, our optimization framework reduces the average execution time of 10 experimental NNs by 9.09$\%$, improves the throughput by 11.27$\%$, reduces the energy by 12.62$\%$, and reduces the time-energy-product (TEP) by 14.49$\%$. The results also show that the performance enhancement is related to the coefficient of variation of the neural network to be computed.
Yongqi Xue, Jinlun Ji, Xinming Yu, Shize Zhou, Tong Cheng, Shiping Li, Kai Chen 0034, Zhonghai Lu, Li Li 0003
IEEE Trans. Computers11
2024 HAS-RL: A Hierarchical Approximate Scheme Optimized With Reinforcement Learning for NoC-Based NN Accelerators
abstract
Network-on-Chip (NoC) is a scalable on-chip communication architecture for the NN accelerator, but with the increase in the number of nodes, the communication delay becomes higher. Applications such as machine learning have a certain resilience to noisy/erroneous transmitted data. Therefore, approximate communication becomes a promising solution to improving performance by reducing traffic loads under the constraint of the acceptable maximum accuracy loss of neural networks. It is a key issue to balance the result quality and the communication delay for approximate NoC systems. The traditional approximate NoC only considers the node-to-node approximation-based dynamic traffic regulation. However, the dynamically changing traffic patterns across different nodes, different times, and different applications lead to a huge search space, which makes it hard to explore an optimal global approximation solution. In this paper, we propose a quality model for different neural networks, which presents the relationship between the quality loss and the data approximate rate. Then, a hierarchical approximate scheme optimized with reinforcement learning (HAS-RL) is proposed and we reduce the complexity of the HAS-RL by reducing the state space and action space, which will reduce the resource overhead as well. After that, we embed a global approximate controller in the NoC system, in which we deploy a policy network trained with the offline reinforcement learning algorithm to adjust the data approximate rates of each node at run time. Compared with the state-of-the-art method, the proposed scheme reduces the average network delay by 13.5% while their accuracies are similar. The proposed HAS-RL only causes an additional area overhead of 1.24% and power consumption of 0.77% compared with the traditional router design.
Shize Zhou, Yongqi Xue, Wenjie Fan 0004, Tong Cheng, Jinlun Ji, Chenyang Dai, Wenqing Song, Qinyu Chen, Chang Gao 0002, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.11
2023 Low-Cost High-Precision Architecture for Arbitrary Floating-Point Nth Root Computation
abstract
In this paper, we propose a feasible architecture with high precision and low resource consumption to compute the$N$th root of a floating-point number, which is mainly based on radix-4 SRT and 2-based Coordinate Rotation Digital Computer (CORDIC). Simulation results show that our method can achieve a relative error of the magnitude of 10−7. Under the same precision requirements, the hardware implementation results show a better performance of our design in terms of area, power, and absolute delay compared with the method based on the generalized hyperbolic CORDIC. After synthesizing it under the TSMC$40n$m CMOS technology, it can be obtained that our design can achieve an area consumption of$125465.80\ \mu m^{2}$and power consumption of 97.8062 mW at the highest frequency of 3.12 GHz.
Wanyuan Hong, Hui Chen 0015, Lianghua Quan, Li Li 0003
ISCAS5
2023 A DSP-Purposed REconfigurable Acceleration Machine (DREAM) for High Energy Efficiency MIMO Signal Processing
abstract
The wireless baseband processing algorithms are still developing and show a great diversity. The development of ASIC implementations cannot quickly adapt to the evolution of algorithms and standards. Meanwhile, the general-purpose processors cannot meet the real-time requirements in some scenarios. This paper proposes a DSP-purposed REconfigurable Acceleration Machine (DREAM) core for wireless baseband digital signal processing, which has a good trade-off between flexibility and performance. First, we abstract a set of shared operators with a moderate granularity from a variety of wireless MIMO signal processing algorithms. Then, we propose a two-step configuration process to reduce the size of the required reconfiguration bits. Besides, we design a conflict-free address generator to transfer data between the on-chip scratchpad memory and reconfiguration processing elements with high efficiency and high throughput. Finally, the prototype DREAM core has been implemented in TSMC CMOS 28 nm, and its area and power consumption have been analyzed. The chip has great flexibility in supporting a variety of wireless MIMO processing algorithms and a wide range of MIMO scales. The proposed DREAM core can achieve the normalized area efficiency and the normalized energy efficiency of$0.67~Gbps/MGE$and$15.05~Gbps/W$, which are$1.56\times $and$4.18\times $those of state-of-the-art reconfigurable implementations when running the WeJi-based MIMO detection algorithm.
Kai Chen 0034, Wenqing Song, Guoqiang He, Sirui Shen, Huizheng Wang, Chuan Zhang 0001, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.8
2022 Work in Progress: ACAC: An Adaptive Congestion-aware Approximate Communication Mechanism for Network-on-Chip Systems
abstract
Data-intensive applications, such as machine learning and pattern recognition, result in heavy Network-on-Chip (NoC) communication loads and a tremendous increase in the communication latency. At the same time, the error-tolerant nature of these applications makes approximate communication an effective way to relieve the sharp increase of the network latency. This paper proposes an adaptive congestion-aware approximate communication mechanism (ACAC) that can alleviate the communication congestion of NoC systems in heavy communication loads. Our cycle-accurate simulations have shown that the proposed ACAC effectively reduces the network latency similar to ABDTR under a 22% to 52% lower data approximate ratio and significantly decreases the additional compression control traffic volume under real applications.
Shize Zhou, Yongqi Xue, Jinlun Ji, Tong Cheng, Li Li 0003
CASES6
2022 AOME: Autonomous Optimal Mapping Exploration Using Reinforcement Learning for NoC-based Accelerators Running Neural Networks
abstract
Hardware mapping plays a critical role in the performance of NoC-based accelerators running large-scale neural networks (NN). Confronted with enormous mapping exploration space, traditional algorithms may find sub-optimal solutions. We conduct preliminary experiments to investigate the impact of different hardware mappings on communication latencies. Then, this paper proposes an Autonomous Optimal Mapping Exploration (AOME) architecture based on two reinforcement learning algorithms. Combining soft and hard constraints, AOME transforms the mapping process into a sequential decision problem and targets to explore the optimal mapping of the NoC system. We evaluate the performance of AOME on ten NNs. The results show that compared with the direct X mapping, the direct Y mapping, GA-base mapping, and NN-aware mapping, AOME reduces the average communication latency of ten NNs by 27.30%, 33.33%, 4.27% and 12.46% using A2C, by 27.19%, 33.21%, 4.11% and 12.31% using PPO, and improves the average communication throughput by 43.24%, 63.60%, 5.17% and 14.83% using A2C, by 43.18%, 63.68%, 5.23% and 14.87% using PPO.
Yongqi Xue, Jinlun Ji, Shize Zhou, Tong Cheng, Li Li 0003
ICCD7
2022 A Hierarchical Parallel Discrete Gaussian Sampler for Lattice-Based Cryptography
abstract
Discrete Gaussian sampling is one of the important components in lattice-based cryptosystems which are promising candidates for post-quantum cryptographic algorithms. For sufficient security and satisfactory performance, the Knuth-Yao algorithm is an efficient way to implement discrete Gaussian samplers. Nevertheless, most polynomials in lattice-based cryptography have 256 coefficients or more, which suffers from long latency to complete the sample generation. In this paper, the first parallel discrete Gaussian sampler with hierarchical structure is proposed, while keeping statistical distance to the actual distribution. Based on the imbalanced visiting frequency of the probability matrix, a three-stage generation strategy is adopted with hierarchical bit search units (BSUs) that can greatly reduce area consumption of the repeated costly lookup tables. Besides the architecture improvement, a lowest-set-bit scanning scheme is introduced to BSUs. Moreover, the parallelism of our design provides obfuscation ability against side-channel attacks (SCAs). A practical hardware implementation of discrete Gaussian distributions with $\sigma$=3.33 on the Xilinx Virtex-5 XC5VLX30 FPGA device spends 26.12 ns on average to generate 256 samples, consuming 994 slices. Results have verified its advantages of area efficiency over the state-of-the-arts (SOAs).
Sirui Shen, Wenqing Song, Xinyu Wang 0027, Xinyu Shao, Zhonghai Lu, Li Li 0003
ISCAS7
2022 Unsupervised Learning Based on Temporal Coding Using STDP in Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have been recognized as one of the next generation of Neural Networks (NNs), showing a great potential in a variety of applications. Spiking-Timing Dependent Plasticity (STDP) underlies the brain’s learning mechanisms, and trains SNNs with great energy efficiency. In this paper, we propose a low-cost spike-time based unsupervised learning method. It constructs a SNN with one fully-connected excitatory layer structure without inhibitory layer, and trains the SNN with STDP using a first-spike-based temporal coding scheme where input information is directly encoded into spike times. It only updates the synaptic weights connected to the neuron that first generates a spike in a forward propagation step, which reduces the frequency of the synaptic weight updates significantly. The forward propagation process can be stopped once a neuron fires whether in the training mode or the inference mode, by which many unnecessary computations are just avoided and the latency in the inference mode is reduced. The method was used to train on the classification task on MNIST dataset and achieved an accuracy of 90.4% with 800 excitatory neurons.
Congyi Sun, Qinyu Chen, Kai Chen 0034, Guoqiang He, Li Li 0003
ISCAS6
2022 Huicore: A Generalized Hardware Accelerator for Complicated Functions
abstract
Emerging advanced System-on-Chip (SoC) designs contain more and more complicated functions to be accelerated. This presents a challenge to conventional design approaches which use different hardware architectures or separate hardware accelerators to implement the various functions. To tackle this challenge, for the first time, we propose a generalized hardware accelerator called “Huicore” to speed up diverse functions on the same substrate. Through the analysis and transformation of mathematical characteristics, we reveal the commonality of many complicated functions using the CORDIC algorithm. Then we explore a reconfigurable architecture to implement them. The proposed reconfigurable accelerator can not only accelerate the implementation of many complicated functions, but also has small area, low power consumption and high precision. It is very suitable for integration in a SoC system to accelerate the implementation of various applications.
Hui Chen 0015, Zongguang Yu, Zhonghai Lu, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.7
2022 An Energy Efficient STDP-Based SNN Architecture With On-Chip Learning
abstract
In this paper, we propose a spike-time based unsupervised learning method using spiking-timing dependent plasticity (STDP). A simplified linear STDP learning rule is proposed for the energy efficient weight updates. To reduce unnecessary computations for the input spike values, a stop mechanism of the forward pass is introduced in the forward pass. In addition, a hardware-friendly input quantization scheme is used to reduce the computational complexities in both the encoding phase and the forward pass. We construct a two-layer fully-connected spiking neuron network (SNN) based on the proposed method. Compared to general rate-based SNNs trained by STDP, the proposed method reduces the complexity of network architecture (an extra inhibitory layer is not needed) and the computations of synaptic weight updates. According to the fixed-point simulation with 9-bit synaptic weights, the proposed SNN with 6144 excitatory neurons achieves 96% of recognition accuracy on MNIST dataset without any supervision. An SNN processor that contains 384 excitatory neurons with on-chip learning capability is designed and implemented with 28 nm CMOS technology based on the proposed low complexity methods. The SNN processor achieves an accuracy of 93% on MNIST dataset. The implementation results show that the SNN processor achieves a throughput of 277.78k FPS with$0.50~\mu \text{J}$/inference energy consuming in inference mode, and a throughput of 211.77k FPS with$0.66~\mu \text{J}$/learning energy consuming in learning mode.
Congyi Sun, Haohan Sun, Jianing Han, Xinyuan Wang 0007, Xinyu Wang 0027, Qinyu Chen, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.9
2022 Low-Latency Low-Complexity Method and Architecture for Computing Arbitrary Nth Root of Complex Numbers
abstract
This paper presents a new architecture, based on CORDIC and parabolic synthesis methodology, for computing Nth root of a complex number. The proposed architecture uses the pretreatment for normalization and parabolic synthesis method to calculate the Nth root of modulus of the input complex number and performs the conversion between the plane coordinate form and the polar coordinate form of the complex number by CORDIC, which not only ensures the accuracy but also has an ultra-low computation latency. MATLAB simulation result indicates that our proposed method can calculate the Nth root of the complex numbers in the form of fixed-point number with an error of$2.16 \boldsymbol {\times {10^{ - 6}}}$. Under TSMC 40nm CMOS technology, the report shows that the area consumption is$27390.72 \boldsymbol {\mu m^{2}}$at the frequency of 1GHz and the power consumption is 2.3549mW. More importantly, the computation latency of the proposed architecture is only 60.18% of the latest architecture in the same calculation accuracy.
Hui Chen 0015, Guoqiang He, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.5
2021 A General Methodology and Architecture for Arbitrary Complex Number Nth Root Computation
abstract
As the existing complex number Nth root computation methods are relatively discrete, we propose a general method and architecture based on coordinate rotation digital computer (CORDIC) to compute arbitrary complex number Nth root for the first time. Our method performs the tasks of computing complex modulus, complex phase angle, real Nth root, sine function and cosine function, which can be implemented by circular CORDIC, linear CORDIC and hyperbolic CORDIC. Based on these CORDICs, our proposed architecture can not only improve the hardware efficiency just through shift-add operations, but also flexibly adjust the precision and the input range of complex number Nth root. To prove its feasibility, we conduct a software simulation and implement an example circuit in hardware. Under the TSMC 28nm CMOS technology, we synthesize it and get the report that it has the area of 6561μm2and the power of 3.95mW at the frequency of 1.5GHz.
Hui Chen 0015, Zhonghai Lu, Li Li 0003, Zongguang Yu
ISCAS5
2021 Dynamic and Traffic-Aware Medium Access Control Mechanisms for Wireless NoC Architectures
abstract
Wireless NoC (WiNoC) has low latency and simple wiring, which can reduce the energy consumption caused by the metal interconnection in traditional NoC architectures. However, traditional time division based media access control (MAC) mechanism in WiNoC is not aware of different wireless interfaces' (WIs) traffic demands, resulting in an unreasonable distribution of wireless communication channels and degradation in performance. Hence, in order to dynamically allocate wireless channels to the WIs based on their traffic demands, a dynamic and traffic-aware MAC mechanism is required. In this paper, we design a traffic demand predictor for each WI based on its current and history traffic conditions. According to the predicted demands, we are able to allocate access to wireless channels dynamically and switch between two kinds of time division based MAC mechanisms. Simulations under various conditions indicate that the average delay decreases by 30% and 20% on average compared with a traditional MAC mechanism and an existing dynamic time division based one, respectively. Moreover, the network with the dynamic and traffic-aware MAC enters the saturation point at a higher packet injection rate.
Wenqing Song, Zhonghai Lu, Li Li 0003
ISCAS4
2021 Adaptive Successive Cancellation Priority Decoder for 5G Polar Codes
abstract
As two common successive cancellation (SC)-based decoding algorithms of polar codes, the SC list (SCL) and SC stack (SCS) decoder can achieve satisfactory error correction performance, especially with increased list size or stack depth. Nevertheless, a large list size or stack depth will lead to high computational complexities and hardware resources. To this end, successive cancellation priority (SCP) decoding with priority- first searching strategy and trellis-like storage is proposed to offer one solution. In this paper, an efficient SCP decoder is first proposed to verify its advantages over SCL and SCS decoders. Furthermore, an adaptive node-inserting scheme is proposed to reduce the number of bits insert into the priority queue. Numerical results have shown that for the polar code with transmission length 1024 and rate 1/2, the proposed adaptive SCP (ASCP) decoder can achieve significant time complexity reduction on average compared with the standard SCL decoder. The hardware architecture of SCP decoding is implemented using 65-nm CMOS technology and the results show better throughput compared with the SCS decoder.
Wenqing Song, Yifei Shen 0003, Chuan Zhang 0001, Li Li 0003
ISCAS5
2021 Optimizing Vertical Link Placement and Congestion Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCs
abstract
The fully connected 3D-NoCs in which all routers are vertically connected with their neighbors above and below need a lot of Through-Silicon-Vias (TSVs), and they will occupy a large silicon area and reduce the fabrication yield. Thus, the idea of partially connected 3D-NoCs has emerged. The optimal number and placement of the vertical links (elevators) must be determined at the chip design stage, which is a multiobjective optimization problem of the performance and the cost. However, optimizing the static elevator placement needs a great amount of calculation and we can not examine all possible solutions at design time. Therefore, we propose a hybrid heuristic strategy for the static elevator placement and assignment, in which the genetic algorithm and the tabu search are combined. The dynamic assignment method is essential for the partially connected 3D-NoCs, and it leads to different traffic distributions and therefore has a huge impact on performance. Many previous static assignment methods can not dynamically change the elevator assignment according to the real-time states of the network, thus it may lead to network congestion. A congestion-aware dynamic assignment (CDA) scheme is proposed in this article, which considers the impact of the distance factor and the congestion factor on the network performance. Experiments show that the proposed CDA method can improve the network performance by 67%-86% compared with the random selection algorithm and can improve the reliability of the partially connected 3D-NoC as well. The key component for the CDA method, the path selection module (PSM), is implemented in FPGA, and the results show that its area cost is negligible compared with a router.
Chuan Zhang 0001, Wenqing Song, Qinyu Chen, Hui Chen 0015, Li Li 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2021 Symmetric-Mapping LUT-Based Method and Architecture for Computing XY-Like Functions
abstract
We propose a new method and hardware architecture to compute the functions expressed as XY (X and Y are arbitrary floating-point numbers), which can support arbitrary Nth root, exponential and power operations. Because of the complexity of direct computation, we usually convert it to logarithm, multiplication, and antilogarithm operations. Traditional approaches suffer from long latency, large area and high power consumption. To solve this problem, we propose a symmetric-mapping lookup table (SM-LUT) to be capable of computing log2x (x ∈ [1, 2]) and 2x(x ∈ [0, 1]) simultaneously. It lays the foundation for computing XY. To further improve hardware performance of our architecture, we propose a multi-region address searcher to speed up the calculation of SM-LUT. In addition, we use an optimized Vedic multiplier to shorten the critical path and improve the efficiency of multiplication, which is included in computing XY. Under the TSMC 40nm CMOS technology, we design and synthesize a reference circuit to compute XY with a maximum relative error of 10-3. The report shows that the reference circuit achieves the area of 14338.50 μm2and the power consumption of 4.59 mW at the frequency of 1 GHz. In comparison with the state-of-the-art work under the same input range and similar precision, it saves 78.57% area and 80.42% power consumption for N√R computation and 82.89% area and 81.89% power consumption for RN computation averagely. On top of that, our architecture reduces the computation latency by 62.77% averagely and has one more order of magnitude of energy efficiency than others.
Hui Chen 0015, Heping Yang, Wenqing Song, Zhonghai Lu, Li Li 0003, Zongguang Yu
IEEE Trans. Circuits Syst. I Regul. Pap.6
2021 Low-Complexity High-Precision Method and Architecture for Computing the Logarithm of Complex Numbers
abstract
This paper proposes a low-complexity method and architecture to compute the logarithm of complex numbers based on coordinate rotation digital computer (CORDIC). Our method takes advantage of the vector mode of circular CORDIC and hyperbolic CORDIC, which only needs shift-add operations in its hardware implementation. Our architecture has lower design complexity and higher performance compared with conventional architectures. Through software simulation, we show that this method can achieve high precision for logarithm computation, reaching the relative error of 10-7. Finally, we design and implement an example circuit under TSMC 28nm CMOS technology. According to the synthesis report, our architecture has smaller area, lower power consumption, higher precision and wider operation range compared with the alternative architectures.
Hui Chen 0015, Zongguang Yu, Yonggang Zhang 0005, Zhonghai Lu, Li Li 0003
IEEE Trans. Circuits Syst. I Regul. Pap.6
2020 A CORDIC-Based Architecture with Adjustable Precision and Flexible Scalability to Implement Sigmoid and Tanh Functions
abstract
In the artificial neural networks, tanh (hyperbolic tangent) and sigmoid functions are widely used as activation functions. Past methods to compute them may have shortcomings such as low precision or inflexible architecture that is difficult to expand, so we propose a CORDIC-based architecture to implement sigmoid and tanh functions, which has adjustable precision and flexible scalability. It just needs shift-add-or-subtract operations to compute high-accuracy results and is easy to expand the input range through scaling the negative iterations of CORDIC without changing the original architecture. We adopt the control variable method to explore the accuracy distribution through software simulation. A specific case (ARCH. (1, 15, 18), RMSE: 10−6) is designed and synthesized under the TSMC 40nm CMOS technology, the report shows that it has the area of 36512.78μm2and power of 12.35mW at the frequency of 1GHz. The maximum work frequency can reach 1.5GHz, which is better than the state-of-the-art methods.
Hui Chen 0015, Yuanyong Luo, Zhonghai Lu, Li Li 0003, Zongguang Yu
ISCAS6
2020 An Efficient Accelerator for Multiple Convolutions From the Sparsity Perspective
abstract
Convolutional neural networks (CNNs) have emerged as one of the most popular ways applied in many fields. These networks deliver better performance when going deeper and larger. However, the complicated computation and huge storage impede hardware implementation. To address the problem, quantized networks are proposed. Besides, various convolutional structures are designed to meet the requirements of different applications. For example, compared with the traditional convolutions (CONVs) for image classification, CONVs for image generation are usually composed of traditional CONVs, dilated CONVs, and transposed CONVs, leading to a difficult hardware mapping problem. In this brief, we translate the difficult mapping problem into the sparsity problem and propose an efficient hardware architecture for sparse binary and ternary CNNs by exploiting the sparsity and low bit-width characteristics. To this end, we propose an ineffectual data removing (IDR) mechanism to remove both the regular and irregular sparsity based on dual-channel processing elements (PEs). Besides, a flexible layered load balance (LLB) mechanism is introduced to alleviate the load imbalance. The accelerator is implemented with 65-nm technology with a core size of 2.56 mm2. It can achieve 3.72-TOPS/W energy efficiency at 50.1 mW, which makes it a promising design for embedded devices.
Qinyu Chen, Wenqing Song, Zhonghai Lu, Li Li 0003
IEEE Trans. Very Large Scale Integr. Syst.7
2019 Smilodon: An Efficient Accelerator for Low Bit-Width CNNs with Task Partitioning
abstract
Convolutional Neural Networks (CNNs) have been widely applied in various fields such as image and video recognition, recommender systems, and natural language processing. However, the massive size and intensive computation loads prevent its feasible deployment in practice, especially on the embedded systems. As a highly competitive candidate, low bit-width CNNs are proposed to enable efficient implementation. In this paper, we propose Smilodon, a scalable, efficient accelerator for low bit-width CNNs based on a parallel streaming architecture, optimized with a task partitioning strategy. We also present the 3D systolic-like computing arrays fitting for convolutional layers. Our design is implemented on Zynq XC7Z020 FPGA, which can satisfy the needs of real-time with a frame rate of 1, 622 FPS throughput, while consuming 2 1 Watt. To the best of our knowledge, our accelerator is superior to the state-of-the-art works in the tradeoff among throughput, power efficiency, and area efficiency.
Qinyu Chen, Kaifeng Cheng, Wenqing Song, Zhonghai Lu, Li Li 0003, Chuan Zhang 0001
ISCAS6
2019 Congestion-Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCs
abstract
The combination of Network-on-Chips (NoCs) and 3D IC technology, 3D NoCs, has been proven to be able to achieve a great improvement in both network performance and power consumption compared to 2D NoCs. In the traditional 3D NoC, all routers are vertically connected. Due to the large overhead of Through-Silicon-Via (TSV, e.g., low fabrication yield and the occupied silicon area), the partially connected 3D NoC has emerged. The assignment method determines the traffic loads of the vertical links (elevators), thus has a great impact on 3D-NoCs' performance. In this paper, we propose a congestion-aware dynamic elevator assignment (CDA) scheme, which takes both the distance factors and network congestion information into account. Experiments show that the performance of the proposed CDA scheme is improved by 67% to 87% compared to the random selection scheme, 8% to 25% compared to SelByDis-1, and 13% to 18% compared to SelByDis-2.
Qinyu Chen, Guoqiang He, Kai Chen 0034, Zhonghai Lu, Chuan Zhang 0001, Li Li 0003
ISCAS7
2019 Thermal Sensor Placement and Thermal Reconstruction Under Gaussian and Non-Gaussian Sensor Noises for 3-D NoC
abstract
On-chip thermal sensors are essential for temperature management in 3-D network-on-chip (NoC) systems. However, due to the physical (area and power) or economical constraints, the number of sensors is limited. Therefore, the two critical issues we face are: 1) how to figure out an efficient thermal sensor placement with the limited number of sensors and 2) how to reconstruct the entire thermal profile based on sensor observations. Another major issue for the thermal reconstruction is the sensor measurement accuracy. Thus, online accurate full-chip thermal reconstruction under Gaussian and non-Gaussian noises is another great challenge. In this paper, a greedy thermal sensor placement algorithm maximizing the rank of the observability Gramian is proposed. A good placement algorithm always relies on a specific reconstruction method. The proposed placement algorithm is designed for the state-space-based thermal model, thus the combination of the proposed placement algorithm and the Kalman filter-based reconstruction method provides a high reconstruction accuracy under Gaussian noise. For accurate temperature reconstruction under non-Gaussian noise, the Gaussian-Sum filter is applied to 3-D NoC. Compared with the Kalman filter, the Gaussian-Sum filter can reduce the root-mean-squared-error and the max error by 29.27%–35% and 33.26%–40.6%, respectively. A reusable architecture for the Kalman filter and the Gaussian-Sum filter has been proposed. Its hardware implementation details are presented in this paper. Besides, the performance and the area are evaluated as well.
Li Li 0003, Hongbing Pan, Kun Wang 0005, Qinyu Chen, Chuan Zhang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Efficient Successive Cancellation Stack Decoder for Polar Codes
abstract
As an improved version of successive cancellation (SC) polar decoder, an SC stack (SCS) decoder has been proposed for performance improvement. However, the existing SCS polar decoder suffers a lot from high time complexity at low signal-to-noise ratio (SNR) region and space complexity compared with the SC decoder. To this end, two improved decoders are proposed to reduce time and space complexity in both low and high SNR regions. The first one is the segmented cyclic redundancy check (CRC)-aided SCS (SCA-SCS) decoder, which is based on segmented parity checkers. The second one is the adaptive SCS (ASCS) decoder, which has the flexibility of stack depth and searching width. Furthermore, a channel condition estimator is proposed to select appropriate decision criteria for different SNR scenarios. Results have shown that for the polar code of length 1024 and rate 1/2, two improved SCS decoders can perform better than the traditional SCS decoder. The proposed SCA-SCS decoder and the ASCS decoder can achieve 10.8% and 11.42% time complexity reduction and 31.68% and 60.85% space complexity reduction on average over binary-input additive white Gaussian noise channels (BI-AWGNCs), respectively. Efficient parallel hardware architecture of the SCS polar decoder is first proposed and implemented with 90- and 65-nm technologies. Results have verified its advantages over the state of the art (SOA).
Wenqing Song, Huayi Zhou 0002, Kai Niu 0001, Zaichen Zhang, Li Li 0003, Xiaohu You 0001, Chuan Zhang 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2019 Analysis and Design of a Large Dither Injection Circuit for Improving Linearity in Pipelined ADCs
abstract
In this paper, a new large dither injection technique is proposed for improving linearity in pipelined analog-to-digital converters (ADCs), without losing the dynamic range of the ADCs and deteriorating the corresponding amplifier's linearity. First, analyses of a proper pipelined ADC's architecture are performed for large dither injection. Then, a 9-bit capacitive digital-toanalog converter (DAC) with split architecture is developed to inject the dither ranging from -511/1024 least significant bit (LSB) to 511/1024 LSB of the first stage. To counteract the consumption of the correction range by the capacitive injection dither, the novel 6-bit complementary DACs embedded in the comparator threshold generation circuit are proposed to realize comparator dither injection. In addition, the dither injection amplitude is configurable for investigating different amplitude's effects on the linearity of the ADC. Finally, the proposed dither injection circuit, together with a 16-bit 150 million samples per second (MSPS) ADC, is implemented in a 0.18-μm CMOS technology. The measured results demonstrate the effectiveness of the proposed techniques. The optimum dither is the 9-bit dither, improving not only the spurious free dynamic range (SFDR) of the small signal by at least 13 dB but also that of the large signal by more than 8 dB compared to the case without dither injection. Moreover, dither injection makes the noise floor clean.
Congyi Zhu, Renrong Liang, Jun Lin 0001, Zhongfeng Wang 0001, Li Li 0003
IEEE Trans. Very Large Scale Integr. Syst.5
2017 Kalman Predictor-Based Proactive Dynamic Thermal Management for 3-D NoC Systems With Noisy Thermal Sensors
abstract
Thermal sensor noise has a great impact on the efficiency and effectiveness of a dynamic thermal management (DTM) strategy. To address the problem of forecasting temperature based on noisy thermal sensors, we first propose a Kalman-based runtime thermal prediction scheme. To obtain accurate temperature predictions, a multivariate linear power model and a physically-based state space thermal model for 3-D network-on-chip are also proposed. Simulation results show that it reduces the standard deviations of the prediction error by 46%–53% compared with the auto-regressive based one under sensor noise with${\sigma =2}$. Conventional reactive DTM techniques suffer from significant performance degradation due to their pessimistic reaction, thus, based on the proposed prediction scheme, we further propose a proactive DTM strategy that primarily consists of a thermal-aware routing algorithm and a proactive throttling scheme: 1) to take into account both thermal and congestion issues, we propose a proactive congestion and thermal aware routing algorithm. Simulation results demonstrate that it can achieve better throughput as well as approach better thermal balance. Specifically, under uniform traffic, the proposed scheme reduces the maximum chip temperature by about 3.9 °C and achieves 78.3% higher throughput compared with the competing thermal optimization approach based on dynamic programming network and 2) when the temperature exceeds the threshold, existing coarse-grained reactive throttling schemes cool down the overheated nodes at the penalty of significant performance loss. In this paper, a proactive quota-based throttling scheme is proposed. Simulation results show that it improves the throughput up to 11.1% compared with the reactive throttling schemes.
Li Li 0003, Kun Wang 0005, Chuan Zhang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 Accurate runtime thermal prediction scheme for 3D NoC systems with noisy thermal sensors
abstract
Thermal sensor noise has great impact on the efficiency and effectiveness of a dynamic thermal management (DTM) strategy. Conventional reactive thermal management techniques suffer significant performance degradation due to the pessimistic reaction. In this paper, to address the problem of forecasting temperatures based on noisy thermal readings, we propose a Kalman predictor based runtime thermal prediction scheme, which can predict temperatures N step ahead. An activity-based power model for 3D NoC power estimation is also proposed; the model is an essential prerequisite of accurate temperature predictions. Besides that, we propose a distributed multi-input single-output (MISO) thermal model for 3D NoC systems, which reduces the computational complexity of temperature updating from m2 to m compared with the centralized multi-input multi-output (MIMO) model for the system with m units. The experimental results show that the proposed prediction scheme reduces the mean absolute error (MAE) by 42.8%-72.6% compared with the auto-regressive (AR) based prediction scheme.
Li Li 0003, Hongbing Pan, Kun Wang 0005, Feng Han 0008, Jun Lin 0001
ISCAS2
2016 A high throughput belief propagation decoder architecture for polar codes
abstract
The belief propagation (BP) decoding algorithm not only is an alternative to the successive cancelation (SC) decoders of polar codes, but also provides soft outputs that are necessary for joint detection and decoding. The BP decoders with the flooding schedule achieve high throughput with excessive hardware cost especially when the block length is large. The soft-cancelation (SCAN) decoders for polar codes have reduced memory complexity compared to the BP decoders based on the flooding schedule. The simplified SC aided reduced complexity soft-cancelation (S-RCSC) decoders further reduce the computational and memory complexity of the SCAN decoders at the cost of negligible error performance degradation. Both the SCAN and S-RCSC decoders have limited throughput due to their serial decoding schedules. In this paper, we first propose an improved S-RCSC (IS-RCSC) decoding algorithm and then present a high throughput decoder architecture based on our IS-RCSC algorithm. Our IS-RCSC decoding algorithm performs the message passing on a binary tree representation of a polar code. Compared to the S-RCSC decoding algorithm, our IS-RCSC decoding algorithm accelerates the computing of the returned soft messages when certain types of nodes are activated. The corresponding hardware architecture of our IS-RCSC decoder is also proposed. In terms of area efficiency, the hardware implementation results demonstrate that our IS-RCSC decoders are 19% to 43% better than decoders in the literature.
Jun Lin 0001, Jin Sha 0001, Li Li 0003, Chenrong Xiong, Zhiyuan Yan 0001, Zhongfeng Wang 0001
ISCAS3
2014 Performance and network power evaluation of tightly mixed SRAM NUCA for 3D Multi-core Network on Chips
abstract
Last level cache (LLC) is crucial for the performance of chip multiprocessors (CMPs), while power is a significant design concern for 3D CMPs. In this paper, we focus on the SRAM-based Non-Uniform Cache Architecture (NUCA) for 3D Multi-core Network-on-Chip (McNoC) systems. A tightly mixed SRAM NUCA for 3D mesh NoC is presented and analyzed. We evaluate the performance and network power with benchmarks based on a full system simulation framework. Experiment results on 16-core 3D NoC systems show that the tightly mixed NUCA could provide up to 31.71% and on average 5.95% performance improvement compared to a base 3D NUCA scheme. The tightly mixed 3D NUCA NoC can reduce network power consumption in 1.07%-15.74% and 9.64% on average compared to a baseline 3D NoCs. Our analysis and experimental results provide a guideline to design efficient 3D NoCs with stacking NUCA.
Li Li 0003, Zhonghai Lu, Axel Jantsch, Minglun Gao
ISCAS2
2012 Unified Architecture for Reed-Solomon Decoder Combined With Burst-Error Correction
abstract
Reed-Solomon (RS) codes are widely used as forward correction codes (FEC) in digital communication and storage systems. Correcting random errors of RS codes have been extensively studied in both academia and industry. However, for burst-error correction, the research is still quite limited due to its ultra high computation complexity. In this brief, starting from a recent theoretical work, a low-complexity reformulated inversionless burst-error correcting (RiBC) algorithm is developed for practical applications. Then, based on the proposed algorithm, a unified VLSI architecture that is capable of correcting burst errors, as well as random errors and erasures, is firstly presented for multi-mode decoding requirements. This new architecture is denoted as unified hybrid decoding (UHD) architecture. It will be shown that, being the first RS decoder owning enhanced burst-error correcting capability, it can achieve significantly improved error correcting capability than traditional hard-decision decoding (HDD) design.
Li Li 0003, Bo Yuan 0001, Zhongfeng Wang 0001, Jin Sha 0001, Hongbing Pan, Weishan Zheng
IEEE Trans. Very Large Scale Integr. Syst.1
2010 Low power decoder design for QC-LDPC codes
abstract
This paper presents a low-power decoder design approach for generic quasi-cyclic low-density parity-check (QC-LDPC) codes based on the layered min-sum decoding algorithm. To reduce the energy consumption, a novel message length-shortening scheme is explored. The check node processing unit (CNU) is accordingly optimized using bit-serial architecture. This low cost design scheme can greatly lower the power consumption while maintaining the necessary throughput required by mobile applications. We further demonstrate the benefits of the proposed techniques by applying the new architecture to the QC-LDPC code in CMMB standard.
Jin Sha 0001, Li Li 0003, Zhongfeng Wang 0001
ISCAS3
2010 Application-level pipelining on Hierarchical NoC
abstract
Multiprocessor System-on-Chip is a promising solution for the high performance Embedded System. This paper is based on an independent research about Hierarchical NoC (Network-on-chip). By integrating 16 ARM cores in the FPGA board, we can bring out the four-channel fade-in and fade-out for real-time streaming media. We present two parallel models for our multiprocessor. One is fine-grained parallelization, with which the speed-up is 7.6, the other module is coarse-grained parallelization, with which the speed-up is higher than 9.2.
Hongbing Pan, Li Li 0003, Minglun Gao, Ning Hou, Gaoming Du, Duoli Zhang
ISCAS4
2010 Layered decoding for non-binary LDPC codes
abstract
In this paper, we present a layered decoding algorithm for non-binary LDPC codes. Differing from the flooding message-passing schedule in conventional designs, the proposed scheme updates check node messages in serial. Furthermore, since fully serial decoding will lead to long latency, the locally-parallel globally-serial schedule is adopted. Basically the check nodes can be divided into several groups, i.e. layers. The layers are processed one by one while the check nodes in each layer are processed in parallel. Simulation results show that the proposed algorithm not only brings some improvement in error correcting performance but also gives some advantage in VLSI implementation of efficient partially parallel decoders.
Jin Sha 0001, Li Li 0003, Zhongfeng Wang 0001
ISCAS3
2009 Towards an Optimal Trade-off of Viterbi Decoder Design
abstract
Viterbi decoder (VD) is widely used in modern communication systems. For low power applications, trace-back approach (TBA) is usually employed for the survivor memory unit (SMU) of VD. However, TBA suffers from long latency and low throughput. Employing multiple memory banks can resolve the throughput issue on a great extent. In this paper, we present efficient schemes to improve the latency issue of conventional TBA by exploiting pre-trace-back method. In the meantime, we adopt buffer-based TBA method to reduce memory access times, thus reduce power assumption significantly. Simulation results show that the proposed decoding schemes cause either zero or negligible performance loss.
Jinjin He, Zhongfeng Wang 0001, Zhiqiang Cui, Li Li 0003
ISCAS4
2009 LDPC Decoder Design for IEEE 802.15 Standard
abstract
This paper presents an efficient decoder design for the LDPC codes in IEEE 802.15 standard. This decoder features by high parallel level, low message memory requirement and code rate flexibility. By processing 72 columns and 72 rows in parallel, it can reach a throughput of 2.8 Gbps to fulfill the standard requirement. Furthermore, the decoder supports three different code rates by employing flexible check node processor units.
Jin Sha 0001, Jun Lin 0001, Li Li 0003, Minglun Gao, Zhongfeng Wang 0001
ISCAS3
2009 Area-efficient Reed-Solomon Decoder Design for 10-100 Gb/s Applications
abstract
With the extensive applications in high-speed communication systems, the current high-throughput Reed-Solomon decoders are required to achieve the target data rates from 10 Gb/s to 100 Gb/s with low hardware complexity. In this paper, pipeline interleaving inversionless Berlekamp-Massey (PI-iBM) algorithm and pipeline interleaving reformulated inversionless Berlekamp-Massey (PI-RiBM) algorithms for decoding Reed-Solomon codes are presented. Based on these two new algorithms PI-iBM and PI-RiBM Reed-Solomon decoders targeted at 10-100 Gb/s applications are developed. Compared with previously published works, the proposed designs can achieve very high throughput with relatively low hardware complexity. Thus they are well suited for modern high data rate communication systems.
Bo Yuan 0001, Li Li 0003, Jin Sha 0001, Zhongfeng Wang 0001
ISCAS2
2009 High-throughput GCM VLSI Architecture for IEEE 802.1ae Applications
abstract
This paper presents a high-throughput GCM VLSI architecture fully compliant to IEEE 802.1ae applications, which can be operated in all modes specified in the standard. Unlike previous works, with the modified parallel GHASH module, the design implements encryption efficiently without knowing the total number of data blocks in advance. Furthermore, a fully subpipelined version of loop-free key expansion architecture is employed to support constant key changes in each clock cycle. An encryptor design example with 2-parallel modified GHASH module is implemented and fabricated in Fujitsu 0.13 mum 1.2 V 1P8M CMOS technology. The ASIC implementation results demonstrate that the maximum operating frequency can reach 764.5 MHz and our design can obtain 97.9 Gb/s throughput with 547 k gates.
Chuan Zhang 0001, Li Li 0003, Jun Xu 0013, Zhongfeng Wang 0001
ISCAS2
2009 Multi-Gb/s LDPC Code Design and Implementation
abstract
Low-density parity-check (LDPC) code, a very promising near-optimal error correction code (ECC), is being widely considered in next generation industry standards. The VLSI implementation of high-speed LDPC decoder remains a big challenge. This paper presents the construction of a new class of implementation-oriented LDPC codes, namelyshift-LDPCcodes. With girth optimization, this kind of codes can perform as well as computer generated random codes. More importantly, the decoder can be efficiently implemented to obtain very high decoding speeds. In addition, more than 50% of message memory can be generally saved over conventional partially parallel decoder architectures. We demonstrate the benefits of the proposed techniques with an application-specific integrated circuit (ASIC) design (in 0.18-mum CMOS) for a 8192-bit regular LDPC code, which can achieve 5 Gb/s throughput at 15 iterations.
Jin Sha 0001, Zhongfeng Wang 0001, Minglun Gao, Li Li 0003
IEEE Trans. Very Large Scale Integr. Syst.4
2007 On the Implementation of Virtual Array Using Configuration Plane
Yongsheng Yin, Li Li 0003, Minglun Gao, Gaoming Du, Yu-Kun Song
APPT2