Jun Lin 0001

dblp:55/1226-1 · DBLP profile ↗
← Back
67ranked-venue papers
10as first author
27since 2021 · last 2026
0009-0005-3505-4847ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 55 · 9 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorArtificial intelligence and machine learning · 4 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 HoloLUT: An Efficient LUT-Based Engine via Holistic Data Processing for Low-bit LLM Inference
abstract
Weight-only quantization enhances the efficiency of large language models (LLMs) by storing weights in low-bit integers while retaining activations at a higher precision. However, the absence of efficient mixed-precision computation support on general-purpose hardware tends to hinder the potential computational gains from low-bit LLMs. While lookup table (LUT)-based methods offer a promising alternative, conventional bit-serial architectures introduce shift-and-accumulate bottlenecks, limiting throughput and energy efficiency. To overcome these limitations, we propose HoloLUT, a novel LUT-based engine that incorporates the Unitary Data Operation paradigm. This paradigm processes weights holistically rather than bit-serially, thereby eliminating shift-and-accumulate operations. Furthermore, a precision-adaptive mapping strategy combined with a unified LUT generator allows HoloLUT to flexibly and efficiently handle various precisions with negligible hardware overhead. Implemented in 28nm CMOS technology, HoloLUT achieves 1.86 × and 2.18 × improvements in area and power efficiency, respectively, compared to state-of-the-art LUT-based accelerators, demonstrating its strong potential for deploying low-bit LLMs in resource-constrained scenarios.
Hui Wang 0083, Weize Ma, Jinming Lu, Jun Lin 0001
ACM Great Lakes Symposium on VLSI4
2026 WiFlow: A Precision-Scalable DNN Training Accelerator Through Winograd Algorithm and Dataflow Co-Design
abstract
To address performance degradation from the domain shift and to support user-specific services while considering privacy, security, and communication overhead, there is an urgent need for efficient on-device training accelerators for deep neural networks (DNNs). Given limited computing resources and battery capacity constraints, implementing complex DNN training on edge devices is extremely challenging. To address these issues, we introduce a Winograd-Integrated Gradient Optimization Framework (WIGOF) for cross-phase operand sharing in the Winograd domain, which significantly reduces the number of multiplications and additions. Additionally, we develop WiFlow, an efficient, precision-scalable on-device training accelerator, minimizing area and power overheads of the dedicated Winograd transformation unit. The WiFlow supports 16-bit floating point (FP16), 16-bit brain floating point (BF16), and 8 and 4-bit fixed point (INT8 and INT4), demonstrating scalable improvements in both computational throughput (TOPS) and energy efficiency (TOPS/W) at low precision. A novel data rearrangement pattern, named channel augmentation, addresses the imperfect decomposition to enhance the utilization of processing element units. Furthermore, we propose a Winograd interleaved block-execution dataflow (WInBlock), along with Hierarchical Adaptive Reuse Memory Optimization (HARM) to improve data reuse and reduce both the amount of DRAM and SRAM access. The end-to-end training of WiFlow is achieved on Xilinx XCVU440 FPGA. WiFlow is also synthesized with a 28nm CMOS technology, achieving an area efficiency of 624 GOPS/mm2and an energy efficiency of 4.4 TOPS/W at a supply voltage of 0.9V and an operating frequency of 500 MHz. WiFlow accomplishes$7.75\times $higher area efficiency and$2.91\times $higher energy efficiency in actual DNN training compared with the state-of-the-art on-device training accelerators.
Hui Wang 0083, Jinming Lu, Weize Ma, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Circuits Syst. I Regul. Pap.6
2026 MCRA: Multicolumn Residue Accumulation Analog Compute-in-Memory Architecture With Time-Domain M-Input ΣΔ ADC
abstract
Analog compute-in-memory (ACIM) architectures offer significant throughput and energy benefits by performing multiplication-and-accumulation (MAC) operations directly within memory arrays. However, their overall efficiency is fundamentally constrained by the high power consumption of the per-column high-resolution analog-to-digital converters (ADCs) required to support modern DNNs (e.g., transformers), many of which demand both high computational precision and large-column throughput. In conventional ADC designs, energy in the noise-limited regime scales near-exponentially, typically by$4\times $per additional bit, making high-resolution ADCs on every column power-prohibitive. This article proposes a high-precision and power-efficient multicolumn residue accumulation (MCRA) ACIM architecture to efficiently support precision-demanding modern DNNs. Each column uses a low-resolution coarse ADC (cADC), while the per-column residuals are accumulated and further quantized by an energy-efficient time-domain multi-input incremental sigma–delta (Mi-$\Sigma \Delta $) fine ADC (fADC). This approach amortizes the near-exponential energy growth across columns, while exploiting the more favorable power-resolution scaling of the time-domain Mi-$\Sigma \Delta $quantization. Postlayout simulations demonstrate a 66.2-dB signal-to-noise-and-distortion ratio (SNDR) per column at only 1/21 the energy of the baseline with per-column high-resolution ADCs, and$0.405\times $(1/2.47) the energy of an energy-saving ADC per column, which achieves a similar SNDR. Circuit- and system-level simulations demonstrate that our MCRA CIM architecture achieves negligible accuracy degradation in precision-demanding ViT tasks while delivering high energy and area efficiency.
Wenlun Zhang, Shimpei Ando, Zhongfeng Wang 0001, Jun Lin 0001, Kentaro Yoshioka
IEEE Trans. Very Large Scale Integr. Syst.5
2025 AiSpGEMM: Accelerating Imbalanced SpGEMM on FPGAs with Flexible Interconnect and Intra-row Parallel Merging
abstract
The row-wise product algorithm shows significant potential for sparse matrix-matrix multiplication (SpGEMM) on hardware accelerators. Recent studies have made notable progress in accelerating SpGEMM using this algorithm. However, several challenges remain in accelerating imbalanced SpGEMM, where the distribution of non-zero elements across different rows is imbalanced. These challenges include: (1) the fixed dataflow of the merger tree, which leads to lower PE utilization, and (2) highly imbalanced data distributions, such as single rows with numerous non-zero elements, which result in intensive computations. This imbalance significantly challenges SpGEMM acceleration, leading to time-consuming processes that dominate overall computation time. In this paper, we propose AiSpGEMM to accelerate imbalanced SpGEMM on FPGAs. First, we improved the C2SR format to adapt it for imbalanced SpGEMM acceleration based on the row-wise product algorithm. This reduces off-chip memory bank conflicts and increases data reuse of matrix B. Secondly, we design a reconfigurable merger (R-merger) with flexible interconnects to improve PE utilization. Additionally, we propose an intra-row parallel merging algorithm and its corresponding hardware architecture, the parallel merger (P-merger), to accelerate intensive operations. Experimental results demonstrate that AiSpGEMM achieves a geometric mean (geomean) speedup of 5.8× compared to the state-of-the-art FPGA-based SpGEMM accelerator. In Geomean, AiSpGEMM achieves a 3.0× speedup and a 9.8× improvement in energy efficiency compared to the NVIDIA cuSPARSE library running on an NVIDIA A6000 GPU. Moreover, AiSpGEMM-21 demonstrated a 4× increase in average throughput compared to the same GPU.
Enhao Tang, Hao Zhou 0008, Guohao Dai 0001, Jun Lin 0001, Kun Wang 0005
DATE5
2025 A High-Speed 8-bit Single-Channel SAR ADC with Tailored Bit Intervals
abstract
This paper presents a high-speed 8-bit asynchronous successive approximation register (SAR) analog-to-digital converter (ADC) featuring tailored bit intervals (TBI). The design employs built-in SAR logic to set delays autonomously, eliminating the need for digital assistance, and thereby reducing both power and area consumption. This approach also effectively shortens the waiting time before lower-bit comparisons, enabling faster conversions. The ADC is simulated in the 16 nm process, occupying only 0.0012 mm2, with post-simulation conducted under various extreme process and temperature conditions. Compared to prior works, our design exhibits notable performance advantages, achieving an ENOB of 7.38 bits at TT 25°C with a power consumption of 6.94 mW. Furthermore, the designed TBI-ADC attains a sampling rate of 1.6 GS/s at FF -40°C, representing a 33% increase over the fastest previously reported single-channel, 1b/cycle, 8-bit SAR ADC.
Ruida Wang, Congyi Zhu, Zhongfeng Wang 0001, Jun Lin 0001
ISCAS6
2025 HWSA: A High-Ratio Weight Sparse Accelerator for Efficient CNN Inference
abstract
Pruning has emerged as an effective technique for compressing convolutional neural networks (CNNs) by eliminating redundant weights, achieving lightweight models with negligible loss in inference accuracy. To leverage the sparsity for acceleration, many accelerators built for sparse CNNs have been developed. Existing hardware accelerators can perform well with structured or customized sparsity patterns. However, when facing the unstructured sparsity which can achieve higher compression rates, the corresponding hardware always suffers from insufficient utilization of computational resources, severe load balance between process elements, and significant overhead in logic resources, which results in a reduction of throughput, making it difficult to leverage high ratio sparsity for acceleration effectively. An efficient CNN inference accelerator that can handle both structured and unstructured sparse networks is proposed to address these issues. By flexibly employing multiple parallel computation methods combined with carefully developed sorting algorithms, the proposed architecture mitigates the hardware utilization inefficiencies caused by unstructured sparsity. Through a hardware-software methodology, new sparsity rules are introduced, nearly eliminating the load imbalance issue. The proposed Processing Element (PE) architecture can effectively select inputs for sparse networks while reducing the overhead of logic resources. The proposed architecture is implemented on the XCVU9P FPGA, achieving a frequency of 200 MHz. It achieves a computational throughput of 350.49 GOPs and 326.01 GOPs on ResNet-50 and ResNet-152, respectively, demonstrating a 1.2-$2.1\times $DSP efficiency improvement compared to previous works.
Xuejing Dai, Jinze Zhang, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 A RISC-V Domain-Specific Processor for Deep Learning-Based Channel Estimation
abstract
Channel estimation (CE) is a critical component in the massive multi-input multi-output (MIMO) communication systems. Compared with conventional CE algorithms, deep learning (DL)-based approach becomes a promising alternative, due to its capability of offering enhanced performance and robustness across diverse scenarios. However, efficient DL-based CE algorithms have two key properties that make them challenging for implementation in existing architectures at the edge side: the diversity of deep neural networks (DNNs) and CE strategies, and the involvements of multiple computation-intensive tasks that compass conventional signal processing, artificial intelligence (AI) inference, and online learning. To address these challenges, a domain-specific processor based on an extended RISC-V instruction set architecture (ISA) is proposed to perform these DL-based CE algorithms. First, a dedicated RISC-V ISA extension is developed to support all essential operations required by a DL-based CE algorithm, such as matrix inversion, in a flexible manner. Building on the customized ISA extension, a highly adaptable and scalable RISC-V processor is developed, featuring scalar and vector posit arithmetic units to alleviate high computational and memory demands of DNNs during both inference and training phase. Additionally, a coarse-grained matrix accelerator is integrated to expedite various matrix operations ensuring high throughput. In this way, both high flexibility and computational efficiency are achieved. Finally, our processor is implemented on a TSMC 28-nm technology. Implementation results show that the processor achieves a speedup of$5.16\sim 6.80\times $for all matrix operations compared with the state-of-the-art work. Moreover, the proposed processor provides an area efficiency improvement of$1.61\times $and an energy efficiency enhancement of$6.6\sim 15.4\times $compared to the open-source vector processor Ara. Notably, this work is the first RISC-V domain-specific processor tailored for diverse DL-based CE algorithms.
Chuanning Wang, Yangcan Zhou, Shaowei Wang 0001, Chuan Zhang 0001, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 DiffAccel: Accelerating Diffusion Models Through Adaptive Feature Optimization and Dynamic Hardware Adaptation
Enhao Tang, Weize Ma, Yudan Jiang, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2025 SPEED: A Scalable RISC-V Vector Processor Enabling Efficient Multiprecision DNN Inference
abstract
Deploying deep neural networks (DNNs) on those resource-constrained edge platforms is hindered by their substantial computation and storage demands. Quantized multiprecision DNNs (MP-DNNs), denoted as MP-DNNs, offer a promising solution for these limitations but pose challenges for the existing RISC-V processors due to complex instructions, suboptimal parallel processing, and inefficient dataflow mapping. To tackle the challenges mentioned above, SPEED, a scalable RISC-V vector (RVV) processor, is proposed to enable efficient MP-DNN inference, incorporating innovations in customized instructions, hardware architecture, and dataflow mapping. First, some dedicated customized RISC-V instructions are introduced based on RVV extensions to reduce the instruction complexity, allowing SPEED to support processing precision ranging from 4- to 16-bit with minimized hardware overhead. Second, a parameterized multiprecision tensor unit (MPTU) is developed and integrated within the scalable module to enhance parallel processing capability by providing reconfigurable parallelism that matches the computation patterns of diverse MP-DNNs. Finally, a flexible mixed dataflow method is adopted to improve computational and energy efficiency according to the computing patterns of different DNN operators. The synthesis of SPEED is conducted on TSMC 28-nm technology. Experimental results show that SPEED achieves a peak throughput of 737.9 GOPS and an energy efficiency of 1383.4 GOPS/W for 4-bit operators. Furthermore, SPEED exhibits superior area efficiency compared with prior RVV processors, with the enhancements of$5.9\sim 26.9\times $and$8.2\sim 18.5\times $for 8-bit operator and best integer performance, respectively, which highlights SPEED’s significant potential for efficient MP-DNN inference.
Chuanning Wang, Chao Fang 0005, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2024 A Precision-Scalable RISC-V DNN Processor with On-Device Learning Capability at the Extreme Edge
abstract
Extreme edge platforms, such as in-vehicle smart devices, require efficient deployment of quantized deep neural networks (DNNs) to enable intelligent applications with limited amounts of energy, memory, and computing resources. However, many edge devices struggle to boost inference throughput of various quantized DNNs due to the varying quantization levels, and these devices lack floating-point (FP) support for on-device learning, which prevents them from improving model accuracy while ensuring data privacy. To tackle the challenges above, we propose a precision-scalable RISC-V DNN processor with on-device learning capability. It facilitates diverse precision levels of fixed-point DNN inference, spanning from 2-bit to 16-bit, and enhances on-device learning through improved support with FP16 operations. Moreover, we employ multiple methods such as FP16 multiplier reuse and multi-precision integer multiplier reuse, along with balanced mapping of FPGA resources, to significantly improve hardware resource utilization. Experimental results on the Xilinx ZCU102 FPGA show that our processor significantly improves inference throughput by 1.6$\sim 14.6\times$ and energy efficiency by 1.1$\sim 14.6\times$ across various DNNs, compared to the prior art, XpulpNN. Additionally, our processor achieves a $16.5\times$ higher FP throughput for on-device learning.
Longwei Huang, Chao Fang 0005, Jun Lin 0001, Zhongfeng Wang 0001
ASPDAC4
2024 A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
abstract
RISC-V processors encounter substantial challenges in deploying multi-precision deep neural networks (DNNs) due to their restricted precision support, constrained throughput, and suboptimal dataflow design. To tackle these challenges, a scalable RISC-V vector (RVV) processor, namely SPEED, is proposed to enable efficient multi-precision DNN inference by innovations from customized instructions, hardware architecture, and dataflow mapping. Firstly, dedicated customized RISC-V instructions are proposed based on RVV extensions, providing SPEED with fine-grained control over processing precision ranging from 4 to 16 bits. Secondly, a parameterized multi-precision systolic array unit is incorporated within the scalable module to enhance parallel processing capability and data reuse opportunities. Finally, a mixed multi-precision dataflow strategy, compatible with different convolution kernels and data precision, is proposed to effectively improve data utilization and computational efficiency. We perform synthesis of SPEED in TSMC 28nm technology. The experimental results demonstrate that SPEED achieves a peak throughput of 287.41 GOPS and an energy efficiency of 1335.79 GOPS/W at 4-bit precision condition, respectively. Moreover, when compared to the pioneer open-source vector processor Ara, SPEED provides an area efficiency improvement of 2.04× and 1.63× under 16-bit and 8-bit precision conditions, respectively, which shows SPEED’s significant potential for efficient multi-precision DNN inference.
Chuanning Wang, Chao Fang 0005, Zhongfeng Wang 0001, Jun Lin 0001
ISCAS5
2024 An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
abstract
Convolutional neural networks (CNNs) with large kernels, drawing inspiration from the key operations of vision transformers (ViTs), have demonstrated impressive performance in various vision-based applications. To address the issue of computational efficiency degradation in existing designs for supporting large-kernel convolutions, an FPGA-based inference accelerator is proposed for the efficient deployment of CNNs with arbitrary kernel sizes. Firstly, a Z-flow method is presented to optimize the computing data flow by maximizing data reuse opportunity. Besides, the proposed design, incorporating the kernel-segmentation (Kseg) scheme, enables extended support for large-kernel convolutions, significantly reducing the storage requirements for overlapped data. Moreover, based on the analysis of typical block structures in emerging CNNs, vertical-fused (VF) and horizontal-fused (HF) methods are developed to optimize CNN deployments from both computation and transmission perspectives. The proposed hardware accelerator, evaluated on Intel Arria 10 FPGA, achieves up to 3.91 × better DSP efficiency than prior art on the same network. Particularly, it demonstrates efficient support for large-kernel CNNs, achieving throughputs of 169.68 GOPS and 244.55 GOPS for RepLKNet-31 and PyConvResNet-50, respectively, both of which are implemented on hardware for the first time.
Miaoxin Wang, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS3
2024 WinTA: An Efficient Reconfigurable CNN Training Accelerator With Decomposition Winograd
abstract
Convolutional neural networks (CNNs) are expected to bridge the domain shift between the training data and real-world tasks. Moreover, the efficient training of CNNs on resource-constrained platforms has become more important because of communication latency and privacy concerns. However, deploying CNN training on edge devices is challenging due to the intensive computation and diverse computational patterns. In this work, we firstly propose a hybrid decomposition Winograd (HDW) method that significantly reduces the number of multiplications and flexibly handles various convolution operations during training. Secondly, we design a reconfigurable CNN training accelerator, named WinTA, utilizing a set of unified transformation units to support various Winograd operations. Thirdly, we implement an efficient and flexible data access scheme using a hierarchical barrel shifter network (HBSN). Experimental results on the Xilinx Alveo U50 FPGA Card demonstrate that WinTA effectively accelerates CNN training. Compared to CPU and GPU implementations, WinTA achieves speedups of 7.1 texttimes and 1.65 texttimes, respectively, while improving energy efficiency by 26.6 texttimes and 10.4 texttimes, respectively. Additionally, our design provides 1.24 texttimes and 2.04 texttimes improvements in terms of throughput and resource efficiency compared to prior-art FPGA-based training accelerator.
Jinming Lu, Hui Wang 0083, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 TECO: A Unified Feature Map Compression Framework Based on Transform and Entropy
abstract
The massive memory accesses of feature maps (FMs) in deep neural network (DNN) processors lead to huge power consumption, which becomes a major energy bottleneck of DNN accelerators. In this article, we propose a unified framework named Transform and Entropy-based COmpression (TECO) scheme to efficiently compress FMs with various attributes in DNN inference. We explore, for the first time, the intrinsic unimodal distribution characteristic that widely exists in the frequency domain of various FMs. In addition, a well-optimized hardware-friendly coding scheme is designed, which fully utilizes this remarkable data distribution characteristic to encode and compress the frequency spectrum of different FMs. Furthermore, the information entropy theory is leveraged to develop a novel loss function for improving the compression ratio and to make a fast comparison among different compressors. Extensive experiments are performed on multiple tasks and demonstrate that the proposed TECO achieves compression ratios of in ResNet-50 on image classification, in UNet on dark image enhancement, and in Yolo-v4 on object detection while keeping the accuracy of these models. Compared with the upper limit of the compression ratio for original FMs, the proposed framework achieves the compression ratio improvement of 21%, 157%, and 152% on the above models.
Yubo Shi, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Amoeba: An Efficient and Flexible FPGA-Based Accelerator for Arbitrary-Kernel CNNs
abstract
Inspired by the key operation of vision transformers (ViTs), convolutional neural networks (CNNs) have widely adopted arbitrary-kernel convolutions to achieve high performance in diverse vision-based tasks. However, existing hardware efforts primarily focus on implementing CNN models that consist of a stack of small kernels, which poses challenges in supporting large-kernel convolutions. To address this limitation, we propose Amoeba, a flexible field-programmable gate array (FPGA)-based inference accelerator designed for efficiently supporting CNNs with arbitrary kernel sizes. Specifically, we present an optimized dataflow approach in collaboration with the Z-flow method and kernel-segmentation (Kseg) scheme, which enables flexible support for arbitrary-kernel convolutions without sacrificing efficiency. Additionally, we incorporate vertical-fused (VF) and horizontal-fused (HF) methods into the layer execution schedule to optimize the computation and data transfer process. To further enhance the CNN deployment performance, we employ the loop tiling scheme search (LTSS) method, guided by a fine-grained performance model, during the early design phase. The proposed Amoeba accelerator is evaluated on Intel Arria 10 SoC FPGA. The experimental results demonstrate excellent performance on prevalent and emerging CNNs, achieving a throughput of up to 286.2 GOPs. Notably, Amoeba achieves 4.36$\times$better DSP efficiency compared to prior works on the same network, highlighting its superior utilization of hardware resources for CNN inference tasks.
Miaoxin Wang, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2023 An Efficient Massive MIMO Detector Based on Approximate Expectation Propagation
abstract
Among expectation propagation (EP)-based massive multiple-input–multiple-output (MIMO) detection algorithms, EP with weighted Neumann-series approximation (EPA-wNSA) has the lowest computational complexity while requiring many iterations to guarantee the detection performance, which severely limits the throughput of hardware implementations. Through the joint optimization of algorithm and hardware architecture, we propose an EP-based detector with higher throughput and area efficiency. First, the second-order Richardson iteration (SORI) algorithm is employed to replace the wNSA algorithm for higher convergence speed. Then three algorithmic transformations are proposed to minimize the overall complexity. Simulation results show that the proposed EPA-SORI algorithm requires much fewer iterations to achieve comparable or even better detection performance compared with EPA-wNSA. Furthermore, an efficient detector architecture is delicately designed by incorporating multiple optimization methods, such as reverse data flow, advanced addition, and rounding cells. Implemented with the Taiwan Semiconductor Manufacturing Company (TSMC) 28-nm CMOS technology, the proposed detector has$2.2 \times $higher throughput than the state-of-the-art EP-based detector.
Yangyang Chen 0005, Suwen Song, Zhongfeng Wang 0001, Jun Lin 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2022 An Efficient Hardware Accelerator for Sparse Transformer Neural Networks
abstract
Transformers have been an indispensable staple in deep learning. However, it is challenging to realize efficient deployment for Transformer-based model due to their substantial computation and memory demands. To address this issue, we present an efficient sparse Transformer accelerator on FPGA, namely STA, by exploiting N:M fine-grained structured sparsity. Our design features not only a unified computing engine capable of performing both sparse and dense matrix multiplications with high computational efficiency, but also a scalable softmax module eliminating the latency from intermediate off-chip data communication. Experimental results show that our implementation achieves the lowest latency compared to CPU, GPU, and prior FPGA-based accelerators. Moreover, compared with the state-of the-art FPGA-based accelerators, it can achieve up to $12.28\times$ and $51.00\times$ improvement on energy efficiency and MAC efficiency, respectively.
Chao Fang 0005, Shouliang Guo, Jun Lin 0001, Zhongfeng Wang 0001, Ming Kai Hsu
ISCAS4
2022 Accelerate Three-Dimensional Generative Adversarial Networks Using Fast Algorithm
abstract
Three-dimensional generative adversarial networks (3D-GAN) have attracted widespread attention in three-dimension (3D) visual tasks. 3D deconvolution (DeConv), as an important computation of 3D-GAN, significantly increases computational complexity compared with 2D DeConv. 3D DeConv has become a bottleneck for the acceleration of 3D-GAN. Previous accelerators suffer from several problems, such as large memory requirements and resource underutilization. To handle the above issues, a fast algorithm for 3D DeConv (F3DC) is proposed in this paper. F3DC applies a fast algorithm to reduce the number of multiplications and achieves a significant algorithmic strength reduction. Besides, F3DC removes the extra memory requirement for overlapped partial sums and avoids computational imbalance to fully utilize resources. Moreover, we design an F3DC-based hardware architecture, which consists of four fast processing units (FPUs). Each FPU includes a pre-process module, a EWMM module and a post-process module for F3DC transformation. By implementing our design on the Xilinx VC709 platform for 3D-GAN, we achieve a throughput up to 1700 GOPS and 4× computational efficiency improvement compared with prior works.
Ziqi Su, Wendong Mao, Zhongfeng Wang 0001, Jun Lin 0001
ISCAS4
2022 A Reconfigurable Approach for Deconvolutional Network Acceleration with Fast Algorithm
abstract
Recently, deconvolutional neural network (DeCNN) has attracted widespread attention in various applications. The deconvolution (DeConv), as the main operation in DeCNN, has become the bottleneck of acceleration, due to its high computational complexity. Previous works have introduced fast algorithms such as the cascaded fast FIR algorithm (CFFA) and the Winograd algorithm to reduce the computational complexity of DeConv for the applications on mobile devices. Since these fast algorithms need different computing parameters to accelerate various operations, directly applying these methods to process DeCNNs with different kernels usually causes limited flexibility. To address this problem, we propose a reconfigurable scheme based on the fast transformation algorithm (FTA) to accelerate multiple types of DeConvs, minimizing the hardware overhead for reconfigurability. Based on this scheme, a reconfigurable hardware architecture is developed to support several types of DeConvs. In addition, an adaptive dataflow is proposed to handle different convolutional layers. The presented design can support several types of operations and achieve up to 222.54 GOPS under 210 MHz on the Intel Arria 10SX FPGA platform, which shows our design can obtain better flexibility and computational efficiency compared with prior arts.
Peixiang Yang, Wendong Mao, Zhongfeng Wang 0001, Jun Lin 0001
ISCAS4
2022 Efficient Software Implementation of the SIKE Protocol Using a New Data Representation
abstract
Thanks to relatively small public and secret keys, the Supersingular Isogeny Key Encapsulation (SIKE) protocol made it into the third evaluation round of the post-quantum standardization project of the National Institute of Standards and Technology (NIST). Even though a large body of research has been devoted to the efficient implementation of SIKE, its latency is still undesirably long for many real-world applications. Most existing implementations of the SIKE protocol use the Montgomery representation for the underlying field arithmetic since the corresponding reduction algorithm is considered the fastest method for performing multiple-precision modular reduction. In this paper, we propose a new data representation for supersingular isogeny-based Elliptic-Curve Cryptography (ECC), of which SIKE is a sub-class. This new representation enables significantly faster implementations of modular reduction than the Montgomery reduction, and also other finite-field arithmetic operations used in ECC can benefit from our data representation. We implemented all arithmetic operations in C using the proposed representation such that they have constant execution time and integrated them to the latest version of the SIKE software library. Using four different parameters sets, we benchmarked our design and the optimized generic implementation on a 2.6 GHz Intel Xeon E5-2690 processor. Our results show that, for the prime of SIKEp751, the proposed reduction algorithm is approximately 2.61 times faster than the currently best implementation of Montgomery reduction, and our representation also enables significantly better timings for other finite-field operations. Due to these improvements, we were able to achieve a speed-up by a factor of about 1.65, 2.03, 1.61, and 1.48 for SIKEp751, SIKEp610, SIKEp503, and SIKEp434, respectively, compared to state-of-the-art generic implementations.
Jing Tian 0004, Piaoyang Wang, Zhe Liu 0001, Jun Lin 0001, Zhongfeng Wang 0001, Johann Großschädl
IEEE Trans. Computers4
2022 A Proximal Iteratively Reweighted Approach for Efficient Network Sparsification
abstract
The huge size of deep neural networks makes it difficult to deploy on the embedded platforms with limited computation resources directly. In this article, we propose a novel trimming approach to determine the redundant parameters of the trained deep neural network in a layer-wise manner to produce a compact neural network. This is achieved by minimizing a nonconvex sparsity-inducing term of the network parameters while maintaining the response close to the original one. We present a proximal iteratively reweighted method to resolve the resulting nonconvex model, which approximates the nonconvex objective by a weighted l1 norm of the network parameters. Moreover, to alleviate the computational burden, we develop a novel termination criterion during the subproblem solution, significantly reducing the total pruning time. Global convergence analysis and a worst-case O(1/k) ergodic convergence rate for our proposed algorithm is established. Numerical experiments demonstrate the proposed approach is efficient and reliable.
Hao Wang 0045, Yuanming Shi, Jun Lin 0001
IEEE Trans. Computers4
2022 Rethinking Adaptive Computing: Building a Unified Model Complexity-Reduction Framework With Adversarial Robustness
abstract
Adaptive computing (AC) is a technique to dynamically select the layers to pass in a prespecified deep neural network (DNN) according to the input samples. In previous literature, AC was deemed as a standalone complexity-reduction skill. This brief studies AC through a different lens: we investigate how this strategy interacts with mainstream compression techniques in a unified complexity-reduction framework and whether its "input sample related" feature helps with the improvement of model robustness. Following this direction, we first propose a defensive accelerating branch (DAB) based on the AC strategy that can reduce the average computational cost and inference time of DNNs with higher accuracy compared with its counterparts. Then, the proposed DAB is jointly applied with the mainstream parameterwise compression skills, pruning and quantization, to build a unified complexity-reduction framework. Extensive experiments are conducted, and the results reveal quasi-orthogonality between the input-related and parameterwise complexity-reduction skills, which means that the proposed AC can be integrated into an off-the-shelf compressed model without hurting its accuracy. Besides, the robustness of the proposed compression framework is explored, and the experimental results demonstrate that DAB can be used as both the detector and the defensive tool when the model is under adversarial attacks. All these findings shed light on the great potential of DAB in building a unified complexity-reduction framework with both a high compression ratio and great adversarial robustness.
Liulu He, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2021 Evaluations on Deep Neural Networks Training Using Posit Number System
abstract
The training of Deep Neural Networks (DNNs) brings enormous memory requirements and computational complexity, which makes it a challenge to train DNN models on resource-constrained devices. Training DNNs with reduced-precision data representation is crucial to mitigate this problem. In this article, we conduct a thorough investigation on training DNNs with low-bit posit numbers, a Type-III universal number (Unum). Through a comprehensive analysis of quantization with various data formats, it is demonstrated that the posit format shows great potential to be employed in the training of DNNs. Moreover, a DNN training framework using 8-bit posit is proposed with a novel tensor-wise scaling scheme. The experiments show the same performance as the state-of-the-art (SOTA) across multiple datasets (MNIST, CIFAR-10, ImageNet, and Penn Treebank) and model architectures (LeNet-5, AlexNet, ResNet, MobileNet-V2, and LSTM). We further design an energy-efficient hardware prototype for our framework. Compared to the standard floating-point counterpart, our design achieves a reduction of 68, 51, and 75 percent in terms of area, power, and memory capacity, respectively.
Jinming Lu, Chao Fang 0005, Mingyang Xu, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Computers4
2021 Design of High-Performance and Area-Efficient Decoder for 5G LDPC Codes
abstract
Low-density parity-check (LDPC) code as a very promising error-correction code has been adopted as the channel coding scheme in the fifth-generation (5G) new radio. However, it is very challenging to design a high-performance decoder for 5G LDPC codes because their inherent numerous degree-1 variable-nodes are very prone to be erroneous. In this article, the problem is solved gracefully by developing a low-complexity check-node update function, greatly improving the reliability of check-to-variable messages. By further incorporating the proposed column degree adaptation strategy, our decoder could offer a 0.4dB performance gain over the existing ones. In addition, this article presents an efficient 5G LDPC decoder architecture. Benefiting the specific structure of 5G LDPC codes, layer merging, split storage method, and selective-shift structure are introduced to facilitate a significant reduction of decoding delay and area consumption. Implementation result on 90-nm CMOS technology demonstrates that the proposed decoder architecture yields an impressive improvement in throughput-to-area ratio, achieving up to 173.3% compared to conventional design.
Hangxuan Cui, Fakhreddine Ghaffari, Khoa Le, David Declercq, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2021 Low-Latency Hardware Accelerator for Improved Engle-Granger Cointegration in Pairs Trading
abstract
Pairs trading is a solidly profitable strategy in the algorithmic trading area, and an important step of this strategy is selecting pairs of stocks. Compared with other existing pairs selection approaches, the Engle-Granger cointegration is more stable and reliable. Nowadays, as trading is becoming faster and faster in stock markets all over the world, it is necessary to accelerate the pairs selection process to increase potential profits. However, intensive computations and complicated data flow in the cointegration approach bring challenges to hardware acceleration. In this paper, for the first time, we propose an efficient and hardware-friendly computing scheme to accelerate the Engle-Granger cointegration. Besides, a novel algorithmic strength reduction strategy and approximation methods are used to significantly reduce the complexity of the proposed scheme. Based on the improved algorithm, both FPGA and ASIC accelerators are developed. The implementation results show that our FPGA and ASIC accelerators perform 36× and 207× faster than GPU, respectively. Thus, our design can significantly reduce the latency of the pairs selection process, and make more profit for investors and traders.
Shuang Liang 0006, Siyuan Lu 0002, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 An Efficient and Flexible Accelerator Design for Sparse Convolutional Neural Networks
abstract
Designing hardware accelerators for convolutional neural networks (CNNs) has recently attracted tremendous attention. Plenty of existing accelerators are built for dense CNNs or structured sparse CNNs. By contrast, unstructured sparse CNNs can achieve higher compression ratio with equivalent accuracy. However, their corresponding hardware implementations generally suffer from load imbalance and conflict access to on-chip buffers, which results in under utilization of processing elements (PEs). To tackle these issues, we propose a hardware/power-efficient and highly flexible architecture to support both unstructured and structured sparse CNNs with various configurations. Firstly, we propose an efficient weight reordering algorithm to preprocess compressed weights and balance the workload of PEs. Secondly, an adaptive on-chip dataflow, namely hybrid parallel (HP) dataflow, is introduced to promote weight reuse. Thirdly, the partial fusion scheme, which was first introduced in one of our prior works, is incorporated as the off-chip dataflow. Benefited from dataflow optimizations, the repetitive data exchanges between on-chip buffers and external memories are significantly reduced. We implement the design on the Intel Arria10 SX660 platform and evaluate with MobileNet-v2, ResNet-50, and ResNet-18 on ImageNet dataset. Compared to existing sparse accelerators on FPGAs, the proposed accelerator can achieve 1.35 ~ 1.81× improvement in power efficiency with the same sparsity. Compared to prior dense accelerators, this accelerator can achieve an improvement of 1.92 ~ 5.84× in DSP efficiency.
Xiaoru Xie, Jun Lin 0001, Zhongfeng Wang 0001, Jinghe Wei
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 Fast Modular Multipliers for Supersingular Isogeny-Based Post-Quantum Cryptography
abstract
As one of the postquantum protocol candidates, the supersingular isogeny key encapsulation (SIKE) protocol delivers promising public and secret key sizes over other candidates. Nevertheless, the considerable computations form the bottleneck and limit its practical applications. The modular multiplication operations occupy a large proportion of the overall computations required by the SIKE protocol. The VLSI implementation of the high-speed modular multiplier remains a big challenge. In this article, we propose three improved modular multiplication algorithms based on an unconventional radix for this protocol, all of which cost about 20% fewer computations than the prior art. Besides, a multiprecision scheme is also introduced for the proposed algorithms to improve the scalability in hardware implementation, resulting in three new algorithms. We then present very efficient high-speed constant-time modular multiplier architectures for the six algorithms. It is shown that these new architectures can be extensively pipelined and highly optimized to obtain high throughput and low latency. The field-programmable gate array (FPGA) implementation results show that all proposed multipliers achieve much higher throughput than previous designs, but the increase in resources is relatively small. In addition, the multipliers without the multiprecision scheme have very low latency, which is very friendly to high-speed applications of the SIKE protocol.
Jing Tian 0004, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2020 In-Memory Computing: The Next-Generation AI Computing Paradigm
abstract
To overcome the memory bottleneck of von-Neuman architecture, various memory-centric computing techniques are emerging to reduce the latency and energy consumption caused by data communication. The great success of artificial intelligence (AI) algorithms, which involve a large number of computations and data movements, has motivated and accelerated the recent researches of in-memory computing (IMC) techniques to significantly reduce or even diminish the accesses of off-chip data, where memory is not only storing data but can also directly output computation results. For example, the multiply-and-accumulate (MAC) operations in deep learning algorithms can be realized by accessing the memory using the input activations. This paper will investigate the recent trends of IMC from techniques (SRAM, flash, RRAM and other types of non-volatile memory) to architecture and to applications, which will serve as a guide to the future advances on computing in-memory (CIM).
Yufei Ma 0002, Yuan Du, Jun Lin 0001, Zhongfeng Wang 0001
ACM Great Lakes Symposium on VLSI4
2020 LSTM-Based Quantitative Trading Using Dynamic K-Top and Kelly Criterion
abstract
With the strong capability of modeling time sequence, long short-term memory (LSTM) networks have been widely applied to predicting financial time series. This has attracted tremendous attention in the quantitative trading area. A complete quantitative trading system usually has three tasks, including market timing, stock selection, and portfolio management. In this paper, we present an LSTM-based quantitative trading system and optimize this system from the following two aspects. Firstly, in the process of stock selection, we first introduce the dynamic K-top method in the LSTM-based quantitative trading system to follow the market change. Secondly, concerning portfolio management, we further incorporate the Kelly Criterion to attain an appropriate position ratio. Taking CSI300 constituent stocks as the study example, extensive experiments have been carried out to show the superiority of the proposed method. In comparison with the straight forward LSTM-based trading strategy, the improved LSTM-based trading strategy with the dynamic K-top method and the Kelly Criterion can achieve an increase of 44.97% over ten days in terms of accumulative return. In addition, our novel method can gain a win ratio of 55.95%, a monthly alpha of 0.16, a monthly Sharpe ratio of 2.17, and a monthly Sortino ratio of 2.96 disregarding the transaction costs.
Binjing Li, Keli Xie, Siyuan Lu 0002, Jun Lin 0001, Zhongfeng Wang 0001
IJCNN4
2020 Hardware Accelerator for Engle-Granger Cointegration in Pairs Trading
abstract
Pairs trading is a classic strategy in the algorithmic trading area and has achieved great success in the stock market. It consists of two stages: pairs selection and trading based on the selected stock pairs. The process of pairs selection is the key to higher returns. Among existing pairs selection methods, pairs trading based on Engle-Granger cointegration has been proven to be superior. However, the cointegration approach is computationally expensive and brings high latency which may greatly affect the returns. In this paper, the Engle-Granger cointegration algorithm is drastically simplified. Meanwhile, a low latency hardware architecture is proposed for the modified algorithm. In the experiment of selecting stocks of length 5000, our hardware design is more than 1290× faster than CPU and 190× faster than GPU. To the best of our knowledge, this is the first work on hardware accelerator for Engle-Granger cointegration in open literature.
Shuang Liang 0006, Siyuan Lu 0002, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS3
2020 A Three-Level Scoring System for Fast Similarity Evaluation Based on Smith-Waterman Algorithm
abstract
The Smith-Waterman (S-W) algorithm is widely adopted by the state-of-the-art DNA sequence aligners in next-generation sequencing (NGS). Prevailing read aligners, such as BWA-MEM and Bowtie 2, use the S-W algorithm to implement the seed-and-extend paradigm. In this work, we further extend the functionality of the S-W algorithm to evaluate the similarity between a pair of sequences without going through traceback process, and design a three-level hardware scoring system to compute final result efficiently. The system is made reconfigurable to align pairs of sequences of various length with a restriction of maximum number of errors. Experimental results show that the system can achieve a throughput of 685Mb/s at 69 iterations in the case of 126bp and the accuracy rate of the outputs is over 98% campared with software results. To the best of our knowledge, this is the first hardware implementation for a similarity evaluation system based on the S-W algorithm.
Jiajun Wu 0025, Minghao Li 0001, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS4
2020 A lightweight face detector by integrating the convolutional neural network with the image pyramid
Jiapeng Luo, Jiaying Liu 0001, Jun Lin 0001, Zhongfeng Wang 0001
Pattern Recognit. Lett.3
2020 Information Storage Bit-Flipping Decoder for LDPC Codes
abstract
Tabu-list random-penalty gradient descent bit-flipping (TRGDBF) decoder is the state-of-the-art hard-decision low-density parity-check (LDPC) decoder in terms of error-correction performance on binary symmetric channel (BSC). However, the TRGDBF decoder suffers from a long critical path caused by the global maximum-finding operation, limiting the achievable throughput. This brief proposes an information storage bit-flipping (ISBF) decoder to solve this problem. Different from the existing bit-flipping (BF) decoders which adopt serial decoding manner, in the ISBF decoder, by storing the previous decoding information, the global maximum-finding operation can be executed in parallel to other decoding operations, significantly shortening the critical path. Moreover, a nonuniform flipping rule is incorporated to achieve a better decoding performance. We also present an efficient architecture to implement the ISBF decoder. The design example demonstrates that compared to other hard-decision BF decoders, the ISBF decoder could provide both the best decoding performance and throughput on BSC.
Hangxuan Cui, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2020 F-DNA: Fast Convolution Architecture for Deconvolutional Network Acceleration
abstract
Deconvolutional neural network (DeCNN), such as fully convolutional network (FCN) and generative adversarial network (GAN), has shown great potential in various vision tasks. Convolution and deconvolution, the two major operations of DeCNN, both require real-time hardware acceleration. However, some previous designs for deconvolutions require large memory for overlapped results, while others incur computation imbalance and cause resource underutilization. In this article, we propose an efficient method to convert deconvolutions to convolutions, which enables balanced computations to make full use of processing elements. Based on the fast FIR algorithm, a reconfigurable conv-deconv unit (RCU) with low complexity is designed, which can support various types of convolutions and deconvolutions. By exploiting the computing characteristics of RCUs, a computation-balance scheme is developed to eliminate large memory requirements caused by overlapped results. In addition, a fast convolution architecture for deconvolutional network acceleration (F-DNA) is proposed. The dataflow of F-DNA improves the computation efficiency through input data reuse. The architecture is implemented on Xilinx Virtex-UltraScale, for two typical DeCNNs, DCGAN and FSRCNN. Implementation results show that the proposed design outperforms existing works significantly, particularly in terms of computation efficiency and memory requirements.
Wendong Mao, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2019 An Enhanced Offset Min-Sum decoder for 5G LDPC Codes
abstract
This paper presents an Enhanced Offset Min-Sum (EOMS) decoder for Low-Density Parity-Check (LDPC) codes used in the 5th generation (5G) mobile communications. It is observed that a significant part of Variable Nodes (VNs) in the 5G LDPC codes are with degree-1 and are very sensitive to be erroneous, leading to the fact that the decoding performance is generally reduced. In the EOMS decoding, the core check nodes (CN) and extension CNs are processed with different update rules. A new CN -update criterion is also proposed by making use of the third minimum value. As a result, the offset factors are adaptively selected and the error probability of degree-1 VNs is significantly reduced. Simulation results show that the proposed EOMS decoder offers a much better error-correction performance than the state-of-the-art benchmarks for several 5G LDPC codes with a negligible complexity overhead.
Hangxuan Cui, Khoa LeTrung, Fakhreddine Ghaffari, David Declercq, Jun Lin 0001, Zhongfeng Wang 0001
APCC5
2019 A Low-latency Sparse-Winograd Accelerator for Convolutional Neural Networks
abstract
Low-latency and low-power implementations of Convolutional Neural Network (CNN) are highly desired for budget-restricted scenarios. Pruning and Winograd algorithm are two representative approaches to reduce the computation complexity of CNNs. Coupling them is very attractive, but the Winograd transformation removes data sparsity brought by pruning. In this paper, we present a low-latency sparse-Winograd CNN accelerator (LSW-CNN) for pruned Wino-grad CNN models. The ReLU-modified algorithm is employed to solve the zero refilling issue. Our design fully leverages the sparsity in both weights and activations, and thus eliminates all unnecessary computation and cycles. Moreover, a novel fast mask indexing algorithm for sparse data compression is developed. Accumulation buffers are scaled to reduce the latency brought by irregular serial channel merging. On VGG-16, experimental results demonstrate that the latency of LSW-CNN is reduced by 5.1 and 1.7 times, respectively, compared with state-of-the-art dense-Winograd and sparse-Winograd accelerators. Besides, the consumed hardware resource is also significantly reduced.
Jun Lin 0001, Zhongfeng Wang 0001
ICASSP4
2019 A New Fast-SSC-Flip Decoding of Polar Codes
abstract
Polar codes are a great breakthrough in coding theory, which have been standardized for the next generation mobile communication and are promising to improve the reliability of MLC NAND flash. The successive-cancellation (SC) decoding is low-complexity, while its error-correction performance is not satisfactory. The SC flip (SCF) decoding offers a better error-correction performance than the SC decoding while keeps a similar complexity to the SC decoding. To reduce the latency, recently the fast simplified SC (Fast-SSC) decoding is merged with the SCF, resulting in the Fast-SSC-Flip decoding. In this paper, we propose a new Fast-SSC-Flip decoding algorithm for polar codes. A novel decision LLR calculation method and bit-flipping scheme in single-parity-check (SPC) nodes are presented, leading to a better error-correction performance than the prior art. Besides, more types of special nodes in the decoding tree are considered in the proposed algorithm to reduce the decoding latency. Moreover, our algorithm can find the first erroneous bit of an invalid codeword more effectively than the prior algorithm, which can reduce the number of flipping trials and also makes contribution to a lower decoding latency.
Yangcan Zhou, Jun Lin 0001, Zhongfeng Wang 0001
ICC2
2019 TIE: energy-efficient tensor train-based inference engine for deep neural network
abstract
In the era of artificial intelligence (AI), deep neural networks (DNNs) have emerged as the most important and powerful AI technique. However, large DNN models are both storage and computation intensive, posing significant challenges for adopting DNNs in resource-constrained scenarios. Thus, model compression becomes a crucial technique to ensure wide deployment of DNNs.
Chunhua Deng, Fangxuan Sun, Xuehai Qian, Jun Lin 0001, Zhongfeng Wang 0001, Bo Yuan 0001
ISCA4
2019 A New Probabilistic Gradient Descent Bit Flipping Decoder for LDPC Codes
abstract
Probabilistic gradient descent bit-flipping (PGDBF) is the state-of-the-art hard-decision algorithm for decoding low-density parity-check (LDPC) codes on binary symmetric channel (BSC). However, there still exists a considerable performance gap between the PGDBF algorithm and soft-decision algorithms, especially in the error-floor region. To bridge this performance gap, a tabu-list aided PGDBF (T-PGDBF) algorithm is proposed in this paper. In the T-PGDBF algorithm, a tabu-list is employed to help the decoding escape from trapping sets, which is the main cause of the error-floor phenomenon. The bits which are flipped in the current iteration will be added to the tabu-list to prevent them being flipped in the next iteration. Simulation results show that the T-PGDBF algorithm offers a significant performance gain when compared to the PGDBF algorithm, which can reach that of soft-decision algorithms. We also present the hardware architecture to implement the T-PGDBF algorithm. Synthesis results show that the improved performance offered by the T-PGDBF algorithm can be obtained with a small hardware overhead.
Hangxuan Cui, Jun Lin 0001, Suwen Song, Zhongfeng Wang 0001
ISCAS2
2019 USCA: A Unified Systolic Convolution Array Architecture for Accelerating Sparse Neural Network
abstract
Due to the intensive computational complexity and various types of convolution, it is a challange to implement different CNN models on a specific hardware. Many previous works focus on data reuse and sparsity exploration to accelerate computation but fail to support various types of convolution efficiently. When dealing with variants of conventional convolution, such as deconvolution or dilated convolution, previous accelerators waste time on padding zeroes and convolving with padded feature maps. In this paper, we propose a unified convolution algorithm to intelligently combine several convolution types together and exploit the sparsity in activations. The padding process can be skipped by the proposed algorithm. Moreover, a unified systolic convolution array (USCA) architecture is developed based on the algorithm. The USCA architecture is implemented with a TSMC 28nm CMOS technology. The implementation results demonstrate that the architecture costs 206k logic gates and 114.7kB on-chip memory. It can reach a peak performance of 374.7GOPs and comsumes 201.1mW at a frequency of 1449MHz. Compared to similar works, USCA architecture achieves 3 × energy efficiency, which is measured by the number of GOPS per watt. Besides, to the best of our knowledge, USCA is the first architecture that can simultaneously support conventional convolution, deconvolution, and dilated convolution in an efficient way.
Jun Lin 0001, Zhongfeng Wang 0001
ISCAS2
2019 Methodology for Efficient Reconfigurable Architecture of Generative Neural Network
abstract
Generative neural networks have been developing rapidly in the field of deep learning nowadays. Generative models have obtained much popularity in various applications such as image generation, reading comprehension and style transfer. Convolutional (CONV) and deconvolutional (DeCONV) layers are typical components of generative neural networks. The use of traditional convolution accelerators will cause problems of overlapping and resource under-utilization while doing deconvolutions. There is little research on acceleration of deconvolution implementations. In this paper, we propose efficient reconfigurable architecture of generative neural networks. Firstly, the fast reconfigurable unit (FRU) based on cascaded fast FIR algorithm (CFFA) is proposed to support both convolutions and deconvolutions. The problems of overlapping and resource under-utilization are solved. Secondly, the reconfigurable architecture on the basis of FRUs for CONV and DeCONV layers is proposed accordingly. Thirdly, a novel shift scale quantization method is proposed to uniformly quantize CONV and DeCONV layers. Only integer computations are required with the quantization method. Finally, we choose a typical generative neural network and implement it on Xilinx Zynq ZC706. It is estimated that the performance reaches 62.85 GOPS under 330MHz working frequency on Xilinx ZC706. In brief, the proposed design outperforms existing works significantly, particularly surpasses related reconfigurable design by more than 20 times in terms of performance density.
Wendong Mao, Jichen Wang, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS3
2019 A Novel Low-Complexity Joint Coding and Decoding Algorithm for NB-LDPC Codes
abstract
Non-binary low-density parity-check (NB-LDPC) codes exhibit a much better performance than their binary counterparts, especially for moderate codeword length and high-order modulation. However, their decoding algorithms suffer from very high computational complexity. In this paper, a low-complexity algorithm is proposed, named parity-check erased algorithm (PCEA), where an additional parity check bit is added to each symbol of the codeword when encoding and a series of simple operations are performed based on these bits during decoding. As a universal joint coding and decoding algorithm, the PCEA can be combined with arbitrary NB-LDPC encoding schemes and decoding algorithms based on message passing. The proposed algorithm facilitates significant improvement of decoding performance with a small decrease of the code rate. Additionally, it usually has an even better performance than a nearly same-rate code constructed by the original method, and requires much lower decoding complexity due to smaller size of the parity check matrix.
Suwen Song, Jing Tian 0004, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS3
2019 Analysis and Design of a Large Dither Injection Circuit for Improving Linearity in Pipelined ADCs
abstract
In this paper, a new large dither injection technique is proposed for improving linearity in pipelined analog-to-digital converters (ADCs), without losing the dynamic range of the ADCs and deteriorating the corresponding amplifier's linearity. First, analyses of a proper pipelined ADC's architecture are performed for large dither injection. Then, a 9-bit capacitive digital-toanalog converter (DAC) with split architecture is developed to inject the dither ranging from -511/1024 least significant bit (LSB) to 511/1024 LSB of the first stage. To counteract the consumption of the correction range by the capacitive injection dither, the novel 6-bit complementary DACs embedded in the comparator threshold generation circuit are proposed to realize comparator dither injection. In addition, the dither injection amplitude is configurable for investigating different amplitude's effects on the linearity of the ADC. Finally, the proposed dither injection circuit, together with a 16-bit 150 million samples per second (MSPS) ADC, is implemented in a 0.18-μm CMOS technology. The measured results demonstrate the effectiveness of the proposed techniques. The optimum dither is the 9-bit dither, improving not only the spurious free dynamic range (SFDR) of the small signal by at least 13 dB but also that of the large signal by more than 8 dB compared to the case without dither injection. Moreover, dither injection makes the noise floor clean.
Congyi Zhu, Renrong Liang, Jun Lin 0001, Zhongfeng Wang 0001, Li Li 0003
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Eadnet: Efficient Architecture for Decomposed Convolutional Neural Networks
abstract
Convolutional neural networks (CNNs) are widely used in various intelligent tasks. However, the huge computational complexity of CNNs makes it hard to be implemented in many real-time embedded devices. Various methods have been employed to reduce the model size of CNNs, where the Canonical Polyadic Decomposition (CPD) has shown its capability to reduce both the computational complexity and the storage requirement with negligible accuracy loss. In this paper, an efficient configurable hardware architecture called EadNet is proposed for CPD-CNNs. In detail, to minimize the on-chip memory access, different data reuse patterns are first analyzed. Based on the chosen optimal reuse scheme, a much improved computation flow is also developed for efficiently caching activations. The EadNet is implemented with a TSMC 90nm CMOS technology. The implementation results indicate that EadNet achieves considerable improvements on computation efficiency compared to the state-of-the-art CNN accelerator architectures.
Fangxuan Sun, Jun Lin 0001, Zhongfeng Wang 0001
ICASSP2
2018 Approximate Belief Propagation Decoder for Polar Codes
abstract
Polar code is increasing its popularity recently for its capacity-achieving property for B-DMCs. However, when designing decoders for polar code, it has always been an inevitable concern for us to balance the decoding performance and the hardware consumption. In this paper, we propose an approximate belief propagation (BP) decoder for polar code for the first time. By introducing the approximate computation schemes, we reduced the critical path delay (CPD) and the hardware consumption of the conventional BP decoders. Simulation results show that the proposed approximate BP decoder achieves nearly the same decoding performance as the conventional one. Advantages of the proposed decoder has been verified by FPGA implementation.
Menghui Xu, Shusen Jing, Jun Lin 0001, Weikang Qian, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001
ICASSP3
2018 An Efficient NB-LDPC Decoding Algorithm for Next-Generation Memories
abstract
Due to the aggressive technology scaling, the memory reliability has been seriously degraded, which poses a challenge to the widely used low-density parity-check (LDPC) codes. Non-binary LDPC (NB-LDPC) codes present larger coding gain and lower error floor than their binary counterparts in many cases, which show a great potential to be used in the next-generation memories. However, the excessive computational complexity of current NB-LDPC decoding algorithms form a bottleneck and limit their applications. In this paper, a novel algorithm, called dual-threshold-based shrinking based improved trellis-based min-sum algorithm (simply TIT-MSA), is proposed to deal with this problem. The improvements include two steps. The first step is for the check node processing (CNP). Based on the CNP of the simplified min-sum algorithm (SMSA) and that of the trellis-based extended min-sum algorithm (T-EMSA), an improved trellis-based min-sum algorithm (IT-MSA) is developed, which achieves better error performance and lower computational complexity than its origins. The second step is for the whole decoding process. Based on the IT-MSA, the TIT-MSA is proposed, for which two constant thresholds are introduced to remove redundant messages by constructing two subsets of the Galois field. Simulation results show that the error performance of the TIT-MSA is nearly the same as that of the EMSA. Meanwhile, the proposed algorithm can save almost 90% computations compared to the SMSA and T-EMSA.
Jing Tian 0004, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS2
2018 A New Soft-input Hard-output decoding algorithm for Turbo Product Codes
abstract
Turbo product codes (TPCs) are being considered as a competitive forward error correction scheme for ultra high speed network communication beyond 100Gbps [1], [2]. Conventional soft-input soft-output (SISO) decoders for turbo product codes (TPCs) need to exchange extrinsic soft messages between row and column component decoders. However, the extrinsic information increases the difficulty of hardware implementation. First of all, for a two dimensional message memory, row and column component decoders access soft messages in horizontal and vertical directions, respectively, resulting in potential memory access conflict. Secondly, the computation complexity of calculating the soft extrinsic information is very high. Finally, storing the extrinsic information needs additional memory and updating it frequently increas the decoding latency and power consumption. In this paper, we propose a new soft-input hard-output (SIHO) decoding algorithm for TPCs. The proposed SIHO decoding algorithm requires the exchange of only hard information between row and column component decoders, leading to simpler decoder architectures. Moreover, each component decoder of a SIHO TPC decoder does not generate any soft information. In terms of error correction performance, the SIHO decoding algorithm is in between the SISO and the hard-input hard-output decoding algorithm. It is believed that the SIHO decoding algorithm is a good tradeoff between high net coding gain (NCG) and low complexity implementation.
Jun Lin 0001, Zhongfeng Wang 0001
ISCAS2
2018 An Efficient Convolution Core Architecture for Privacy-Preserving Deep Learning
abstract
Cloud service for deep learning (DL) has been widely used except for applications involving medical, financial, or other sensitive data due to privacy and security requirements. By employing homomorphic encryption, trained deep convolutional neural networks (CNNs) can be converted to CryptoNets, which is suitable for privacy-preserving DL cloud service. However, the high computation complexity of CryptoNets leads to tremendous implementation challenge. In this paper, to the best of our knowledge, efficient hardware acceleration of CryptoNets is discussed for the first time in open literature. In more detail, without compromise of security and inference accuracy, encryption parameters and modular multiplication algorithm are carefully selected to reduce the computation complexity of polynomial multiplication. Besides, based on the negative wrapped convolution and fast finite impulse filter schemes, an efficient algorithm for convolutions in CryptoNets is developed. Moreover, a dedicated low complexity convolution core architecture for CryptoNets is proposed and implemented with a 90nm CMOS technology. Compared to a well optimized CPU implementation of CryptoNets, this architecture is 11.9× faster while consuming a power of only 537mW.
Yizhi Wang 0003, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS2
2018 Hardware-Oriented Compression of Long Short-Term Memory for Efficient Inference
abstract
Long short-term memory (LSTM) and its variants have been widely adopted in processing sequential data. However, the intrinsic large memory requirement and high computational complexity make it hard to be employed in embedded systems. This incurs the need of model compression and dedicated hardware accelerator for LSTM. In this letter, efficient clipped gating and top-k pruning schemes are introduced to convert the dense matrix computations in LSTM into structured sparse-matrix-sparse-vector multiplications. Then, mixed quantization schemes are developed to eliminate most of the multiplications in LSTM. The proposed compression scheme is well suited for efficient hardware implementations. Experimental results show that the model size and the number of matrix operations can be reduced by 32× and 18.5×, respectively, at a cost of less than 1% accuracy loss on a word-level language modeling task.
Jun Lin 0001, Zhongfeng Wang 0001
IEEE Signal Process. Lett.2
2018 An Energy-Efficient Architecture for Binary Weight Convolutional Neural Networks
abstract
Binary weight convolutional neural networks (BCNNs) can achieve near state-of-the-art classification accuracy and have far less computation complexity compared with traditional CNNs using high-precision weights. Due to their binary weights, BCNNs are well suited for vision-based Internet-of-Things systems being sensitive to power consumption. BCNNs make it possible to achieve very high throughput with moderate power dissipation. In this paper, an energy-efficient architecture for BCNNs is proposed. It fully exploits the binary weights and other hardware-friendly characteristics of BCNNs. A judicious processing schedule is proposed so that off-chip I/O access is minimized and activations are maximally reused. To significantly reduce the critical path delay, we introduce optimized compressor trees and approximate binary multipliers with two novel compensation schemes. The latter is able to save significant hardware resource, and almost no computation accuracy is compromised. Taking advantage of error resiliency of BCNNs, an innovative approximate adder is developed, which significantly reduces the silicon area and data path delay. Thorough error analysis and extensive experimental results on several data sets show that the approximate adders in the data path cause negligible accuracy loss. Moreover, algorithmic transformations for certain layers of BCNNs and a memory-efficient quantization scheme are incorporated to further reduce the energy cost and on-chip storage requirement. Finally, the proposed BCNN hardware architecture is implemented with the SMIC 130-nm technology. The postlayout results demonstrate that our design can achieve an energy efficiency over 2.0TOp/s/W when scaled to 65 nm, which is more than two times better than the prior art.
Yizhi Wang 0003, Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2017 Efficient approximate layered LDPC decoder
abstract
Energy efficient and high throughput LDPC decoders are highly demanded, especially in the coming 5-th generation (5G) mobile communication era. In this paper, to the best of our knowledge, approximate computing units (ACUs) for the update of soft messages in row layered LDPC decoders are proposed for the first time. Under the TSMC 90nm CMOS technology, the synthesis results demonstrate that for typical LDPC codes employed in industrial standards, the corresponding ACUs achieve significant reduction in critical path delay (CPD), area and energy consumption. Numerical results show that the presented ACUs cause negligible degradation in the error-correction performance.
Yangcan Zhou, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS2
2017 Efficient Soft Cancelation Decoder Architectures for Polar Codes
abstract
The flooding belief propagation (FO-BP) and the soft-cancelation (SCAN) algorithms are the two most popular soft-output BP algorithms for the decoding of capacity-achieving polar codes. The FO-BP algorithm has high throughput at the cost of performance degradation in high signal-to-noise ratio (SNR) region or with large block length. The SCAN algorithm has much better decoding performance while suffering from long decoding latency and low throughput. In this paper, an improved BP algorithm, named reduced complexity soft-cancelation (RCSC) algorithm, is proposed. Compared with the SCAN algorithm, the number of memory entries required by the RCSC algorithm is reduced by more than 50% in general, while achieving comparable or even better (e.g., when block size N = 215) decoding performance. When block size is large (e.g., N ≥ 215), the proposed RCSC algorithm reduces the required memory entries by more than 23% compared with the state-of-the-art FO-BP algorithm. The numerical results show that the error performance improvement of the RCSC algorithm is more significant when the SNR increases. For a different tradeoff, a reduced latency soft-cancelation (RLSC) algorithm is proposed to reduce the decoding latency and increase the throughput of the RCSC algorithm while slightly sacrificing decoding performance. Finally, the optimized VLSI architectures are presented for the RCSC and RLSC algorithms, respectively. The synthesis results demonstrate the efficiency of the proposed algorithms and architectures.
Jun Lin 0001, Zhiyuan Yan 0001, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2017 Accelerating Recurrent Neural Networks: A Memory-Efficient Approach
abstract
Recurrent neural networks (RNNs) have achieved the state-of-the-art performance on various sequence learning tasks due to their powerful sequence modeling capability. However, RNNs usually require a large number of parameters and high computational complexity. Hence, it is quite challenging to implement complex RNNs on embedded devices with stringent memory and latency requirement. In this paper, we first present a novel hybrid compression method for a widely used RNN variant, long-short term memory (LSTM), to tackle these implementation challenges. By properly using circulant matrices, forward nonlinear function approximation, and efficient quantization schemes with a retrain-based training strategy, the proposed compression method can reduce more than 95% of memory usage with negligible accuracy loss when verified under language modeling and speech recognition tasks. An efficient scalable parallel hardware architecture is then proposed for the compressed LSTM. With an innovative chessboard division method for matrix-vector multiplications, the parallelism of the proposed hardware architecture can be freely chosen under certain latency requirement. Specifically, for the circulant matrix-vector multiplications employed in the compressed LSTM, the circulant matrices are judiciously reorganized to fit in with the chessboard division and minimize the number of memory accesses required for the matrix multiplications. The proposed architecture is modeled using register transfer language (RTL) and synthesized under the TSMC 90-nm CMOS technology. With 518.5-kB on-chip memory, we are able to process a 512×512 compressed LSTM in 1.71 μs, corresponding to 2.46 TOPS on the uncompressed one, at a cost of 30.77-mm2chip area. The implementation results demonstrate that the proposed design can achieve significantly high flexibility and area efficiency, which satisfies many real-time applications on embedded devices. It is worth mentioning that the memory-efficient approach of accelerating LSTM developed in this paper is also applicable to other RNN variants.
Jun Lin 0001, Zhongfeng Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Error performance analysis of the symbol-decision SC polar decoder
abstract
Polar codes are the first provably capacity-achieving forward error correction codes. To improve decoder throughput, the symbol-decision SC algorithm makes hard-decision for multiple bits at a time. In this paper, we prove that for polar codes, the symbol-decision SC algorithm is better than the bit-decision SC algorithm in terms of the frame error rate (FER) performance because the symbol-decision SC algorithm performs a local maximum likelihood decoding within a symbol. Moreover, the bigger the symbol size, the better the FER performance. Finally, simulation results over both the additive white Gaussian noise channel and the binary erasure channel confirm our theoretical analysis.
Chenrong Xiong, Jun Lin 0001, Zhiyuan Yan 0001
ICASSP2
2016 Accurate runtime thermal prediction scheme for 3D NoC systems with noisy thermal sensors
abstract
Thermal sensor noise has great impact on the efficiency and effectiveness of a dynamic thermal management (DTM) strategy. Conventional reactive thermal management techniques suffer significant performance degradation due to the pessimistic reaction. In this paper, to address the problem of forecasting temperatures based on noisy thermal readings, we propose a Kalman predictor based runtime thermal prediction scheme, which can predict temperatures N step ahead. An activity-based power model for 3D NoC power estimation is also proposed; the model is an essential prerequisite of accurate temperature predictions. Besides that, we propose a distributed multi-input single-output (MISO) thermal model for 3D NoC systems, which reduces the computational complexity of temperature updating from m2 to m compared with the centralized multi-input multi-output (MIMO) model for the system with m units. The experimental results show that the proposed prediction scheme reduces the mean absolute error (MAE) by 42.8%-72.6% compared with the auto-regressive (AR) based prediction scheme.
Li Li 0003, Hongbing Pan, Kun Wang 0005, Feng Han 0008, Jun Lin 0001
ISCAS6
2016 A high throughput belief propagation decoder architecture for polar codes
abstract
The belief propagation (BP) decoding algorithm not only is an alternative to the successive cancelation (SC) decoders of polar codes, but also provides soft outputs that are necessary for joint detection and decoding. The BP decoders with the flooding schedule achieve high throughput with excessive hardware cost especially when the block length is large. The soft-cancelation (SCAN) decoders for polar codes have reduced memory complexity compared to the BP decoders based on the flooding schedule. The simplified SC aided reduced complexity soft-cancelation (S-RCSC) decoders further reduce the computational and memory complexity of the SCAN decoders at the cost of negligible error performance degradation. Both the SCAN and S-RCSC decoders have limited throughput due to their serial decoding schedules. In this paper, we first propose an improved S-RCSC (IS-RCSC) decoding algorithm and then present a high throughput decoder architecture based on our IS-RCSC algorithm. Our IS-RCSC decoding algorithm performs the message passing on a binary tree representation of a polar code. Compared to the S-RCSC decoding algorithm, our IS-RCSC decoding algorithm accelerates the computing of the returned soft messages when certain types of nodes are activated. The corresponding hardware architecture of our IS-RCSC decoder is also proposed. In terms of area efficiency, the hardware implementation results demonstrate that our IS-RCSC decoders are 19% to 43% better than decoders in the literature.
Jun Lin 0001, Jin Sha 0001, Li Li 0003, Chenrong Xiong, Zhiyuan Yan 0001, Zhongfeng Wang 0001
ISCAS1
2016 Stage-combined belief propagation decoding of polar codes
abstract
A novel modification is introduced in this paper for the belief propagation decoder of polar codes, wherein adjacent two processing stages are efficiently combined together to speed up decoding. Corresponding path based belief estimation method is presented in detail. The proposed decoder halves the number of stages of the conventional decoder and thus can significantly reduce the decoding latency and lower message memory requirement.
Jin Sha 0001, Jun Lin 0001, Zhongfeng Wang 0001
ISCAS2
2016 A High Throughput List Decoder Architecture for Polar Codes
abstract
While long polar codes can achieve the capacity of arbitrary binary-input discrete memoryless channels when decoded by a low complexity successive-cancellation (SC) algorithm, the error performance of the SC algorithm is inferior for polar codes with finite block lengths. The cyclic redundancy check (CRC)-aided SC list (SCL) decoding algorithm has better error performance than the SC algorithm. However, current CRC-aided SCL decoders still suffer from long decoding latency and limited throughput. In this paper, a reduced latency list decoding (RLLD) algorithm for polar codes is proposed. Our RLLD algorithm performs the list decoding on a binary tree, whose leaves correspond to the bits of a polar code. In existing SCL decoding algorithms, all the nodes in the tree are traversed, and all possibilities of the information bits are considered. Instead, our RLLD algorithm visits much fewer nodes in the tree and considers fewer possibilities of the information bits. When configured properly, our RLLD algorithm significantly reduces the decoding latency and, hence, improves throughput, while introducing little performance degradation. Based on our RLLD algorithm, we also propose a high throughput list decoder architecture, which is suitable for larger block lengths due to its scalable partial sum computation unit. Our decoder architecture has been implemented for different block lengths and list sizes using the TSMC 90-nm CMOS technology. The implementation results demonstrate that our decoders achieve significant latency reduction and area efficiency improvement compared with the other list polar decoders in the literature.
Jun Lin 0001, Chenrong Xiong, Zhiyuan Yan 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2016 A Multimode Area-Efficient SCL Polar Decoder
abstract
Polar codes are of great interest, since they are the first provably capacity-achieving forward error correction codes. To improve throughput and to reduce decoding latency of polar decoders, maximum likelihood (ML) decoding units are used by successive cancellation list (SCL) decoders as well as SC decoders. This paper proposes an approximate ML (AML) decoding unit for SCL decoders first. In particular, we investigate the distribution of frozen bits of polar codes designed for both the binary erasure and additive white Gaussian noise channels, and take advantage of the distribution to reduce the complexity of the AML decoding unit, improving the throughput-area efficiency of the SCL decoders. Furthermore, a multimode (MM) SCL decoder with variable list sizes and parallelism is proposed. If high throughput or small latency is required, the decoder decodes multiple received words in parallel with a small list size. However, if error performance is of higher priority, the MM-SCL decoder switches to a serial mode with a bigger list size. Therefore, the MM-SCL decoder provides a flexible tradeoff between latency, throughput, and error performance at the expense of small overhead. Hardware implementation and synthesis results show that our polar decoders not only have a better throughput-area efficiency but also easily adapt to different communication channels and applications.
Chenrong Xiong, Jun Lin 0001, Zhiyuan Yan 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2015 A hybrid partial sum computation unit architecture for list decoders of polar codes
abstract
Although the successive cancelation (SC) algorithm works well for very long polar codes, its error performance for shorter polar codes is much worse. Several SC based list decoding algorithms have been proposed to improve the error performances of both long and short polar codes. A significant step of SC based list decoding algorithms is the updating of partial sums for all decoding paths. In this paper, we first proposed a lazy copy partial sum computation algorithm for SC based list decoding algorithms. Instead of copying partial sums directly, our lazy copy algorithm copies indices of partial sums. Based on our lazy copy algorithm, we propose a hybrid partial sum computation unit architecture, which employs both registers and memories so that the overall area efficiency is improved. Compared with a recent partial sum computation unit for list decoders, when the list size L = 4, our partial sum computation unit achieves an area saving of 23% and 63% for block length 213and 215, respectively.
Jun Lin 0001, Zhiyuan Yan 0001
ICASSP1
2015 An Efficient List Decoder Architecture for Polar Codes
abstract
Long polar codes can achieve the symmetric capacity of arbitrary binary-input discrete memoryless channels under a low-complexity successive cancelation (SC) decoding algorithm. However, for polar codes with short and moderate code lengths, the decoding performance of the SC algorithm is inferior. The cyclic-redundancy-check (CRC)-aided SC-list (SCL)-decoding algorithm has better error performance than the SC algorithm for short or moderate polar codes. In this paper, we propose an efficient list decoder architecture for the CRC-aided SCL algorithm, based on both algorithmic reformulations and architectural techniques. In particular, an area efficient message memory architecture is proposed to reduce the area of the proposed decoder architecture. An efficient path pruning unit suitable for large list size is also proposed. For a polar code of length 1024 and rate 1/2, when list size L=2 and 4, the proposed list decoder architecture is implemented under a Taiwan Semiconductor Manufacturing Company (TSMC) 90-nm CMOS technology. Compared with the list decoders in the literature, our decoder achieves 1.24-1.83 times the area efficiency.
Jun Lin 0001, Zhiyuan Yan 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2014 Efficient list decoder architecture for polar codes
abstract
Long polar codes achieve the capacity of binary-input discrete memoryless channels when decoded with a successive cancelation (SC) algorithm. For polar codes with short or moderate length, the decoding performance of the SC algorithm is inferior, and the cyclic redundancy check (CRC) aided successive cancelation list (SCL) algorithm achieves significantly improved performance. In this paper, we propose an efficient list decoder architecture for the CRC aided SCL algorithm. Three list decoders with list size L = 2, 4 and 8, respectively, are implemented with a 90nm CMOS technology. Compared to list decoders with L = 2 and 4 in the literature, the proposed list decoders achieve 1.42 and 2.84 times, respectively, higher hardware efficiency. The implementation with list size L = 8 demonstrates that our decoder architecture works for large list sizes.
Jun Lin 0001, Zhiyuan Yan 0001
ISCAS1
2014 An Efficient Fully Parallel Decoder Architecture for Nonbinary LDPC Codes
abstract
Nonbinary low-density parity-check (NB-LDPC) codes outperform their binary counterparts in some cases, but their high decoding complexity is a significant hurdle to their applications. In this paper, we propose a decoding algorithm with reduced computational complexities and smaller memory requirements for NB-LDPC codes. First, a simplified algorithm is proposed to reduce the computational complexity of variable node processing. To reduce the memory requirements, existing NB-LDPC decoders often truncate the message vectors to a limited number nmof values. However, the memory requirements of these decoders remain high when the field size is large. In this paper, an improved trellis-based check node processing algorithm is proposed to significantly reduce the memory requirement. The number of elements in a variable-to-check message is reduced to nv(nvm). The sorted log likelihood ratio (LLR) vector of a check-to-variable (c-to-v) message is approximated using a piecewise linear function. For each a priori message, most of the LLRs are approximated with a linear function. Two low complexity LLR generation units (LGUs) are proposed to compute LLR vectors for c-to-v messages. A fully parallel NB-LDPC decoder over GF(256) is implemented with 28-nm CMOS technology. The decoder over GF(256) achieves a throughput of 546 Mb/s and an energy efficiency of 0.178 nJ/b/iter.
Jun Lin 0001, Zhiyuan Yan 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2013 A decoding algorithm with reduced complexity for non-binary LDPC codes over large fields
abstract
Non-binary low-density parity-check (NB-LDPC) codes outperform their binary counterparts in some cases, but their high decoding complexity is a significant hurdle to their applications. In this paper, we propose a decoding algorithm with reduced computational complexities and smaller memory requirements for NB-LDPC codes over large fields. First, a simplified algorithm is proposed to reduce the computational complexity of variable node processing. To reduce memory requirements, existing NB-LDPC decoders often truncate the message vectors to a limited number nmof values. However, the memory requirements of these decoders remain high when the field size is large, since nmneeds to be large enough to alleviate error performance degradation. In this paper, an improved trellised based check node processing algorithm is proposed to significantly reduce the memory requirement. The number of elements in a variable-to-check message is reduced to nv(nvm). The sorted log likelihood ratio (LLR) vector of a check-to-variable message is approximated using a piece-wise linear function. Thus, only few LLRs are stored and other LLRs are computed on-the-fly when needed. For each a priori message, most LLRs are approximated with a linear function. Our numerical results demonstrate that the proposed decoding algorithm outperforms existing algorithms. Two LLR generation units (LGUs) are proposed to compute LLR vectors for check-to-variable messages, and the two LGUs require only a fraction of the area needed to store nmLLRs.
Jun Lin 0001, Zhiyuan Yan 0001
ISCAS1
2013 Efficient Shuffled Decoder Architecture for Nonbinary Quasi-Cyclic LDPC Codes
abstract
In this brief, a shuffled schedule (SS) of the min-max decoding algorithm is proposed for nonbinary low-density parity-check (LDPC) codes. To increase the throughput and reduce the memory requirement, a modified SS (MSS) with much simpler check node processing is also proposed, based on a new shuffled merge algorithm. Numerical simulations for three LDPC codes with different lengths and rates over GF(32) show: 1) both SS and MSS converge faster and have slightly better error performance than the flooding schedule, and 2) the degradation of the MSS in error performance as well as convergence rate is negligible. Finally, an efficient decoder architecture based on the MSS is proposed for quasi-cyclic LDPC codes. The proposed decoder architecture further enhances the decoding throughput with improved check and variable node processing units. The implementation results of an (837, 726) LDPC decoder over GF(32) demonstrate that the proposed architecture outperforms those in previous works.
Jun Lin 0001, Zhiyuan Yan 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2012 Modified shuffled schedule for nonbinary low-density parity-check codes
abstract
In this paper, a shuffled schedule of the Min-Max decoding algorithm is proposed for nonbinary LDPC codes. To reduce the latency and memory requirement, a modified shuffled schedule with much simpler check node processing is also proposed, based on a new shuffle sort algorithm. Numerical simulations for three LDPC codes over GF(32) with different lengths and rates show: (1) both the shuffled and modified shuffled schedule converge faster and have slightly better error performance than the flooding schedule, and (2) the degradation of the modified shuffled schedule in error performance as well as convergence rate is negligible.
Jun Lin 0001, Zhiyuan Yan 0001
ISCAS1
2009 LDPC Decoder Design for IEEE 802.15 Standard
abstract
This paper presents an efficient decoder design for the LDPC codes in IEEE 802.15 standard. This decoder features by high parallel level, low message memory requirement and code rate flexibility. By processing 72 columns and 72 rows in parallel, it can reach a throughput of 2.8 Gbps to fulfill the standard requirement. Furthermore, the decoder supports three different code rates by employing flexible check node processor units.
Jin Sha 0001, Jun Lin 0001, Li Li 0003, Minglun Gao, Zhongfeng Wang 0001
ISCAS2