EDBT 2026 Demo / reviewers in the wild / expert
Jing Tian 0004
dblp:69/4394-4
· DBLP profile ↗
21ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-2402-523XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 5 first-author · 18 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Efficient Sine/Cosine Design Using Piecewise Quadratic Approximation for FOC Applications
Yuan Gao 0007, Yuan Cao 0003, Jianjun Zhuang, Rongkai Pan, Jing Tian 0004 |
ISCAS | 7 |
| 2026 | Fast Scloud+: A High-Speed Hardware Implementation for Unstructured-LWE-Based Post-Quantum Cryptography
Jing Tian 0004, Yaodong Wei, Dejun Xu, Anyu Wang 0001, Zhiyuan Qiu, Fu Yao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | An Improved Two-Step Attack on Lattice-Based Cryptography: A Case Study of KyberabstractAfter three rounds of post-quantum cryptography (PQC) strict evaluations conducted by NIST, CRYSTALS-Kyber was successfully selected in July 2022 and standardized in August 2024. It becomes urgent to further evaluate Kyber’s physical security for the upcoming deployment phase. In this brief, we present an improved two-step attack on Kyber to quickly recover the full secret key, s, by using much fewer power traces and less time. In the first step, we use the correlation power analysis (CPA) to obtain a portion of guess values of s with a small number of power traces. The CPA is enhanced by utilizing both Pearson and Kendall’s rank correlation coefficients and modifying the leakage model to improve the accuracy. In the second step, we adopt the lattice attack to recover s based on the results of CPA. The success rate is largely built up by constructing a trial-and-error method. We deploy the reference implementations of Kyber-512, -768, and -1024 on an ARM Cortex-M4 target board and successfully recover s in approximately$9\sim 10$min with at most 15 power traces, using a Xeon Gold 6342-equipped machine for the attack. Dejun Xu, Jing Tian 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Fast Hardware Architecture With Efficient Matrix Computations for the Key Generation of Classic McElieceabstractClassic McEliece, with a remarkably stable security level, has been selected as one of the four key-establishment algorithms in the fourth-round evaluation of the post-quantum cryptography (PQC) standardization process of national institute of standards and technology (NIST). However, its memory-intensive and time-consuming key generation poses an obstacle to widespread use. In this paper, we propose a fast hardware implementation of the key generation incorporating several architectural optimizations. For the Gaussian elimination, we optimize the scheduling of computing resources and the memory access process and present a high-performance and flexible systemizer with multiple low fan-out systolic arrays. Besides, an algorithmic-level parallelized design for entry generation and Gaussian elimination is proposed to reduce the redundant computation time. A compact entry generator with a multi-level feedback mechanism and a 2-D high-speed FFT module facilitates continuous streaming the generated entries into the systemizer.FPGA implementation results show that our designs for the key generation improve time-area efficiency by 11.9% to 43.2% compared to the state-of-the-arts. Moreover, compared to the hardware implementations for the key generation of the other two quasi-cyclic code-based PQC algorithms, ours for Classic McEliece based on the random code achieves close to or better results in several metrics. Xinyuan Qiao, Jing Tian 0004, Suwen Song, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | HRCIM-NTT: An Efficient Compute-in-Memory NTT Accelerator With Hybrid-Redundant NumbersabstractRecently, four NIST-approved Post-Quantum Cryptography (PQC) algorithms are selected to be standardized. Three of them are lattice-based cryptographic schemes and feature the number-theoretic transform (NTT) as the computing bottleneck compelling fast and low-power hardware implementations. In this work, a high-speed and power-efficient NTT accelerator is presented leveraging the compute-in-memory (CIM) technique with bottom-up optimizations. Firstly, a carry-free modular multiplication (CFMM) algorithm is proposed, which utilizes on-the-fly reduction and hybrid-redundant representation to optimize the butterfly unit operation, the cornerstone of NTT. Based on the optimized algorithm, an efficient butterfly unit in memory (BUIM) is developed by co-designing with SRAM circuit, which saves the memory access energy, decreases operation cycles, and obtains ultra-short critical path. Additionally, the data pattern of CIM array is also improved to avoid redundant memory read/write operations, which further reduces memory access overhead. Finally, a combination of pipelined operation flow and constant interstage data mapping strategy is employed to bestow the proposed hybrid-redundant CIM NTT (HRCIM-NTT) architecture with minimized computing cycles and reduced routing overhead. The implementation under 45nm CMOS technology demonstrates that HRCIM-NTT achieves the highest throughput and lowest latency among the existing CIM-based NTT accelerators. Xu Zhang 0040, Yaodong Wei, Minghao Li 0001, Jing Tian 0004, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Efficient FPGA-Based Accelerator of the L-BFGS Algorithm for IoT ApplicationsabstractThe Internet of Things (IoT)-centric applications, such as augmented reality and self-driven cars, require real-time task processing, large bandwidth, and low data transmission latency. FPGA-based edge computing is considered an effective solution to tackle these challenges. As an excellent tool in these applications, nonlinear optimization methods involve computation-intensive and data-dependency operations leading to limited real-time applications. The limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm ranks among the most efficient algorithms for large-scale optimization problems. In this paper, we propose, for the first time, a high-parallel FPGA-based architecture for the two key parts of the L-BFGS algorithm: the search direction computation and line searching. Compared with the implementation on CPU, the search direction computation and line searching implementation on FPGA achieve$\mathbf{39.73}\times$and$\mathbf{5.50}\times$speedups, respectively. Compared with the straightforward implementation on GPU, the search direction computation on FPGA obtains a speedup of$\mathbf{31.03}\times$. Huiyang Xiong, Bohang Xiong, Jing Tian 0004, Hao Zhu 0004, Zhongfeng Wang 0001 |
ISCAS | 4 |
| 2023 | Low-latency Hardware Architecture for VDF Evaluation in Class GroupsabstractThe verifiable delay function (VDF), as a kind of cryptographic primitives, has recently been adopted quite often in decentralized systems. Highly correlated to the security of VDFs, the fastest implementation for VDF evaluation is generally desired to be publicly known. In this paper, for the first time, we propose a low-latency hardware implementation for the complete VDF evaluation in the class group by jointly exploiting optimizations. On one side, we reduce the required computational cycles by decreasing the hardware-unfriendly divisions and increase the parallelism of computations by reducing the data dependency. On the other side, we provide low-latency large-number divisors, multipliers, and adders, respectively, while those operators are generally very hard to be accelerated. Besides, we carefully schedule the sub-modules and devise the low-latency architecture for the complete VDF evaluation. Finally, the proposed design is coded and synthesized under the TSMC 28-nm CMOS technology. The experimental results show that our design can achieve a speedup of 3.5x compared to the optimal C++ implementation for the VDF evaluation over an advanced CPU. Moreover, compared to the state-of-the-art hardware implementation for the squaring, a key step of VDF, we achieve about 2x speedup. Danyang Zhu, Jing Tian 0004, Minghao Li 0001, Zhongfeng Wang 0001 |
IEEE Trans. Computers | 2 |
| 2023 | Reconfigurable and High-Efficiency Polynomial Multiplication Accelerator for CRYSTALS-KyberabstractRecently, the National Institute of Standards and Technology (NIST) has identified the first four quantum-resistant algorithms for post-quantum cryptography (PQC) standardization. CRYSTALS-Kyber (Kyber) is the only public-key encryption and key-establishment algorithm among them. In this article, we propose a reconfigurable, high-speed, and area-efficient polynomial multiplication accelerator for Kyber to facilitate its practical applications. The cornerstone of polynomial multiplication is the butterfly unit (BU) structure, composed of modular addition, subtraction, and multiplication. For the modular multiplication, we adopt the Barrett reduction method and reduce the size of operands leveraging the form of modulus with a novel formula transformation, which significantly decreases the computational complexity and increases the maximum clock frequency. On the hardware side, we make four BU modules constitute a binomial arithmetic core (Bi-Core) as the basic reconfigurable unit. The memory access scheme tailored for parallel processing is explored with data-reusing and memory-grouping methods, and a compact control logic is devised. The complete polynomial multiplication architecture is coded with Verilog and implemented on a Xilinx Artix-7 xc7a100t-3 device. Experiment results demonstrate that our implementations with different configurations all outperform the state-of-the-art works in area efficiency by up to 39% improvement in terms of area-time product (ATP). Moreover, the proposed design with four Bi-Cores achieves the fastest speed among existing designs. Minghao Li 0001, Jing Tian 0004, Xiao Hu 0007, Zhongfeng Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | AC-PM: An Area-Efficient and Configurable Polynomial Multiplier for Lattice Based CryptographyabstractAs the computation bottleneck in lattice-based cryptography (LBC), the polynomial multiplication based on number theoretic transform (NTT) has been continuously studied for flexible hardware implementations with high area-efficiency. This paper presents an area-efficient and configurable NTT-based polynomial multiplier (AC-PM) incorporating algorithmic and architectural level optimization techniques. For the core operation of polynomial multiplication, two low-complexity and fast modular multiplication algorithms are introduced with loose constraints of LBC-friendly primes. Based on the proposed algorithms, a reconfigurable processing element (RPE) is dedicatedly designed to execute all the operations in an NTT-based polynomial multiplication: NTT, inverse NTT (INTT), and coefficient-wise multiplication (CWM). The proposed AC-PM can be configured with different numbers of RPEs and supports various polynomial degrees without recompilation. Additionally, the dataflow complexity is greatly simplified. More importantly, to the best of our knowledge, the twiddle factors are reused, for the first time, to support both NTT and INTT with multiple polynomial degrees, which leads to increased flexibility of AC-PM with small overhead on hardware resource. FPGA implementation results demonstrate that the proposed AC-PM significantly outperforms the prior arts in both flexibility and area efficiency. Xiao Hu 0007, Jing Tian 0004, Minghao Li 0001, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | A High-Speed FPGA-Based Hardware Implementation for Leighton-Micali SignatureabstractDue to the rapid progress made in quantum computers, modern cryptography faces great challenges. Many digital signature schemes that have resistance to quantum computing are studied and standardized by several influential international organizations. The Leighton-Micali signature (LMS) protocol, one of the hash-based signature schemes, is standardized by both the Internet Engineering Task Force (IETF) and the National Institute of Standards and Technology (NIST) due to its well-studied security and relatively small signature size. However, the heavy computation load and high latency of LMS limits its practical applications. In this paper, for the first time, we propose a full hardware implementation of LMS to accelerate all the three stages:$key~generation$,$signature~generation$, and$verification$. Considering the scalability requirement and the characteristic of the parameter sets of LMS, we extract the coarse-grained basic logic, a hash group, and build a reconfigurable architecture for all available parameters by carefully designing the parallelism degree while achieving low latency and high hardware utilization efficiency. Then, we devise a fusion architecture for$key~generation$and$signature~generation$based on the hash group module. Moreover, for the$signature~verification$stage, we propose a separate architecture by applying the hash group module along with an efficient depth-first Merkle tree module. We code our designs with Verilog language in parameterized style and implement them on a Xilinx XCVU7P FPGA platform. The experimental results show that significant improvements are obtained for different parameter sets by the proposed designs when compared to state-of-the-art works. Yifeng Song, Xiao Hu 0007, Jing Tian 0004, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | A High-Speed Codec Architecture for Lagrange Coded ComputingabstractThe Lagrange Coded Computing (LCC), proposed recently by Yu et at., is regarded as a promising solution for most distributed learning algorithms thanks to its good tradeoff between resiliency, security, and privacy over cloud servers. As a kind of coded computing, LCC also costs extra computations in a local computer for encoding and decoding, which contains many complex operations, such as the continued product operations and divisions. In this paper, we present an efficient high-speed LCC codec architecture based on the linear regression algorithm for the first time. By analyzing the formulas and evaluating the hardware resource, we select a set of optimal parameters and remove most of the complex operations by storing the precomputed coefficients. Besides, the proposed architecture is inherently scalable and can be fully utilized and reused for encoding and decoding. The experimental results on an FPGA show that a significant speedup is achieved compared with the prior art. Bohang Xiong, Jing Tian 0004, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2022 | Reduction-Free Multiplication for Finite Fields and Polynomial Rings
Samira Carolina Oliva Madrigal, Gökay Saldamli, Yue Geng, Jing Tian 0004, Zhongfeng Wang 0001, Çetin Kaya Koç |
WAIFI | 5 |
| 2022 | Efficient Software Implementation of the SIKE Protocol Using a New Data RepresentationabstractThanks to relatively small public and secret keys, the Supersingular Isogeny Key Encapsulation (SIKE) protocol made it into the third evaluation round of the post-quantum standardization project of the National Institute of Standards and Technology (NIST). Even though a large body of research has been devoted to the efficient implementation of SIKE, its latency is still undesirably long for many real-world applications. Most existing implementations of the SIKE protocol use the Montgomery representation for the underlying field arithmetic since the corresponding reduction algorithm is considered the fastest method for performing multiple-precision modular reduction. In this paper, we propose a new data representation for supersingular isogeny-based Elliptic-Curve Cryptography (ECC), of which SIKE is a sub-class. This new representation enables significantly faster implementations of modular reduction than the Montgomery reduction, and also other finite-field arithmetic operations used in ECC can benefit from our data representation. We implemented all arithmetic operations in C using the proposed representation such that they have constant execution time and integrated them to the latest version of the SIKE software library. Using four different parameters sets, we benchmarked our design and the optimized generic implementation on a 2.6 GHz Intel Xeon E5-2690 processor. Our results show that, for the prime of SIKEp751, the proposed reduction algorithm is approximately 2.61 times faster than the currently best implementation of Montgomery reduction, and our representation also enables significantly better timings for other finite-field operations. Due to these improvements, we were able to achieve a speed-up by a factor of about 1.65, 2.03, 1.61, and 1.48 for SIKEp751, SIKEp610, SIKEp503, and SIKEp434, respectively, compared to state-of-the-art generic implementations. Jing Tian 0004, Piaoyang Wang, Zhe Liu 0001, Jun Lin 0001, Zhongfeng Wang 0001, Johann Großschädl |
IEEE Trans. Computers | 1 |
| 2022 | Efficient Homomorphic Convolution Designs on FPGA for Secure InferenceabstractRecently, secure neural network (NN) inference, a combination of homomorphic encryption (HE) and NN, has attracted much attention. Nevertheless, a large number of computations, mainly brought by the HE scheme, form the bottleneck in real-time applications. In this article, we present a hardware accelerator on a field-programmable gate array (FPGA) for the homomorphic convolution layer (HomConvL), which is the most computation-intensive part of the HE-based secure inference. First, we propose a new HomConvL algorithm called packed rotations at inputs (PaRotI), which is suitable for hardware implementation for its inherent high parallelism and low complexity with acceptable noise growth and moderate resource consumption. Then, we present three highly parallel architectures for different parameter sets and application scenarios of state-of-the-art HomConvL algorithms. The new architectures are implemented on a Xilinx VCU110 FPGA board, and the experimental results demonstrate that our designs can achieve 15.31–$19.46\times $speedups compared with the software implementations. Xiao Hu 0007, Minghao Li 0001, Jing Tian 0004, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | DARM: A Low-Complexity and Fast Modular Multiplier for Lattice-Based CryptographyabstractThe lattice-based cryptography (LBC) has been widely used recently in many compute-intensive applications, such as the post-quantum cryptography (PQC) and privacy-preserving deep learning, where the main task for such applications is to improve the computational efficiency. The modular multiplication operations, mainly involved in the number theoretic transform (NTT), comprise a large proportion of the whole computations required by an LBC. This paper presents a novel "decompose-and-reduce" modular multiplication algorithm (DARM), considering primes with the form of q = 22N−δ and δN−2. The inherent structure of the modulus is exploited and the intermediates’ data widths are reduced. Moreover, a low-complexity and fast multiplier is elaborately devised based on DARM. To further validate the performance of our multiplier, an n-point NTT design with DARM is implemented with various configurations. FPGA implementation results demonstrate that compared with the prior arts, the proposed multiplier has 1.12-1.89× speedups with the least DSP utilization. For the case of ⌈log2q⌉ = 60 and n = 4096, the NTT implementation with DARM achieves up to 41.2% and 61.2% reductions in LUTs and DSPs, respectively. Xiao Hu 0007, Minghao Li 0001, Jing Tian 0004, Zhongfeng Wang 0001 |
ASAP | 3 |
| 2021 | High-Speed and Scalable FPGA Implementation of the Key Generation for the Leighton-Micali Signature ProtocolabstractDue to the rapid progress made in quantum computers, modern cryptography faces great challenges. Many new digital signature schemes that have resistance to quantum computing are being presented for Post-Quantum Cryptography (PQC) standardization. The Leighton-Micali signature (LMS), a kind of hash-based signature scheme, is selected as a promising candidate for the PQC signature protocols by the Internet Engineering Task Force (IETF) because of its small private and public key sizes. However, the low-efficiency in key generation forms the bottleneck in practical applications. In this paper, we propose a high-speed architecture for the key generation to accelerate the LMS for the first time. The architecture is delicately devised to be scalable, supporting all the parameter sets for the LMS. The degree of parallelism is carefully designed to achieve low latency and high hardware utilization efficiency. Moreover, the control flow is well managed to accommodate different parameter sets with constant power for the consideration of anti-power analysis attacks. We code our design with Verilog language and implement it on the Xilinx Zynq UltraScale+ FPGA. The experimental results show that, compared with the optimal software implementation running on an Intel(R) Core(TM) i7-6850K 3.60GHz CPU with threading enabled, the new design achieves 55x to 2091x speedup in different parameter configurations. Yifeng Song, Xiao Hu 0007, Jing Tian 0004, Zhongfeng Wang 0001 |
ISCAS | 4 |
| 2021 | Low-Latency Architecture for the Parallel Extended GCD Algorithm of Large NumbersabstractThe extended Greatest Common Divisor (GCD) is an extension of the GCD operation, which computes not only the GCD of integers a and b but also the Bezout's coefficients that are integers x and y such that ax + by =3D GCD(a,b). Recently, the large-number extended GCD algorithm is used in the core function of the next-generation blockchain systems and served as the most time-consuming operation. Considering the efficiency, speeding up this operation is urgently desired. However, the extended GCD, which is rarely explored in literature, is extremely hard to parallelize because of long serial operations with strong data dependency. In this paper, we propose a low- latency architecture for the extended GCD of large numbers by utilizing many algorithmic transformations and architectural optimizations. Firstly, a parallel extended GCD algorithm is well studied and modified to be practical in hardware. Secondly, a high-parallel architecture is designed for the selected extended GCD, where the trade-off is well evaluated between computation latency and power consumption. Finally, the architecture is coded using Verilog language and synthesized under the TSMC 28- nm CMOS technology. The experimental results for the 1024-bit extended GCD show that our design significantly outperforms the prior arts. Danyang Zhu, Jing Tian 0004, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2021 | High-Speed FPGA Implementation of SIKE Based on an Ultra-Low-Latency Modular MultiplierabstractThe supersingular isogeny key encapsulation (SIKE) protocol, as one of the post-quantum protocol candidates, is widely regarded as the best alternative for curve-based cryptography. However, the long latency, caused by the serial large-degree isogeny computation which is dominated by modular multiplications, has made it less competitive than most popular post-quantum candidates. In this paper, we propose a high-speed and low-latency architecture for our recently presented optimized SIKE algorithm. Firstly, we design a new field arithmetic logic unit (FALU) with many algorithmic transformations and architectural optimizations. Especially, for the FALU, an extremely low-latency modular multiplier is devised based on a modified algorithm by fully parallelizing and highly optimizing the small-size multipliers and the reduction submodules. Secondly, we develop a compact control logic and update the instructions based on the benchmark provided in the newest SIKE library, fitting well with our design. Thirdly, an efficient memory access method is proposed by scheduling the input and output of the arithmetic logic unit (ALU) in two identical RAMs, which can significantly reduce the latency. Finally, we code the proposed architectures using the Verilog language and integrate them into the SIKE library. The implementation results on a Xilinx Virtex-7 FPGA show that for SIKEp751, our design only costs 9.3 ms with a frequency of 155.8 MHz, about 2× faster than the state-of-the-art, and achieves the best area efficiency among existing works. Particularly, the modular multiplier merely needs 16 clock cycles, reducing the delay by nearly one order of magnitude with a small factor of increase in hardware resource. Jing Tian 0004, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Fast Modular Multipliers for Supersingular Isogeny-Based Post-Quantum CryptographyabstractAs one of the postquantum protocol candidates, the supersingular isogeny key encapsulation (SIKE) protocol delivers promising public and secret key sizes over other candidates. Nevertheless, the considerable computations form the bottleneck and limit its practical applications. The modular multiplication operations occupy a large proportion of the overall computations required by the SIKE protocol. The VLSI implementation of the high-speed modular multiplier remains a big challenge. In this article, we propose three improved modular multiplication algorithms based on an unconventional radix for this protocol, all of which cost about 20% fewer computations than the prior art. Besides, a multiprecision scheme is also introduced for the proposed algorithms to improve the scalability in hardware implementation, resulting in three new algorithms. We then present very efficient high-speed constant-time modular multiplier architectures for the six algorithms. It is shown that these new architectures can be extensively pipelined and highly optimized to obtain high throughput and low latency. The field-programmable gate array (FPGA) implementation results show that all proposed multipliers achieve much higher throughput than previous designs, but the increase in resources is relatively small. In addition, the multipliers without the multiprecision scheme have very low latency, which is very friendly to high-speed applications of the SIKE protocol. Jing Tian 0004, Jun Lin 0001, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | A Novel Low-Complexity Joint Coding and Decoding Algorithm for NB-LDPC CodesabstractNon-binary low-density parity-check (NB-LDPC) codes exhibit a much better performance than their binary counterparts, especially for moderate codeword length and high-order modulation. However, their decoding algorithms suffer from very high computational complexity. In this paper, a low-complexity algorithm is proposed, named parity-check erased algorithm (PCEA), where an additional parity check bit is added to each symbol of the codeword when encoding and a series of simple operations are performed based on these bits during decoding. As a universal joint coding and decoding algorithm, the PCEA can be combined with arbitrary NB-LDPC encoding schemes and decoding algorithms based on message passing. The proposed algorithm facilitates significant improvement of decoding performance with a small decrease of the code rate. Additionally, it usually has an even better performance than a nearly same-rate code constructed by the original method, and requires much lower decoding complexity due to smaller size of the parity check matrix. Suwen Song, Jing Tian 0004, Jun Lin 0001, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2018 | An Efficient NB-LDPC Decoding Algorithm for Next-Generation MemoriesabstractDue to the aggressive technology scaling, the memory reliability has been seriously degraded, which poses a challenge to the widely used low-density parity-check (LDPC) codes. Non-binary LDPC (NB-LDPC) codes present larger coding gain and lower error floor than their binary counterparts in many cases, which show a great potential to be used in the next-generation memories. However, the excessive computational complexity of current NB-LDPC decoding algorithms form a bottleneck and limit their applications. In this paper, a novel algorithm, called dual-threshold-based shrinking based improved trellis-based min-sum algorithm (simply TIT-MSA), is proposed to deal with this problem. The improvements include two steps. The first step is for the check node processing (CNP). Based on the CNP of the simplified min-sum algorithm (SMSA) and that of the trellis-based extended min-sum algorithm (T-EMSA), an improved trellis-based min-sum algorithm (IT-MSA) is developed, which achieves better error performance and lower computational complexity than its origins. The second step is for the whole decoding process. Based on the IT-MSA, the TIT-MSA is proposed, for which two constant thresholds are introduced to remove redundant messages by constructing two subsets of the Galois field. Simulation results show that the error performance of the TIT-MSA is nearly the same as that of the EMSA. Meanwhile, the proposed algorithm can save almost 90% computations compared to the SMSA and T-EMSA. Jing Tian 0004, Jun Lin 0001, Zhongfeng Wang 0001 |
ISCAS | 1 |