VLDB 2026 Research / reviewers in the wild / expert
Suwen Song
dblp:239/9026
· DBLP profile ↗
20ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-7654-6582ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 3 first-author · 14 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Scalable Segment-Parallel Architecture for High-Efficiency Lossless Data CompressionabstractThe exponential growth of data volume puts significant pressure on the throughput and CPU resource usage of conventional software-based compression systems, driving the research focus of LZ4 algorithm towards its parallel hardware implementations. However, existing parallel architectures have to make a trade-off between throughput and compression ratio. To address the challenge, this paper presents a novel segment-parallel architecture that simultaneously delivers both high throughput and good compression ratio. The proposed architecture first introduces an interconnected dictionary scheme to preserve compression ratio while enabling parallel processing, limiting compression ratio degradation to less than 8% compared with software benchmarks across various parallelization levels. Second, a priority-based arbitration mechanism for data memory and a hierarchical depth scheduling strategy for hash table are proposed to enhance memory efficiency of multi-port memories. Additionally, the bit-width of hash table entries is reduced by exploiting deterministic address mapping relationships between hash values and input strings. Implemented on FPGA platforms, the proposed architecture achieves a state-of-the-art performance of 17.71 Gbps throughput, matching the performance of our previous work while delivering a$2.91\sim 3.85\times $improvement over other designs. It also demonstrates superior memory efficiency exceeding existing parallel implementations by$1.57\sim 5.57\times $. Suwen Song, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | A High-Performance and Hardware-Efficient Iterative Detection and Decoding Receiver for Polar-Coded Massive MIMO SystemabstractPolar-coded massive multiple-input multiple-output (MIMO) systems have attracted significant attention in wireless communications due to their superior performance with iterative detection and decoding (IDD). However, the practical implementation of polar-coded IDD receivers faces two critical challenges: the high computational complexity of massive MIMO detection and the difficulty in pursuing low-complexity and high-performance soft-output polar decoding. In this paper, we propose a comprehensive system design to address both challenges. For detection, we propose a vectorized Gauss-Seidel (VGS) detector supporting soft-input and soft-output (SISO) operations, achieving$2\times $higher hardware efficiency than existing works when implemented on FPGA. For decoding, we develop a partial-sum-based soft-output successive cancellation list (PS-SSCL) decoder that generates soft outputs without additional decoding procedures. Compared to state-of-the-art soft-output list (SOL) decoders, the PS-SSCL reduces memory usage by 55% and computational complexity by 62% while providing an extra 0.3 dB gain in IDD systems. Finally, an interleaved IDD receiver integrating the SISO VGS detector and PS-SSCL decoder is coded with RTL and synthesized under TSMC 28-nm CMOS technology, achieving an extra coding gain of 1.5 dB at FER$= 10^{-3}$over conventional separate detection and decoding (SDD) receiver with only 5% sacrifice in area efficiency. Huiyu Feng, Suwen Song, Chuan Zhang 0001, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | Enhanced 8B/6T Line Coding Design for High-Reliability Data Communication
Huiyu Feng, Suwen Song, Zhongfeng Wang 0001 |
IEEE Trans. Commun. | 2 |
| 2026 | Rethinking Area Optimization for High-Throughput Hardware Accelerator of Number Theoretic Transform: A Case Study on ML-KEMabstractThe computation-intensive modular multiplications in number theoretic transform (NTT) present significant bottlenecks for lattice-based post-quantum cryptography (PQC). In this article, we propose a systematic area optimization strategy through joint algorithm-hardware codesign. First, we propose two high-throughput architectures to accommodate diverse throughput requirements: a fully parallel NTT and a folded 64-point variant derived from the former, which both feature area-efficient constant modular multipliers. A novel permutation-based algorithm is developed to reuse the NTT architecture for inverse NTT without structural modifications. Second, an area-driven, automated decomposition-based design strategy for constant modular multipliers is developed to reduce hardware overhead. The strategy involves a three-stage process: decomposing, analyzing, and selecting, which are based on finite-field properties and comprehensive circuit modeling. Finally, the proposed architectures are evaluated on Artix-7 AC701 FPGA with the parameters of the module-lattice-Based key-encapsulation mechanism (ML-KEM) from the latest PQC standard. The results show that the optimized modular multipliers deliver an average resource reduction of approximately 46% over the prior art. Furthermore, operating at 189 MHz, the core throughputs reach 47.2 and$94.5\times 10^{6}$operations per second (OPS) for the proposed folded and fully parallel structures, respectively, while both architectures achieve an I/O-limited throughput of$23.6{\,}\times{\,}10^{6}$OPS. Compared to the state-of-the-art highly parallel designs, our work provides a 31.1% area efficiency improvement. Minghao Li 0001, Suwen Song, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2026 | A Low-Latency Hardware Architecture Design for Five-Error-Correcting Reed-Solomon DecoderabstractReed–Solomon (RS) codes are widely used in modern communication and storage systems. However, conventional iteration-based RS decoder architectures are struggling to meet the stringent demands of emerging latency-critical applications. Although noniterative architectures have been proposed for$t \leq 4$RS codes with low latency and low complexity, the design of such architectures remains largely unexplored for five-error-correcting RS codes. Motivated by this gap, we extend the Peterson–Gorenstein–Zierler (PGZ) algorithm to directly compute the error locator polynomial for$t=5$RS codes, thereby eliminating the need for iterative computation in the circuit. Then, a systematic three-step optimization method is further proposed to mitigate the hardware complexity increase caused by PGZ, which significantly decreases the required number of multipliers. Besides, given the prohibitive complexity of noniterative root-finding algorithms for$t = 5$RS codes, we instead adopt a hybrid PGZ-Chien search (PGZ-CS) architecture for the whole decoder. Based on the proposed methods, a low-latency and area-efficient$t=5$RS decoder is finally developed and implemented under the example RS(255, 245, and 5)code. Synthesized under 28-nm CMOS technology, the proposed decoder reduces decoding latency by 65.6%, overall area by 29.7%, and power consumption by 32.6% compared to the compensated simplified reformulated inversionless Berlekamp–Massey (CS-RiBM) decoder. Haobin Xu, Zichuan Qiu, Suwen Song, Zhongfeng Wang 0001, Liyang Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2026 | Low-Complexity Parallel Syndrome Computation and Chien Search Architecture Based on Reed-Muller Transform
Muxun Zhang, Suwen Song, Liyang Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | A Low-Complexity and High-Throughput Hardware Design for Lempel-Ziv 4 Compression AlgorithmabstractThe Lempel-Ziv (LZ) 4 compression algorithm, widely used in data transmission and storage, faces the challenge of high-speed implementation and increased complexity in the era of big data. Therefore, this paper proposes a single-core parallel architecture for LZ4 algorithm with high throughput and low complexity. Firstly, to enhance throughput, two innovative approaches are introduced from the perspective of parallelism and frequency with an acceptable compression ratio loss: each parallelization window is restricted to performing a single match, bridging the gap between actual and theoretical parallelism; the feedback loop in the circuit is broken by utilizing the spatial correlation between adjacent matches for higher frequency. Secondly, two optimization schemes are employed on resource-consuming modules to achieve low complexity. Multi-port hash tables using Live Value Table (LVT) are improved based on inherent data characteristics, significantly reducing the hardware resource consumption while ensuring excellent scalability on hash table depth and frequency. The match comparison operation is moved ahead, further reducing the logic resources by 64.36%. Finally, our design is implemented on FPGA and ASIC platforms. Experimental results on FPGA demonstrate that the proposed architecture achieves a throughput of 17.39 Gb/s, exhibiting a 2.86$\times $improvement over the state-of-the-art, along with a 6.46$\times $enhancement in area efficiency. Further optimizations including Canonic Signed Digit (CSD) coding and computational reuse on the ASIC platform result in a$45\times $improvement in area efficiency. Suwen Song, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | Fast Hardware Architecture With Efficient Matrix Computations for the Key Generation of Classic McElieceabstractClassic McEliece, with a remarkably stable security level, has been selected as one of the four key-establishment algorithms in the fourth-round evaluation of the post-quantum cryptography (PQC) standardization process of national institute of standards and technology (NIST). However, its memory-intensive and time-consuming key generation poses an obstacle to widespread use. In this paper, we propose a fast hardware implementation of the key generation incorporating several architectural optimizations. For the Gaussian elimination, we optimize the scheduling of computing resources and the memory access process and present a high-performance and flexible systemizer with multiple low fan-out systolic arrays. Besides, an algorithmic-level parallelized design for entry generation and Gaussian elimination is proposed to reduce the redundant computation time. A compact entry generator with a multi-level feedback mechanism and a 2-D high-speed FFT module facilitates continuous streaming the generated entries into the systemizer.FPGA implementation results show that our designs for the key generation improve time-area efficiency by 11.9% to 43.2% compared to the state-of-the-arts. Moreover, compared to the hardware implementations for the key generation of the other two quasi-cyclic code-based PQC algorithms, ours for Classic McEliece based on the random code achieves close to or better results in several metrics. Xinyuan Qiao, Jing Tian 0004, Suwen Song, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | A Novel Low-Complexity Massive MIMO Detector with Near-Optimum PerformanceabstractIn massive multiple-input multiple-output (MIMO) detection, likelihood ascent search (LAS) is well-known for its near-optimum performance with low complexity. It employs gradient descent to enhance the performance of suboptimal MIMO detectors, specifically minimum mean-square error (MMSE). In this paper, we introduce several new techniques to improve the MMSE-based LAS in terms of either complexity or performance. The MMSE is first replaced with optimized coordinate descent algorithm (OCD), which performs near MMSE with lower complexity. Then, the conventional OCD and LAS are reformulated and approximated to better reuse the computation for gradient descent, which is required in both algorithms. Besides, we also optimize the search strategy of LAS, leading to the improvement in both complexity and performance. The proposed detector, modulation-based successive gradient descent (MB-SGD) algorithm, outperforms MMSE-LAS and the latest low-complexity near-optimum detector in terms of either complexity or performance for 64×8 and 128×8 MIMO systems under 256-QAM. The corresponding architecture for a 128 × 8 256-QAM MIMO system has 75.8% lower latency, 2.02 × higher area efficiency, and 0.3 dB gain when implemented on a ×Xilinx Virtex-7 FPGA compared to OCD’s. Jinjie Hu, Suwen Song, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2024 | Efficient Soft-Output List Decoding of Polar Codes for Iterative Detection and Decoding in MIMO SystemabstractPolar-coded multiple-input multiple-output (MIMO) system has obtained increasing attention recently, in which iterative detection and decoding (IDD) brings improved performance compared to separate detection and decoding (SDD) schemes. IDD requires the exchange of soft information between the detector and decoder, but the best-performing decoding algorithm for polar codes, the cyclic redundancy check (CRC)-aided successive cancellation list (CA-SCL) algorithm, cannot generate soft outputs by itself. To address this issue, existing high-performance soft-output decoders combine the CA-SCL decoder with the belief propagation (BP) or soft cancellation (SCAN) decoding process to generate soft outputs, leading to high computational complexity and memory requirements. In this work, an efficient partial sum-based soft-output SCL (PS-SSCL) decoder is proposed, which utilizes the inherent partial sums generated during the CA-SCL decoding process to calculate soft outputs without introducing any additional decoding procedure. For the 5G (256,128) polar code, the proposed PS-SSCL decoder achieves a 57% reduction in memory requirements and a 62% reduction in computational complexity compared to the state-of-the-art soft-output list (SOL) polar decoder. Additionally, the PS-SSCL decoder can also bring more than 0.2 dB gain over the SOL decoder, when integrated into the IDD receiver with the linear minimum mean square error (LMMSE) detector. Huiyu Feng, Suwen Song, Zhongfeng Wang 0001 |
PIMRC | 2 |
| 2024 | RISC-V Custom Instructions of Elementary Functions for IoT Endpoint DevicesabstractThe computation of elementary functions is required in many tasks of Internet of Things (IoT) endpoint devices, for example, communications, image processing, and biomedical signal processing. IoT endpoint devices generally adopt software approaches to compute elementary functions, which take many cycles. To improve efficiency, this work proposes custom instructions for elementary functions to the open-source RISC-V instruction set architecture (ISA). In particular, several variants of the custom instructions (fast, intermediate, and tiny variants) are developed to satisfy the needs of various types of IoT devices. Microarchitecture design and VLSI circuit design are then proposed to efficiently support the extended ISA. Both software emulation and on-board evaluation of the new architecture are carried out with testbenches covering typical communication and computation tasks for IoT devices. The custom instructions gain speedups ranging from 3.3 to 18.0 compared to a baseline RV32IM design. ASIC synthesis results under TSMC 28nm technology demonstrate that the power overhead is$ \lt $5% with the tiny variant,$ \lt $17% with the intermediate variant, and$ \lt $26% with the fast variant, which is not significant considering the achieved speedup. The experimental results further confirm that the proposed custom instructions are computation-efficient and versatile to adapt to different IoT devices for various applications. Yuxing Chen 0001, Suwen Song, Lang Feng 0001, Zhongfeng Wang 0001 |
IEEE Trans. Computers | 3 |
| 2024 | Correlated Channel-Oriented Expectation Propagation-Based Detector for Massive MIMO SystemsabstractThe expectation propagation (EP) algorithm is near-optimal in massive multiple-input multiple-output (MIMO) systems but suffers from high computation complexity. Most of the previous works exploit the channel hardening property and introduce iterative matrix inversion algorithms to simplify the EP algorithm, but the performance degrades dramatically in non-ideal channels. In this paper, we propose a more universal EP-based detector, which can perform well in both ideal and non-ideal channels. Firstly, two general methods are proposed to effectively improve the detection performance and convergence speed of iterative matrix inversion algorithms under correlated channels. The proposed diagonal preprocessing (DP) method can improve the detection performance by more than 2-dB compared to not using this method; the novel eigenvalue parameter estimation method guarantees the convergence of all the frames. These two methods are applied to the second-order Richardson iteration (SORI) algorithm to derive the DP-SORI algorithm, which converges more than twice as fast as the state-of-the-art design. Secondly, for another important part of EP-based algorithms, namely the calculation of expectation and variance, complicated operations such as exponentiations, divisions, and inversions are all removed by algorithmic optimization. Moreover, based on the proposed approximate EP with DP-SORI (EPA-DP-SORI) algorithm, an efficient hardware design is developed, combining multiple optimization methods such as efficient matrix multiplication architecture design and low-complexity LDL decomposition. In addition to better detection performance compared with the state-of-the-art design, the presented EPA-DP-SORI detector can also deliver$1.27 \times $and$1.57 \times $higher area and energy efficiency. Yangyang Chen 0005, Huiyu Feng, Suwen Song, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | A Low-Complexity Soft-Output Massive MIMO Detector With Near-Optimum PerformanceabstractIn massive multiple-input multiple-output (MIMO) detection, the likelihood ascent search (LAS) algorithm is well-known for its near-optimum performance and low complexity. It employs gradient descent to enhance the performance of suboptimal MIMO detectors, specifically the minimum mean-square error (MMSE) algorithm. In this paper, we introduce several techniques to improve the MMSE-based LAS (MMSE-LAS) algorithm in terms of both complexity and performance. To reduce complexity, the MMSE is first replaced with the low-complexity optimized coordinate descent (OCD) algorithm at the cost of negligible performance loss. Then, the conventional OCD and LAS algorithms are optimized for better computation reuse. Besides, we derive a new soft-output computation formula for LAS to improve the coded performance. The proposed modulation-based successive gradient descent (MB-SGD) detector outperforms MMSE-LAS and the latest work in terms of either complexity or performance for$64\times 8$and$128\times 8$LDPC-coded MIMO systems with multiple modulations from QPSK to 256-QAM. The corresponding architecture for a$128\times 8$coded MIMO system supporting multiple modulations is implemented on a Xilinx Virtex-7 FPGA and with TSMC 28-nm CMOS technology, exhibiting 74.5% lower latency and 0.24 dB gain compared to OCD on FPGA, and also achieving$14.59\times $energy efficiency and$2.04\times $area efficiency over the state-of-the-art implementation on ASIC. Jinjie Hu, Suwen Song, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | An Efficient Massive MIMO Detector Based on Approximate Expectation PropagationabstractAmong expectation propagation (EP)-based massive multiple-input–multiple-output (MIMO) detection algorithms, EP with weighted Neumann-series approximation (EPA-wNSA) has the lowest computational complexity while requiring many iterations to guarantee the detection performance, which severely limits the throughput of hardware implementations. Through the joint optimization of algorithm and hardware architecture, we propose an EP-based detector with higher throughput and area efficiency. First, the second-order Richardson iteration (SORI) algorithm is employed to replace the wNSA algorithm for higher convergence speed. Then three algorithmic transformations are proposed to minimize the overall complexity. Simulation results show that the proposed EPA-SORI algorithm requires much fewer iterations to achieve comparable or even better detection performance compared with EPA-wNSA. Furthermore, an efficient detector architecture is delicately designed by incorporating multiple optimization methods, such as reverse data flow, advanced addition, and rounding cells. Implemented with the Taiwan Semiconductor Manufacturing Company (TSMC) 28-nm CMOS technology, the proposed detector has$2.2 \times $higher throughput than the state-of-the-art EP-based detector. Yangyang Chen 0005, Suwen Song, Zhongfeng Wang 0001, Jun Lin 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | Performance Analysis of Extended Integrated Interleaved CodesabstractExtended integrated interleaved (EII) codes, as the versatile alternative to locally recoverable codes (LRCs), show great potential in distributed storage systems, in which the output bit-error-rate (BER) below 10−15is usually demanded. However, it is time-consuming to reach such a low BER through normal software simulation, which brings inconvenience to the code construction. To solve the above problem, this work presents an analysis method to evaluate the decoding performance of EII codes, and no simulation is required. Numerical results show that the estimated frame-error-rate (FER) matches well with the simulated FER, so does the BER. Moreover, the failure probability of each decoding stage can be predicted accurately. Therefore, we can dig deep into the decoding behavior of each stage, which guides the adjustment of redundancy distribution, improving the error correction performance. Finally, the theoretical analysis for regular EII codes is simplified to reduce calculations. Keyue Deng, Xinyuan Qiao, Yuxing Chen 0001, Suwen Song, Zhongfeng Wang 0001 |
APCC | 4 |
| 2022 | A Novel Interleaving Scheme for Concatenated Codes on Burst-Error ChannelabstractWith the rapid development of Ethernet, RS (544, 514) (KP4-forward error correction), which was widely used in high-speed Ethernet standards for its good performance-complexity trade-off, may not meet the demands of next-generation Ethernet for higher data transmission speed and better decoding performance. A concatenated code based on KP4-FEC has become a good solution because of its low complexity and excellent compatibility. For concatenated codes, aside from the selection of outer and inner codes, an efficient interleaving scheme is also very critical to deal with different channel conditions. Aiming at burst errors in wired communication, we propose a novel matrix interleaving scheme for concatenated codes which set the outer code as KP4-FEC and the inner code as Bose-Chaudhuri-Hocquenghem (BCH) code. In the proposed scheme, burst errors are evenly distributed to each BCH code as much as possible to improve their overall decoding efficiency. Meanwhile, the bit continuity in each symbol of the RS codeword is guaranteed during transmission, so the number of symbols affected by burst errors is minimized. Simulation results demonstrate that the proposed interleaving scheme can achieve a better decoding performance on burst-error channels than the original scheme. In some cases, the extra coding gain at the bit-error-rate (BER) of 1 × 10−15can even reach 1 dB. Suwen Song, Zhongfeng Wang 0001 |
APCC | 2 |
| 2022 | An Area-Efficient Message Passing Detector for Massive MIMO SystemsabstractRecently, massive multiple-input multiple-output (MIMO) detection schemes based on message passing detection (MPD) have attracted extensive attention due to their good performance-complexity tradeoff. In this paper, to facilitate a high-throughput detector design, we introduce a layered updating schedule and propose an improved layered MPD (ILMPD) algorithm. In the new algorithm, several algorithmic transformations or approximations are derived for lower complexity. For instance, by exploiting the property of quadratic functions, the numbers of multiplications and additions in the constellation matching are both reduced by half; Through reasonable approximations, the multiplication, addition, and sorting operations in the initialization are all removed. Moreover, a lightweight early termination strategy is explored, reducing the number of detection iterations by nearly 20%. Based on the proposed ILMPD algorithm, an area-efficient architecture is devised, where several optimization methods are proposed for fewer resources and higher clock frequency. Compared with the state-of-the-art design, the presented ILMPD detector can deliver a nearly$3\times $higher area efficiency. A reconfigurable version of the proposed detector has also been developed, which can well support modulations QPSK to 256-QAM and still exhibits a superior area efficiency. Suwen Song, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | A Universal Efficient Circular-Shift Network for Reconfigurable Quasi-Cyclic LDPC DecodersabstractQuasi-cyclic low-density parity-check (QC-LDPC) codes for modern communication standards usually have multiple code rates and block lengths. Therefore, reconfigurable LDPC decoders have received widespread attention, which require circular-shift networks to support various expansion factors. Besides, for inputs smaller than the network size, the circular-shift network is desired to process multiple frames in parallel to maximize hardware utilization efficiency. The increasing demands put severe challenges to low-complexity implementations of shift networks, especially for codes with numerous expansion factors, such as 5G LDPC codes. In this brief, we present a universal design of efficient reconfigurable circular-shift networks. Through an ingenious modification on the order of permutations, the generation of control signals is considerably simplified, leading to a significant reduction of area and critical path. Moreover, a hybrid architecture organically integrating different networks is proposed for further complexity reduction. Implementation results under TSMC 90 nm technology demonstrate that the proposed network can achieve 25% area reduction and 46% area-efficiency (AE) improvement over the state-of-the-art ones. Suwen Song, Hangxuan Cui, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | A New Probabilistic Gradient Descent Bit Flipping Decoder for LDPC CodesabstractProbabilistic gradient descent bit-flipping (PGDBF) is the state-of-the-art hard-decision algorithm for decoding low-density parity-check (LDPC) codes on binary symmetric channel (BSC). However, there still exists a considerable performance gap between the PGDBF algorithm and soft-decision algorithms, especially in the error-floor region. To bridge this performance gap, a tabu-list aided PGDBF (T-PGDBF) algorithm is proposed in this paper. In the T-PGDBF algorithm, a tabu-list is employed to help the decoding escape from trapping sets, which is the main cause of the error-floor phenomenon. The bits which are flipped in the current iteration will be added to the tabu-list to prevent them being flipped in the next iteration. Simulation results show that the T-PGDBF algorithm offers a significant performance gain when compared to the PGDBF algorithm, which can reach that of soft-decision algorithms. We also present the hardware architecture to implement the T-PGDBF algorithm. Synthesis results show that the improved performance offered by the T-PGDBF algorithm can be obtained with a small hardware overhead. Hangxuan Cui, Jun Lin 0001, Suwen Song, Zhongfeng Wang 0001 |
ISCAS | 3 |
| 2019 | A Novel Low-Complexity Joint Coding and Decoding Algorithm for NB-LDPC CodesabstractNon-binary low-density parity-check (NB-LDPC) codes exhibit a much better performance than their binary counterparts, especially for moderate codeword length and high-order modulation. However, their decoding algorithms suffer from very high computational complexity. In this paper, a low-complexity algorithm is proposed, named parity-check erased algorithm (PCEA), where an additional parity check bit is added to each symbol of the codeword when encoding and a series of simple operations are performed based on these bits during decoding. As a universal joint coding and decoding algorithm, the PCEA can be combined with arbitrary NB-LDPC encoding schemes and decoding algorithms based on message passing. The proposed algorithm facilitates significant improvement of decoding performance with a small decrease of the code rate. Additionally, it usually has an even better performance than a nearly same-rate code constructed by the original method, and requires much lower decoding complexity due to smaller size of the parity check matrix. Suwen Song, Jing Tian 0004, Jun Lin 0001, Zhongfeng Wang 0001 |
ISCAS | 1 |