VLDB 2026 Research / reviewers in the wild / expert
Wenqing Song
dblp:119/0925
· DBLP profile ↗
21ranked-venue papers
2as first author
13since 2021 · last 2025
0000-0002-8806-8717ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 2 first-author · 11 since 2021Computer networks · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Belief Propagation Decoding for Short Codes on Structured Sparse Parity-Check MatricesabstractAs successfully adopted in standard long code scenarios, belief propagation (BP) decoding has been considered a promising universal decoding candidate for next-generation wireless communications. However, when applied to short codes, BP decoding suffers from poor error correction performance due to harmful cycle structures in the Tanner graph. In this paper, we address this issue by designing a structured, sparse parity-check matrix (ssPCM) framework, composed of multiple cycle-free parity-check row blocks (PCRBs). The resulting ssPCMs feature regular row weights and perform better than the state-of-theart 4 -cycle-free row redundant PCMs across Bose-Chaudhuri-Hocquenghem (BCH) codes of length 63. Yifei Shen 0003, Zongyao Li 0003, Emmanuel Boutillon, Wenqing Song, Yuqing Ren, Chuan Zhang 0001, Xiaohu You 0001, Andreas Peter Burg |
ISIT | 4 |
| 2025 | An Energy Efficient Residual Spiking Neural Network Accelerator With Ternary SpikesabstractSpiking neural networks (SNNs) use discrete binary spikes to transfer information between neurons, which is different from artificial neural networks (ANNs). Although event-based characteristics bring potential computation power and efficiency to SNNs, the long processing time window of discrete spikes leads to high latency. In this brief, a spike version of the residual network using ternary spikes is proposed. A shorter time window is required to achieve competitive performance because the ability to transfer information of the ternary spikes is strengthened. An SNN accelerator based on the proposed residual network with ternary spikes is designed and implemented with 28 nm CMOS technology, and the core area is 0.63 mm2. The proposed SNN accelerator achieves the classification accuracy of 92.07% on CIFAR-10 dataset with SResNet20 and only 6 time steps. The accelerator achieves 0.39 mJ energy consumption per frame with a throughput of 165.7 FPS when running at 500 MHz. Congyi Sun, Wenqing Song, Qinyu Chen, Chenyang Dai, Li Li 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Edge-Spreading Raptor-Like LDPC Codes for 6G Wireless SystemsabstractNext-generation channel coding has stringent demands on throughput, energy consumption, and error rate performance while maintaining key features of 5G New Radio (NR) standard codes such as rate compatibility, which is a significant challenge. Due to excellent capacity-achieving performance, spatially-coupled low-density parity-check (SC-LDPC) codes are considered a promising candidate for next-generation channel coding. In this paper, we propose an SC-LDPC code family called edge-spreading Raptor-like (ESRL) codes. Unlike other SC-LDPC codes that adopt the structure of existing rate-compatible LDPC block codes before coupling, ESRL codes maximize the possible locations of edge placement and focus on constructing an optimal coupled matrix. Moreover, a new graph representation called the unified graph is introduced. This graph offers a global perspective on ESRL codes and identifies the optimal edge reallocation to optimize the spreading strategy. We conduct comprehensive comparisons of ESRL codes and 5G-NR LDPC codes. Simulation results demonstrate that when all decoding parameters and complexity are the same, ESRL codes have obvious advantages in error rate performance and throughput compared to 5G-NR LDPC codes in some specific scenarios (low and high number of iterations), making them a promising solution towards next-generation channel coding. Yuqing Ren, Leyu Zhang, Yifei Shen 0003, Wenqing Song, Emmanuel Boutillon, Alexios Balatsoukas-Stimming, Andreas Peter Burg |
IEEE Trans. Commun. | 4 |
| 2024 | HAS-RL: A Hierarchical Approximate Scheme Optimized With Reinforcement Learning for NoC-Based NN AcceleratorsabstractNetwork-on-Chip (NoC) is a scalable on-chip communication architecture for the NN accelerator, but with the increase in the number of nodes, the communication delay becomes higher. Applications such as machine learning have a certain resilience to noisy/erroneous transmitted data. Therefore, approximate communication becomes a promising solution to improving performance by reducing traffic loads under the constraint of the acceptable maximum accuracy loss of neural networks. It is a key issue to balance the result quality and the communication delay for approximate NoC systems. The traditional approximate NoC only considers the node-to-node approximation-based dynamic traffic regulation. However, the dynamically changing traffic patterns across different nodes, different times, and different applications lead to a huge search space, which makes it hard to explore an optimal global approximation solution. In this paper, we propose a quality model for different neural networks, which presents the relationship between the quality loss and the data approximate rate. Then, a hierarchical approximate scheme optimized with reinforcement learning (HAS-RL) is proposed and we reduce the complexity of the HAS-RL by reducing the state space and action space, which will reduce the resource overhead as well. After that, we embed a global approximate controller in the NoC system, in which we deploy a policy network trained with the offline reinforcement learning algorithm to adjust the data approximate rates of each node at run time. Compared with the state-of-the-art method, the proposed scheme reduces the average network delay by 13.5% while their accuracies are similar. The proposed HAS-RL only causes an additional area overhead of 1.24% and power consumption of 0.77% compared with the traditional router design. Shize Zhou, Yongqi Xue, Wenjie Fan 0004, Tong Cheng, Jinlun Ji, Chenyang Dai, Wenqing Song, Qinyu Chen, Chang Gao 0002, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2023 | A DSP-Purposed REconfigurable Acceleration Machine (DREAM) for High Energy Efficiency MIMO Signal ProcessingabstractThe wireless baseband processing algorithms are still developing and show a great diversity. The development of ASIC implementations cannot quickly adapt to the evolution of algorithms and standards. Meanwhile, the general-purpose processors cannot meet the real-time requirements in some scenarios. This paper proposes a DSP-purposed REconfigurable Acceleration Machine (DREAM) core for wireless baseband digital signal processing, which has a good trade-off between flexibility and performance. First, we abstract a set of shared operators with a moderate granularity from a variety of wireless MIMO signal processing algorithms. Then, we propose a two-step configuration process to reduce the size of the required reconfiguration bits. Besides, we design a conflict-free address generator to transfer data between the on-chip scratchpad memory and reconfiguration processing elements with high efficiency and high throughput. Finally, the prototype DREAM core has been implemented in TSMC CMOS 28 nm, and its area and power consumption have been analyzed. The chip has great flexibility in supporting a variety of wireless MIMO processing algorithms and a wide range of MIMO scales. The proposed DREAM core can achieve the normalized area efficiency and the normalized energy efficiency of$0.67~Gbps/MGE$and$15.05~Gbps/W$, which are$1.56\times $and$4.18\times $those of state-of-the-art reconfigurable implementations when running the WeJi-based MIMO detection algorithm. Kai Chen 0034, Wenqing Song, Guoqiang He, Sirui Shen, Huizheng Wang, Chuan Zhang 0001, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | A Hierarchical Parallel Discrete Gaussian Sampler for Lattice-Based CryptographyabstractDiscrete Gaussian sampling is one of the important components in lattice-based cryptosystems which are promising candidates for post-quantum cryptographic algorithms. For sufficient security and satisfactory performance, the Knuth-Yao algorithm is an efficient way to implement discrete Gaussian samplers. Nevertheless, most polynomials in lattice-based cryptography have 256 coefficients or more, which suffers from long latency to complete the sample generation. In this paper, the first parallel discrete Gaussian sampler with hierarchical structure is proposed, while keeping statistical distance to the actual distribution. Based on the imbalanced visiting frequency of the probability matrix, a three-stage generation strategy is adopted with hierarchical bit search units (BSUs) that can greatly reduce area consumption of the repeated costly lookup tables. Besides the architecture improvement, a lowest-set-bit scanning scheme is introduced to BSUs. Moreover, the parallelism of our design provides obfuscation ability against side-channel attacks (SCAs). A practical hardware implementation of discrete Gaussian distributions with $\sigma$=3.33 on the Xilinx Virtex-5 XC5VLX30 FPGA device spends 26.12 ns on average to generate 256 samples, consuming 994 slices. Results have verified its advantages of area efficiency over the state-of-the-arts (SOAs). Sirui Shen, Wenqing Song, Xinyu Wang 0027, Xinyu Shao, Zhonghai Lu, Li Li 0003 |
ISCAS | 2 |
| 2021 | Dynamic and Traffic-Aware Medium Access Control Mechanisms for Wireless NoC ArchitecturesabstractWireless NoC (WiNoC) has low latency and simple wiring, which can reduce the energy consumption caused by the metal interconnection in traditional NoC architectures. However, traditional time division based media access control (MAC) mechanism in WiNoC is not aware of different wireless interfaces' (WIs) traffic demands, resulting in an unreasonable distribution of wireless communication channels and degradation in performance. Hence, in order to dynamically allocate wireless channels to the WIs based on their traffic demands, a dynamic and traffic-aware MAC mechanism is required. In this paper, we design a traffic demand predictor for each WI based on its current and history traffic conditions. According to the predicted demands, we are able to allocate access to wireless channels dynamically and switch between two kinds of time division based MAC mechanisms. Simulations under various conditions indicate that the average delay decreases by 30% and 20% on average compared with a traditional MAC mechanism and an existing dynamic time division based one, respectively. Moreover, the network with the dynamic and traffic-aware MAC enters the saturation point at a higher packet injection rate. Wenqing Song, Zhonghai Lu, Li Li 0003 |
ISCAS | 2 |
| 2021 | Adaptive Successive Cancellation Priority Decoder for 5G Polar CodesabstractAs two common successive cancellation (SC)-based decoding algorithms of polar codes, the SC list (SCL) and SC stack (SCS) decoder can achieve satisfactory error correction performance, especially with increased list size or stack depth. Nevertheless, a large list size or stack depth will lead to high computational complexities and hardware resources. To this end, successive cancellation priority (SCP) decoding with priority- first searching strategy and trellis-like storage is proposed to offer one solution. In this paper, an efficient SCP decoder is first proposed to verify its advantages over SCL and SCS decoders. Furthermore, an adaptive node-inserting scheme is proposed to reduce the number of bits insert into the priority queue. Numerical results have shown that for the polar code with transmission length 1024 and rate 1/2, the proposed adaptive SCP (ASCP) decoder can achieve significant time complexity reduction on average compared with the standard SCL decoder. The hardware architecture of SCP decoding is implemented using 65-nm CMOS technology and the results show better throughput compared with the SCS decoder. Wenqing Song, Yifei Shen 0003, Chuan Zhang 0001, Li Li 0003 |
ISCAS | 1 |
| 2021 | Efficient Fast-SCAN Flip Decoder for Polar CodesabstractSoft-output decoder is of great importance to be applied in iterative receivers, of which belief propagation (BP) algorithm has been widely studied for 5G low-density parity- check (LDPC) and polar codes. However, for polar codes, BP decoding suffers from high computational complexity and unsatisfactory convergence. To this end, soft cancellation (SCAN) polar decoder has recently drawn attention from academia and can be further improved by using the bit-flipping strategy. Limited by the serial nature of message propagation, the SCAN flip (SCANF) decoder cannot meet a high throughput. In this paper, we accelerate the decoding speed by the fast processing mechanism, conducting Fast-SCANF decoder. The corresponding hardware architecture is designed with memory optimization and implemented by TSMC 40nm technology, delivering a 2.1 Gbps throughput and 65 pJ/b energy. To the knowledge of authors, this is the first SCANF hardware decoder. Leyu Zhang, Yutai Sun, Yifei Shen 0003, Wenqing Song, Xiaohu You 0001, Chuan Zhang 0001 |
ISCAS | 4 |
| 2021 | Optimizing Vertical Link Placement and Congestion Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCsabstractThe fully connected 3D-NoCs in which all routers are vertically connected with their neighbors above and below need a lot of Through-Silicon-Vias (TSVs), and they will occupy a large silicon area and reduce the fabrication yield. Thus, the idea of partially connected 3D-NoCs has emerged. The optimal number and placement of the vertical links (elevators) must be determined at the chip design stage, which is a multiobjective optimization problem of the performance and the cost. However, optimizing the static elevator placement needs a great amount of calculation and we can not examine all possible solutions at design time. Therefore, we propose a hybrid heuristic strategy for the static elevator placement and assignment, in which the genetic algorithm and the tabu search are combined. The dynamic assignment method is essential for the partially connected 3D-NoCs, and it leads to different traffic distributions and therefore has a huge impact on performance. Many previous static assignment methods can not dynamically change the elevator assignment according to the real-time states of the network, thus it may lead to network congestion. A congestion-aware dynamic assignment (CDA) scheme is proposed in this article, which considers the impact of the distance factor and the congestion factor on the network performance. Experiments show that the proposed CDA method can improve the network performance by 67%-86% compared with the random selection algorithm and can improve the reliability of the partially connected 3D-NoC as well. The key component for the CDA method, the path selection module (PSM), is implemented in FPGA, and the results show that its area cost is negligible compared with a router. Chuan Zhang 0001, Wenqing Song, Qinyu Chen, Hui Chen 0015, Li Li 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Symmetric-Mapping LUT-Based Method and Architecture for Computing XY-Like FunctionsabstractWe propose a new method and hardware architecture to compute the functions expressed as XY (X and Y are arbitrary floating-point numbers), which can support arbitrary Nth root, exponential and power operations. Because of the complexity of direct computation, we usually convert it to logarithm, multiplication, and antilogarithm operations. Traditional approaches suffer from long latency, large area and high power consumption. To solve this problem, we propose a symmetric-mapping lookup table (SM-LUT) to be capable of computing log2x (x ∈ [1, 2]) and 2x(x ∈ [0, 1]) simultaneously. It lays the foundation for computing XY. To further improve hardware performance of our architecture, we propose a multi-region address searcher to speed up the calculation of SM-LUT. In addition, we use an optimized Vedic multiplier to shorten the critical path and improve the efficiency of multiplication, which is included in computing XY. Under the TSMC 40nm CMOS technology, we design and synthesize a reference circuit to compute XY with a maximum relative error of 10-3. The report shows that the reference circuit achieves the area of 14338.50 μm2and the power consumption of 4.59 mW at the frequency of 1 GHz. In comparison with the state-of-the-art work under the same input range and similar precision, it saves 78.57% area and 80.42% power consumption for N√R computation and 82.89% area and 81.89% power consumption for RN computation averagely. On top of that, our architecture reduces the computation latency by 62.77% averagely and has one more order of magnitude of energy efficiency than others. Hui Chen 0015, Heping Yang, Wenqing Song, Zhonghai Lu, Li Li 0003, Zongguang Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Hardware Implementation for Belief Propagation Flip Decoding of Polar CodesabstractBelief propagation (BP) decoding has natural advantages in throughput for polar codes to meet high-speed and low-latency requirements. The soft outputs of BP decoding can be utilized further for joint detection and decoding in the baseband communication system. However, its error-correction performance is not comparable with the successive cancellation list (SCL) decoding. Belief propagation flip (BPF) decoding is recently proposed to improve the error-correction performance of BP decoding and indicates the potential to compete with SCL decoding. In this paper, we propose an advanced BPF (A-BPF) scheme that reduces the decoding latency with the help of one critical bit and improves the error-correction performance by the proposed joint detection criterion. To improve area efficiency in the hardware level, an optimized sorting network is proposed and applied for the A-BPF decoder. The decoder is implemented on 65 nm CMOS technology for length-1024 and rate-1/2 polar codes, and the results show that the proposed decoder can achieve a close frame error rate performance to the SCL decoder with four lists and deliver a throughput of 5.17 Gb/s at Eb/N0= 4.0 dB. Houren Ji, Yifei Shen 0003, Wenqing Song, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Efficient Row-Layered Decoder for Sparse Code Multiple AccessabstractSparse code multiple access (SCMA) is a promising technology for the development of wireless communication, which supports a large number of overloading users and enjoys high spectral efficiency. However, conventional SCMA decoders suffer very high complexity in implementations. Changing the updating scheme is a superior approach to reduce complexity, which guarantees the updated information immediately join in the following message propagating of the current iteration and accelerates the decoding convergence. In this paper, a row-layered message passing algorithm (MPA) is proposed, which offers a good trade-off between the hardware complexity and the bit error rate (BER) performance. Simulation results show that the proposed decoder saves 66.7% computation complexity compared with the original MPA with the similar BER performance. Pipelining and folding technology are adopted in VLSI implementations. The synthesis results with 45-nm CMOS technology show that the proposed decoder can achieve higher hardware efficiency and throughput under a high frequency than the existing decoders, achieving 1777.78 Mb/s throughput with 1.112 mm2area consumption. Xu Pang, Wenqing Song, Yifei Shen 0003, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | Bipartite Belief Propagation Polar Decoding With Bit-FlippingabstractFor the scenarios with high throughput requirements, the belief propagation (BP) decoding is one of the most promising decoding strategies for polar codes. By pruning the redundant variable nodes (VNs) and check nodes (CNs) in the original factor graph, the graph is condensed to a sparse bipartite graph which is similar to the graph for low-density parity-check (LDPC) codes. In this paper, we introduce the bit-flipping scheme into the LDPC-like BP (L-BP) decoding and propose two methods to identify the error-prone VNs. By additional decoding attempts, the L-BP flip (L-BPF) decoding improves the error-rate performance with a similar average complexity for high Eb=N0values. The simulation results show that the L-BPF decoding achieves 0:25 dB gain compared with the L-BP decoding. Zihao Gong, Yifei Shen 0003, Houren Ji, Wenqing Song, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ICASSP | 4 |
| 2020 | Improved Belief Propagation Polar Decoders With Bit-Flipping AlgorithmsabstractSince the inherent serial nature of successive cancellation list (SCL) decoding results in a long latency, belief propagation (BP) decoding for polar codes has drawn attention for high-throughput applications. However, its error correction performance is inferior to that of SCL decoding. Therefore, the bit-flipping strategy has been recently applied to BP decoding, which can approach the SCL decoding performance through multiple additional decoding attempts. The original BP flip (BPF) decoding suffers from an inaccurate identification of erroneous bits by a fixed flip set (FS), which has been improved by the generalized BPF (GBPF) decoding. In this article, the GBPF decoding is extended to support multiple bits being flipped in one decoding attempt. In addition, for two types of decoding errors: detected errors and undetected errors, we propose two novel methods to more effectively identify erroneous bits. For detected errors, the concept of loop sets is defined and a loopbased identification method is introduced based on the study of error patterns of BP decoding. On the other hand, a method to generate a more accurate fixed FS is proposed for undetected errors, which considers the bit error distribution under BP decoding. Combining the two methods, the GBPF with merged sets (GBPF-MS) decoding can achieve the SCL-8 performance and outperforms the state-of-the-art BPF, BP list, and SC flip (SCF) decoding, for polar codes with length 1024 and information rate 1/2. Implemented by 40nm CMOS technology, the proposed GBPF-MS decoder with ten flips exhibits an average throughput of 4.19 Gbps at 2.5 dB, which is 1.6× and 1.72× faster than the state-of-the-art SCL-4 and SCF decoders, respectively. Yifei Shen 0003, Wenqing Song, Houren Ji, Yuqing Ren, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Commun. | 2 |
| 2020 | An Efficient Accelerator for Multiple Convolutions From the Sparsity PerspectiveabstractConvolutional neural networks (CNNs) have emerged as one of the most popular ways applied in many fields. These networks deliver better performance when going deeper and larger. However, the complicated computation and huge storage impede hardware implementation. To address the problem, quantized networks are proposed. Besides, various convolutional structures are designed to meet the requirements of different applications. For example, compared with the traditional convolutions (CONVs) for image classification, CONVs for image generation are usually composed of traditional CONVs, dilated CONVs, and transposed CONVs, leading to a difficult hardware mapping problem. In this brief, we translate the difficult mapping problem into the sparsity problem and propose an efficient hardware architecture for sparse binary and ternary CNNs by exploiting the sparsity and low bit-width characteristics. To this end, we propose an ineffectual data removing (IDR) mechanism to remove both the regular and irregular sparsity based on dual-channel processing elements (PEs). Besides, a flexible layered load balance (LLB) mechanism is introduced to alleviate the load imbalance. The accelerator is implemented with 65-nm technology with a core size of 2.56 mm2. It can achieve 3.72-TOPS/W energy efficiency at 50.1 mW, which makes it a promising design for embedded devices. Qinyu Chen, Wenqing Song, Zhonghai Lu, Li Li 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | Smilodon: An Efficient Accelerator for Low Bit-Width CNNs with Task PartitioningabstractConvolutional Neural Networks (CNNs) have been widely applied in various fields such as image and video recognition, recommender systems, and natural language processing. However, the massive size and intensive computation loads prevent its feasible deployment in practice, especially on the embedded systems. As a highly competitive candidate, low bit-width CNNs are proposed to enable efficient implementation. In this paper, we propose Smilodon, a scalable, efficient accelerator for low bit-width CNNs based on a parallel streaming architecture, optimized with a task partitioning strategy. We also present the 3D systolic-like computing arrays fitting for convolutional layers. Our design is implemented on Zynq XC7Z020 FPGA, which can satisfy the needs of real-time with a frame rate of 1, 622 FPS throughput, while consuming 2 1 Watt. To the best of our knowledge, our accelerator is superior to the state-of-the-art works in the tradeoff among throughput, power efficiency, and area efficiency. Qinyu Chen, Kaifeng Cheng, Wenqing Song, Zhonghai Lu, Li Li 0003, Chuan Zhang 0001 |
ISCAS | 4 |
| 2019 | Efficient Successive Cancellation Stack Decoder for Polar CodesabstractAs an improved version of successive cancellation (SC) polar decoder, an SC stack (SCS) decoder has been proposed for performance improvement. However, the existing SCS polar decoder suffers a lot from high time complexity at low signal-to-noise ratio (SNR) region and space complexity compared with the SC decoder. To this end, two improved decoders are proposed to reduce time and space complexity in both low and high SNR regions. The first one is the segmented cyclic redundancy check (CRC)-aided SCS (SCA-SCS) decoder, which is based on segmented parity checkers. The second one is the adaptive SCS (ASCS) decoder, which has the flexibility of stack depth and searching width. Furthermore, a channel condition estimator is proposed to select appropriate decision criteria for different SNR scenarios. Results have shown that for the polar code of length 1024 and rate 1/2, two improved SCS decoders can perform better than the traditional SCS decoder. The proposed SCA-SCS decoder and the ASCS decoder can achieve 10.8% and 11.42% time complexity reduction and 31.68% and 60.85% space complexity reduction on average over binary-input additive white Gaussian noise channels (BI-AWGNCs), respectively. Efficient parallel hardware architecture of the SCS polar decoder is first proposed and implemented with 90- and 65-nm technologies. Results have verified its advantages over the state of the art (SOA). Wenqing Song, Huayi Zhou 0002, Kai Niu 0001, Zaichen Zhang, Li Li 0003, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Hardware Efficient and Low-Latency CA-SCL Decoder Based on Distributed SortingabstractFor polar codes, cyclic redundancy check (CRC)aided successive cancellation list (CA-SCL) decoder has attracted increasing attention from both academia and industry. In this paper, a hardware efficient and low-latency CA-SCL polar decoder based on distributed sorting is first proposed. For path metric (PM) sorting of each level, a distributed sorting (DS) algorithm is proposed to reduce the comparison complexity from (L2) to (L) (L denotes list size), together with the latency from kL2to kL (k is a coefficient independent of L). Employing folding technique, the N-bit folding polar decoder can be implemented based on the basic √N-bit polar decoder. In addition, pipelining technique is employed to refine the timing issue resulting from folding. The CRC is performed for 2L candidate paths serially to reduce hardware cost. According to demo of (1024, 512) code on Altera Stratix V FPGA, the proposed CA-SCL decoders with L = 2 and adjustable L = 2, 4 consume 9% and 50% board resources, respectively. Decoding latencies (in terms of clock cycles) are 2, 528 and 4, 064, respectively. For L = 2 and 4, we can achieve the frame error rate (FER) of 10-2at the signal noise ratio (SNR) of 2.36 dB and 2.06 dB, respectively. Compared with the floating point results, the performance degradation is negligible. Thus, the proposed design is suitable and adjustable for different real-life scenarios. Xiao Liang 0005, Junmei Yang, Chuan Zhang 0001, Wenqing Song, Xiaohu You 0001 |
GLOBECOM | 4 |
| 2016 | Joint detection and decoding for MIMO systems with polar codesabstractAs well known, the near-optimal K-best detection is popular in multiple-input and multiple-output (MIMO) systems. In this paper, we first propose the joint approaches of K-best detection and polar decoding. For joint detection and decoding (JDD) approach, both hard and soft decisions are considered. The simplified successive cancellation (SSC) decoding is exploited for hard decision, and the successive cancellation list (SCL) decoding is used as soft decision. The system setup for JDD is als o introduced, in which the modulation points across several channels are considered together. Simulation results have demonstrated the performance advantage of the JDD algorithms over the separated ones. For 1/2-rate polar codes, JDD schemes show 50% complexity reduction compared to the separated ones. Furthermore, by employing SSC hard decoding, the JDD algorithm is promising for high-throughput and low-complexity application s. Junmei Yang, Chuan Zhang 0001, Wenqing Song, Shugong Xu, Xiaohu You 0001 |
ISCAS | 3 |
| 2016 | Segmented CRC-Aided SC List Polar DecodingabstractBecause of the existence of channel noise, channel coding serves as an indispensable part of mobile communication system and the essential guarantee for the reliable, accurate, and effective transmission of information. As one of the most competitive channel code candidates for the 5th generation (5G) mobile communication, polar codes are the first codes which can provably achieve the symmetric capacity of binary-input discrete memoryless channels (B-DMCs). In this paper, the segmented CRC- aided successive cancellation list (SCA-SCL) polar decoding scheme is proposed for better tradeoff of performance and complexity. Numerical results on binary-input additive white Gaussian noise channel (BI-AWGNC) have shown that, at SNR of 0.5 dB, this approach successfully provides as high as 41.65% complexity reduction and similar decoding performance compared to state-of-the-art ones. Huayi Zhou 0002, Chuan Zhang 0001, Wenqing Song, Shugong Xu, Xiaohu You 0001 |
VTC Spring | 3 |