VLDB 2026 Research / reviewers in the wild / expert
Chuan Zhang 0001
dblp:23/1788-1
· DBLP profile ↗
98ranked-venue papers
11as first author
50since 2021 · last 2026
0000-0002-7736-6487ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 48 · 7 first-author · 24 since 2021Computer networks · 24 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A High-Performance and Hardware-Efficient Iterative Detection and Decoding Receiver for Polar-Coded Massive MIMO SystemabstractPolar-coded massive multiple-input multiple-output (MIMO) systems have attracted significant attention in wireless communications due to their superior performance with iterative detection and decoding (IDD). However, the practical implementation of polar-coded IDD receivers faces two critical challenges: the high computational complexity of massive MIMO detection and the difficulty in pursuing low-complexity and high-performance soft-output polar decoding. In this paper, we propose a comprehensive system design to address both challenges. For detection, we propose a vectorized Gauss-Seidel (VGS) detector supporting soft-input and soft-output (SISO) operations, achieving$2\times $higher hardware efficiency than existing works when implemented on FPGA. For decoding, we develop a partial-sum-based soft-output successive cancellation list (PS-SSCL) decoder that generates soft outputs without additional decoding procedures. Compared to state-of-the-art soft-output list (SOL) decoders, the PS-SSCL reduces memory usage by 55% and computational complexity by 62% while providing an extra 0.3 dB gain in IDD systems. Finally, an interleaved IDD receiver integrating the SISO VGS detector and PS-SSCL decoder is coded with RTL and synthesized under TSMC 28-nm CMOS technology, achieving an extra coding gain of 1.5 dB at FER$= 10^{-3}$over conventional separate detection and decoding (SDD) receiver with only 5% sacrifice in area efficiency. Huiyu Feng, Suwen Song, Chuan Zhang 0001, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Variational Bayesian Message Passing Receiver for Uplink ISAC Systems
Tiancan Xia, Jian Zheng 0003, Xiaosi Tan, Yongming Huang 0001, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Commun. | 6 |
| 2026 | MrcPunc: Modulation and Rate Compatible Puncturing for Few-Iteration 5G LDPC DecodingabstractTo meet the stringent high-throughput demands of enhanced mobile broadband (eMBB), low-density parity-check (LDPC) decoding is limited to few iterations. This constraint usually degrades error-correcting performance, rousing a critical need for puncturing towards algorithmic enhancements. However, existing puncturing methods lack modulation and rate compatibility, converging to local optima due to greedy search strategies. To this end, this paper proposes a modulation and rate compatible puncturing (MrcPunc) for 5G LDPC codes. It incorporates three key techniques: 1)a revised protograph-based extrinsic information transfer (PEXIT) analysisfor accurate performance prediction, 2)a progressive multi-path search (PMPS) algorithmto avoid local optima, and 3)a compatibility optimizationsupporting various modulations and code rates. The MrcPunc yields up to 0.6 dB gain over the 5G standard puncturing under few-iteration condition without extra decoding complexity. It is noted that MrcPunc guarantees robustness against both larger iteration counts and mismatched modulations, alongside reducing the average number of decoding iterations. Qiushi Xu, Huayi Zhou 0002, Mingyang Zhu, Ming Jiang 0012, Chuan Zhang 0001 |
IEEE Trans. Commun. | 5 |
| 2026 | TIP: Turbo Implicit Pursuit Channel Estimator for mmWave MIMO SystemsabstractCompressed-sensing (CS)-based channel estimation is a promising technology for future millimeter wave (mmWave) multiple-input–multiple-output (MIMO) systems, enabling significant pilot reduction and improved estimation accuracy. Channel estimators based on matching pursuit (MP) variants offer lower complexity compared with other CS algorithms, but suffer from high latency due to their iterative nature, hindering efficient hardware implementation. To mitigate this issue, this article introduces a turbo pursuit (TP) strategy that relaxes sequential dependencies in MP variants, enabling parallel processing and pipelined implementation. To demonstrate the effectiveness of TP, this article further introduces turbo implicit pursuit (TIP), a hardware-friendly instance of TP that leverages a prioritized gradient descent (GD) strategy for low-complexity least squares (LS) solving. A hardware auto-generator for TIP is then proposed using a formula representation approach, which constructs a parameterized hardware-algorithm design space and enables hardware-algorithm co-optimization. Our optimized$32 \times 4$TIP MIMO channel estimator ASIC in 65-nm CMOS achieves 0.53-$\mu $s latency under 18.75% measurements. Compared with prior implementations of MP variants, this work achieves higher or comparable estimation accuracy with over 18% latency reduction and over 5$\times$higher throughput to area ratio (TAR) for ASICs, and over 90% latency reduction with over 3$\times$higher hardware efficiency for FPGAs. Changhan Li, Xingchi Zhang, Yutai Sun, Yunwei Mao, Yifang Dai, You You, Yongming Huang 0001, Chuan Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2026 | FCBsP: Fixed-Constellation Belief-Selective Propagation Detection for MIMO Turbo ReceiversabstractThe belief-selective propagation (BsP) algorithm has recently emerged as a promising approach for massive MIMO detection. However, when applied in MIMO turbo receivers, known for their superior performance compared to separated detection and decoding (SDD) receivers, the BsP-based receiver suffers from significant performance degradation and high processing latency. To overcome these limitations, this paper proposes a fixed-constellation BsP (FCBsP) detector tailored for MIMO turbo receivers. By buildingfixed configuration setsand utilizing theapproximate multi-user interferencefor message updates, the proposed FCBsP detector achieves a better trade-off between error performance and computational complexity compared to the BsP. Furthermore, two unexplored features:information compensation and decoding-first mechanismare proposed to fine-tune the exchanged information and lower the processing latency of the FCBsP-based turbo receiver. Numerical results demonstrate that the proposed FCBsP-based turbo receiver earns about 0.7 and 1.8 dB performance gains over the BsP-based turbo receiver at BLER=10−3in an LDPC-coded 32 × 12 64-QAM MIMO system under Rayleigh and practical channels, respectively. Zeqiong Tan, Wenyue Zhou, Kefan Wang, Yongming Huang 0001, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2026 | Efficient Energy Efficiency Optimization Method for Cell-Free Massive MIMO-Enabled URLLC Downlink SystemsabstractThis paper investigates the downlink energy efficiency (EE) optimization for cell-free massive multiple input multiple output (CF-mMIMO) systems subject to ultra-reliable and low-latency communication (URLLC) requirements. To achieve superior performance, we jointly consider the impacts of power allocation, access point (AP)-user association, and AP sleep modes under the finite blocklength (FBL) regime, leading to a challenging mixed-integer (MI) non-convex optimization problem. Utilizing a sequential convex approximation (SCA) framework, we first propose the SCA-Relaxation algorithm to convert the original problem into a series of second-order cone programming (SOCP) sub-problems, which can be efficiently addressed via modern convex programming solvers. Moreover, for further reducing computational complexity, we approximate the original problem as a continuous-variable optimization and tackle it via a combination of the Dinkelbach transformation, penalty functions, as well as an accelerated proximal gradient method with adaptive momentum, resulting in the proposed low complexity EE maximization (LCEE-max) algorithm. Besides, the related convergence and complexity analysis of these two algorithms are also presented in detail. Simulation results demonstrate that compared to the state-of-the-art baseline algorithm, the proposed two algorithms achieve the EE improvements of approximately 40% and 30%, respectively, along with a substantial reduction in complexity, thereby enabling efficient and fast resource allocation in CF-mMIMO-enabled URLLC scenarios. Zheng Wang 0013, Amin Sakzad, Chuan Zhang 0001, Yongming Huang 0001, Derrick Wing Kwan Ng |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | Belief Propagation Decoding for Short Codes on Structured Sparse Parity-Check MatricesabstractAs successfully adopted in standard long code scenarios, belief propagation (BP) decoding has been considered a promising universal decoding candidate for next-generation wireless communications. However, when applied to short codes, BP decoding suffers from poor error correction performance due to harmful cycle structures in the Tanner graph. In this paper, we address this issue by designing a structured, sparse parity-check matrix (ssPCM) framework, composed of multiple cycle-free parity-check row blocks (PCRBs). The resulting ssPCMs feature regular row weights and perform better than the state-of-theart 4 -cycle-free row redundant PCMs across Bose-Chaudhuri-Hocquenghem (BCH) codes of length 63. Yifei Shen 0003, Zongyao Li 0003, Emmanuel Boutillon, Wenqing Song, Yuqing Ren, Chuan Zhang 0001, Xiaohu You 0001, Andreas Peter Burg |
ISIT | 6 |
| 2025 | Fast construction and exploration of performance-cost design space for belief propagation polar decoders
You You, Weikang Qian, Yongming Huang 0001, Chuan Zhang 0001 |
Sci. China Inf. Sci. | 5 |
| 2025 | Toward mobile communication baseband circuit auto-design: a Bayesian model approach
Chuan Zhang 0001, Changhan Li, Yunwei Mao, Yuwei Zeng, You You, Yongming Huang 0001, Xiaohu You 0001 |
Sci. China Inf. Sci. | 1 |
| 2025 | A 40 µs latency cell-free mmWave reliable transmission experimental system via spatiotemporal 2-D coding
Xiaohu You 0001, Dongming Wang 0002, Chuan Zhang 0001, Pengcheng Zhu 0001, Jiamin Li 0001, Bin Kuang, Qinji Jiang |
Sci. China Inf. Sci. | 4 |
| 2025 | Toward Universal Belief Propagation Decoding for Short Binary Block CodesabstractBelief propagation (BP) decoding has been recognized for its capacity-approaching performance and high throughput when decoding long low-density parity-check (LDPC) codes. However, the application of BP decoding for short codes is hindered by dense parity-check matrices (PCMs) and prevalent short cycles in the Tanner graph. In this paper, we introduce a general method to extract an optimized sparse PCM for short binary block codes, which removes length-four cycles and enhances the connectivity of short cycles to enable BP decoding with improved performance. Notably, for short binary codes with lengths up to 64, our BP decoding performance approaches the maximum likelihood bound and surpasses the best-reported BP results with reduced computational complexity. Compared with other universal decoding algorithms, BP decoding using our extracted sparse PCMs is competitive in terms of both error-rate performance and computational complexity. These promising results suggest that our method to improve BP decoding for short codes is a step toward a practical universal BP decoder for next-generation communication systems. Yifei Shen 0003, Zongyao Li 0003, Yuqing Ren, Emmanuel Boutillon, Alexios Balatsoukas-Stimming, Chuan Zhang 0001, Xiaohu You 0001, Andreas Peter Burg |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | A RISC-V Domain-Specific Processor for Deep Learning-Based Channel EstimationabstractChannel estimation (CE) is a critical component in the massive multi-input multi-output (MIMO) communication systems. Compared with conventional CE algorithms, deep learning (DL)-based approach becomes a promising alternative, due to its capability of offering enhanced performance and robustness across diverse scenarios. However, efficient DL-based CE algorithms have two key properties that make them challenging for implementation in existing architectures at the edge side: the diversity of deep neural networks (DNNs) and CE strategies, and the involvements of multiple computation-intensive tasks that compass conventional signal processing, artificial intelligence (AI) inference, and online learning. To address these challenges, a domain-specific processor based on an extended RISC-V instruction set architecture (ISA) is proposed to perform these DL-based CE algorithms. First, a dedicated RISC-V ISA extension is developed to support all essential operations required by a DL-based CE algorithm, such as matrix inversion, in a flexible manner. Building on the customized ISA extension, a highly adaptable and scalable RISC-V processor is developed, featuring scalar and vector posit arithmetic units to alleviate high computational and memory demands of DNNs during both inference and training phase. Additionally, a coarse-grained matrix accelerator is integrated to expedite various matrix operations ensuring high throughput. In this way, both high flexibility and computational efficiency are achieved. Finally, our processor is implemented on a TSMC 28-nm technology. Implementation results show that the processor achieves a speedup of$5.16\sim 6.80\times $for all matrix operations compared with the state-of-the-art work. Moreover, the proposed processor provides an area efficiency improvement of$1.61\times $and an energy efficiency enhancement of$6.6\sim 15.4\times $compared to the open-source vector processor Ara. Notably, this work is the first RISC-V domain-specific processor tailored for diverse DL-based CE algorithms. Chuanning Wang, Yangcan Zhou, Shaowei Wang 0001, Chuan Zhang 0001, Zhongfeng Wang 0001, Jun Lin 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | UniDec: A Unified Factor-Graph-Based Decoder Fully Compatible With 5G NR LDPC/Polar CodesabstractIn comparison to 4G, 5G wireless needs to support a broader range of applications. Therefore, both low-density parity-check (LDPC) codes and polar codes have been standardized by 5G new radio (NR) to fulfill the requirements of data channel and control channel, respectively. Usually, LDPC/polar decodings are implemented by separate hardware, leading to low area efficiency. Though decoders which can handle both codes have been proposed, how to compromise between throughput and efficiency has always been a persistent dilemma due to the absence of a unified and smooth integration methodology. To this end, by fully utilizing the common parts of graph-theoretic algorithms for both codes, this paper presents a unified decoder (UniDec) which is fully compatible with 5G NR LDPC/polar codes. This UniDec enables three key approaches:1) unified processing nodes for both codes,2) configurable permutation networks with multi-parallelism, and3) flexible scheduling for 5G NR parameter configuration, guaranteeing both high data throughput and area efficiency. Implemented in 40nm CMOS, the UniDec attains a maximum of$33.64\times $throughput and$5.98\times $area efficiency compared to its multi-mode counterparts. Even compared with the state-of-the-art (SOA) dedicated ones, the UniDec still maintains a competitive edge in terms of throughput, energy, and area efficiency. It is noted that this methodology can be generalized to other factor-graph based signal processing algorithms. Houren Ji, Yutai Sun, Yongming Huang 0001, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | Explicit Performance Bound of Finite Blocklength Coded MIMO: Time-Domain Versus Spatiotemporal Channel CodingabstractIn the sixth generation (6G), ultra-reliable low-latency communications (URLLC) will be further developed to achieve TKμ extreme connectivity. On the premise of ensuring the same rate and reliability, the spatial domain advantage of multiple-input multiple-output (MIMO) has the potential to further shorten the time-domain code length and is expected to be a key enabler for the realization of TKμ. Different coded MIMO schemes exhibit disparities in exploiting the spatial domain characteristics, so we consider two extreme MIMO coding schemes, namely, time-domain coding in which the codewords on multiple spatial channels are independent of each other, and spatiotemporal coding in which multiple spatial channels are jointly coded. By analyzing the statistical characteristics of information density and utilizing the normal approximation, we provide explicit performance bounds for finite blocklength coded MIMO under time-domain coding and spatiotemporal coding. It is found that, different from the phenomenon in time-domain coding where the performance degrades as the blocklengths decrease, spatiotemporal coding can effectively compensate for the performance loss caused by short blocklengths by improving the spatial degrees of freedom (DoF). These results indicate that spatiotemporal coding can fully exploit the spatial dimension advantages of MIMO systems, enabling extremely low error-rate communication under stringent blocklengths constraint. Feng Ye 0001, Xiaohu You 0001, Jiamin Li 0001, Jinni Chen, Chuan Zhang 0001 |
IEEE Trans. Commun. | 5 |
| 2025 | Adaptive Channel Estimation for RIS-Assisted Systems in Time-Varying mmWave ChannelsabstractTo improve channel estimation (CE) for reconfigurable intelligent surface (RIS)-assisted systems in time-varying mmWave channels, this paper proposes a two-stage adaptive CE scheme. This is the first attempt to develop a CE scheme without assumptions on specific timescales for channel variations. In the first stage, the adaptive scheme incorporates the estimation of partial channel state information and a channel status check process. The introduced check process can monitor the changing status of the channels and provide information for the second stage. In the second stage, based on the results from the check process, the adaptive scheme adaptively selects from two proposed candidate CE algorithms: Two-Phase orthogonal matching pursuit (TP-OMP) and Structured-Shift OMP (SS-OMP). Simulation results show that both TP-OMP and SS-OMP can reduce pilot overhead by around 33%, and respectively lower the computational complexity of existing works by about 55% and 65%. Additionally, the check process obtains an accuracy rate of approximately 92% so that the proposed CE scheme can maintain stable CE performance in time-varying channels. You You, Fengyu Chen, Li Zhang 0011, Yongming Huang 0001, Chuan Zhang 0001 |
IEEE Trans. Commun. | 5 |
| 2025 | Low-Complexity Breadth-First Search Detection for Large-Scale MIMO SystemsabstractThanks to its near-optimal performance, breadth-first search detection (BFSD) finds widespread application in small-scale MIMO systems. However, existing BFSD methods struggle to effectively configure the width (number of candidate nodes) for each layer, resulting in prohibitive complexity in large-scale MIMO systems. To address this, we propose two width optimization schemes for BFSD. We introduce a layer-by-layer optimization framework to reduce the design space of width configurations, and a Monte Carlo-assisted method to link width configurations to detection performance. Using this linking scheme in the reduced design space, we formulate the first width optimization scheme given specific performance constraints. Then, we present another scheme that employs a theoretical linking method as an alternative to the Monte Carlo approach. Although slightly less effective, the second scheme has negligible complexity for width optimization, making it well-suited for communication scenarios with time-varying characteristics. In 128×128 MIMO systems, numerical results demonstrate that the optimized BFSD using our first and second schemes can reduce complexity by up to 82% and 65%, respectively, while achieving superior detection performance compared to state-of-the-art BFSD. Jian Zheng 0003, Yutai Sun, Huayi Zhou 0002, Wenyue Zhou, Yongming Huang 0001, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Commun. | 7 |
| 2025 | Efficient ORBGRAND Implementation With Parallel Noise Sequence GenerationabstractGuessing random additive noise decoding (GRAND) is establishing itself as a universal method for decoding linear block codes, and ordered reliability bits GRAND (ORBGRAND) is a hardware-friendly variant that processes soft-input information. In this work, we propose an efficient hardware implementation of ORBGRAND that significantly reduces the cost of querying noise sequences with slight frame error rate (FER) performance degradation. Different from logistic weight order (LWO) and improved LWO (iLWO) typically used to generate noise sequences, we introduce a reduced-complexity and hardware-friendly method called shift LWO (sLWO), of which the shift factor can be chosen empirically to trade the FER performance and query complexity well. To effectively generate noise sequences with sLWO, we utilize a hardware-friendly lookup-table (LUT)-aided strategy, which improves throughput as well as area and energy efficiency. To demonstrate the efficacy of our solution, we use synthesis results evaluated on polar codes in a 65-nm CMOS technology. While maintaining similar FER performance, our ORBGRAND implementations achieve 53.6-Gbps average throughput ($1.26\times $higher), 4.2-Mbps worst case throughput ($8.24\times $higher), 2.4-Mbps/mm2 worst case area efficiency ($12\times $higher), and$4.66\times 10 ^{{4}}$pJ/bit worst case energy efficiency ($9.96\times $lower) compared with the synthesized ORBGRAND design with LWO for a (128, 105) polar code and also provide$8.62\times $higher average throughput and$9.4\times $higher average area efficiency but$7.51\times $worse average energy efficiency than the ORBGRAND chip for a (256, 240) polar code, at a target FER of$10^{-7}$. Xiaohu You 0001, Chuan Zhang 0001, Christoph Studer |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Stochastic Belief Propagation-Based Iterative Detection and Decoding for MIMO SystemsabstractIn this brief, a stochastic belief propagation (BP)-based iterative detection and decoding (IDD) for multiple-input and multiple-output (MIMO) system is proposed. We modify the algorithm of BP detection to make it more suitable for stochastic computation and enable the soft message to be transmitted between the detector and decoder in the format of stochastic sequences. Through IDD, the required number of iterations and quantization precision for the detector will decrease. By sharing the stochastic number generator, the hardware complexity of both the detector and decoder can be reduced. Hardware architectural optimizations and the corresponding implementation are also given, and we can implement 64 × 32, four-QAM MIMO system with (128, 64) polar codes with 1.283mm2area consumption. Compared with other detector, the hardware efficiency can be improved by 7.8 times. Muhao Li, Houren Ji, Xiaosi Tan, Chuan Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | A Soft Iterative Receiver With Simplified EP Detection for Coded MIMO SystemsabstractExpectation propagation (EP) achieves excellent performance with high-order modulation in massive multiple-input multiple-output (MIMO) detection. The soft output of the EP detector can be iteratively combined with turbo soft decoders to enhance error-correction performance. However, the implementation of EP-based iterative detection and decoding (IDD) receivers suffer from an exponential increase in computational complexity as the number of antennas and modulation order grows. In this brief, we propose a simplified EP approximation-based IDD (sEPA-IDD) scheme for hardware implementation. To alleviate the computational burden, a simplified message update scheme is proposed, reducing complexity by 68% without performance degradation. Additionally, a unified design for extrinsic message computation further improves hardware utilization. Finally, we introduce the first unfolded EP-based IDD architecture to boost throughput. Compared with state-of-the-art (SOA) IDD receivers, the sEPA-IDD receiver implemented on 65 nm CMOS delivers a throughput of 3.07 Gb/s with a maximum 0.5 dB gain, achieving 4.03× higher throughput and 6.04× greater area efficiency. Xiaosi Tan, Xiaohua Xie, Houren Ji, Tiancan Xia, Yongming Huang 0001, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | A 4.86-pJ/b Energy-Efficient Fully Parallel Stochastic LDPC Decoder With Two-Stage Shared MemoryabstractThe complex calculations of the low-density parity-check (LDPC) decoder result in significant energy and hardware consumption. To solve the challenge, this brief describes a fully parallel stochastic LDPC decoder with a two-stage shared memory (TSM) variable node (VN). To enhance cost efficiency, our design incorporates a shared low-cost random number generator (RNG) for all 2160 channels. We introduce a TSM VN function, which demonstrates faster convergence and reduced hardware overhead in comparison with the existing methods. We have taped out the (2160, 1760) stochastic LDPC decoder in the 55-nm process. The measure results exhibit that the proposed design achieves a throughput of 57.6 Gb/s, an efficiency of 33.68 Gb/s/mm2, and a power efficiency of 4.86 pJ/bit, underlining superior performance in terms of decoding throughput, hardware efficiency, and energy conservation. Yakun Zhou, Jienan Chen, Yizhuo Zhou, Zihan Xia 0002, Chuan Zhang 0001, Runsheng Wang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | ISPT-Net: A Noval Transient Backward-Stepping Reduction Policy by Irregular Sequential Prediction TransformerabstractIn the post-layout simulation for large-scale integrated circuits, transient analysis (TA), determining the time-domain response over a specified time interval, is essential and important. However, it tends to be computationally intensive and quite time-consuming without proper settings of NR initial solution and accurate LTE estimation for determining the next transient timestep, which will lead to a mass of backward-steppings. In this paper, an irregular sequential prediction transformer named ISPT-Net is proposed to predict accurately transient solution as NR initial solution and further obtain precise LTE estimation for setting next timestep. The ISPT-Net is strengthened with timestep positional encoding module (TPE), frequency- and timestep-sensitive muti-head self-attention module (FT-MSA) to enhance irregular sequence feature extraction and prediction accuracy. We assess ISPT-Net in the real large-scale industrial circuits on a commercial SPICE simulator, and achieve a remarkable backward stepping reduction: up to 14.43X for NR nonconvergence case and 4.46X for LTE overlimit case while guaranteeing higher solution accuracy. Yichao Dong, Dan Niu, Zhou Jin 0001, Chuan Zhang 0001, Changyin Sun 0001, Zhenya Zhou |
DATE | 4 |
| 2024 | Code Length Compatible Belief Propagation Polar Decoder Based on Folding and UnfoldingabstractThis paper presents a code-length compatible architecture for belief propagation (BP) polar decoders. This decoder incorporates folding and unfolding techniques with control signals, allowing it to decode codes with varying code lengths. By modifying the architecture originally designed for code length N, the proposed decoder can handle codes of length 2iN, where i ∈ Z+using folding, and i ∈ Z−using unfolding. To reduce the critical path and implementation complexity, a new routing design is proposed. Moreover, we introduce a memory architecture utilizing shift registers instead of RAM to increase the throughput. We demonstrate gate-level implementations to illustrate the design’s architecture. Finally, we analyze the throughput, area, and power consumption of the decoders. Compared with traditional single-column designs with N = 1024, the proposed decoder architecture can achieve up to 136% hardware efficiency while consuming 1.6% less area. Muhao Li, Huizheng Wang, Yifei Shen 0003, Xiaosi Tan, Chuan Zhang 0001 |
ISCAS | 5 |
| 2024 | FLAG: Formula-LLM-Based Auto-Generator for Baseband Hardware
Yunwei Mao, You You, Xiaosi Tan, Yongming Huang 0001, Xiaohu You 0001, Chuan Zhang 0001 |
ISCAS | 6 |
| 2024 | A Low-Latency and High-Performance SCL Decoder with Frame-InterleavingabstractIn this paper, we describe a frame-interleaving hardware architecture for a generalized node-based successive cancellation list (SCL) decoder. By efficiently reusing otherwise idle computational units, two independent frames can be decoded simultaneously, resulting in a significant throughput gain. Based on this new architecture, we also exploit graph ensembles to diversify the decoding, enhancing the error-correcting performance by 0.28 dB and reducing the worst-case latency for serial graph processing by over 32%. Implementation results show that the proposed SCL decoder with frame-interleaving architecture achieves a throughput of 7.15 Gbps and an area efficiency of 37.63 Gbps/mm2, which is 1.56× and 1.11× better than the state-of-the-art node-based SCL decoders. Leyu Zhang, Yuqing Ren, Yifei Shen 0003, Wuyang Zhou, Alexios Balatsoukas-Stimming, Chuan Zhang 0001, Andreas Peter Burg |
ISCAS | 6 |
| 2024 | A Node-Based Polar List Decoder With Frame Interleaving and Ensemble Decoding SupportabstractNode-based successive cancellation list (SCL) decoding has received considerable attention in wireless communications for its significant reduction in decoding latency, particularly with 5G New Radio (NR) polar codes. However, the existing node-based SCL decoders are constrained by sequential processing, leading to complicated and data-dependent computational units that introduce unavoidable stalls, reducing hardware efficiency. In this paper, we present a frame-interleaving hardware architecture for a generalized node-based SCL decoder. By efficiently reusing otherwise idle computational units, two independent frames can be decoded simultaneously, resulting in a significant throughput gain. Based on this new architecture, we further exploit graph ensembles to diversify the decoding space, thus enhancing the error-correcting performance with a limited list size. Two dynamic strategies are proposed to eliminate the residual stalls in the decoding schedule, which eventually results in nearly$2 \times $throughput compared to the state-of-the-art baseline node-based SCL decoder. To impart the decoder rate flexibility, we develop a novel online instruction generator to identify the generalized nodes and produce instructions on-the-fly. The corresponding 28nm FD-SOI ASIC SCL decoder with a list size of 8 has a core area of 1.28 mm2 and operates at 692 MHz. It is compatible with all 5G NR polar codes and achieves a throughput of 3.34 Gbps and an area efficiency of 2.62 Gbps/mm2 for uplink (1024, 512) codes, which is$1.41 \times $and$1.69 \times $better than the state-of-the-art node-based SCL decoders. Yuqing Ren, Leyu Zhang, Ludovic Damien Blanc, Yifei Shen 0003, Alexios Balatsoukas-Stimming, Chuan Zhang 0001, Andreas Peter Burg |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | Approximate Belief-Selective Propagation Detector for Massive MIMO SystemsabstractWhen faced with challenging antenna configurations or high-order modulations in realistic propagation environments, the Belief Propagation (BP) MIMO detector outperforms its linear counterparts. To mitigate the error floor issue and lower the complexity, a revised BP detector, named the Belief-selective Propagation (BsP) detector, has recently emerged by selectively utilizing trusted incoming messages for updates. Despite those promising potentials, the straightforward hardware implementation of the BsP detector still suffers from high complexity and necessitates further optimization. To bridge the gap between the BsP algorithm and implementation, this paper introduces anapproximatebut implementation-friendly BsP detector called aBsP, based on which the very first BsP hardware is proposed. Two unexplored features:approximate initializationandsimplified message updatessave the complexity (more than$84$%) with acceptable performance penalization. Multi-level optimization techniques involving group-layered message updating, approximate arithmetic circuits, and hybrid-precise quantization are developed to boost the hardware efficiency A$128\times 8$$256$-QAM aBsP MIMO detector ASIC in$40$nm CMOS occupies an area of$0.68$mm$^2$and reaches a throughput of$790.52$Mbps. Benchmarking with the recent arts, this work achieves$1.08\times$area efficiency and$3.34\times$gate efficiency. Wenyue Zhou, Zhenhao Ji, Zeqiong Tan, Zhuangzhuang You, Xiaosi Tan, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | Analysis of a New Energy-Efficient Model for Future Wireless Communication SystemsabstractEnergy efficiency (EE) is currently one of the primary concerns in wireless communication systems. In future 6G systems, achieving high energy-efficient communication with T-bit transmission rate requirements will heavily depend on resource and power parameters. This paper proposes a novel energy efficiency model that incorporates wireless resources and power for a multi-user multiple-input-multiple-output (MIMO) wireless communication system. Firstly, we analyze the equivalency between frequency and spatial dimensions and introduce a two-dimensional resource domain. Additionally, we propose an energy efficiency model incorporating the wireless resource and power. Using the model, we analyze the interaction between the two parameters aiming at the maximize energy efficiency and present an energy-efficient algorithm. Furthermore, We propose two methods for judging high energy efficiency communication of each user based on the model. Finally, our numerical results illustrate the performance of energy efficiency for each user and the advantages and disadvantages of each measurement method in the multi-user system. Kang Liu 0021, Zaichen Zhang, Chuan Zhang 0001, Jian Dang, Liang Wu 0001, Bingcheng Zhu, Lei Wang 0182 |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Implementation of 6G TKμ Extreme Connectivity via Cell-Free Massive MIMO System: A Theoretical EvaluationabstractThe key performance indicators (KPIs) of the sixth generation (6G) will increase by orders of magnitude compared to the fifth generation (5G), promising extreme connectivity performance with Tbps-scale data rate, Kbps/Hz-scale spectral efficiency (SE) and$\mu \text {s}$-level latency. Cell-free massive MIMO (CF-mMIMO) with rich spatial dimension resources is expected to be a key architecture to realize$\text {TK}\mu $extreme connectivity, but the existing research has not yet given a compact and closed-form approximation to describe the relationship between the spatial dimension and system performance, which makes it difficult to evaluate the KPIs of$\text {TK}\mu $intuitively. This paper derives explicit closed-form expressions for the relationship between system performance and system configuration parameters for finite blocklength CF-mMIMO systems and analyzes the relationship between system performance and spatial dimensions. Based on this, we perform parameter selection and performance evaluation of specific implementations in the three$\text {TK}\mu $KPIs in CF-mMIMO systems. Both theoretical analysis and simulation results show that increasing the spatial degree of freedom (DoF) and deploying antennas more dispersedly can realize latency reduction while guaranteeing the system performance, and the joint collaboration of multi-users and multiple access points (APs) with large DoFs can achieve a continuous increase in SE and data rate. Feng Ye 0001, Xiaohu You 0001, Jiamin Li 0001, Chuan Zhang 0001, Pengcheng Zhu 0001, Dongming Wang 0002, Yongming Huang 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Improved Belief Propagation Decoding of Turbo CodesabstractTurbo codes have been successfully adopted in 4G LTE, which can approach the channel capacity with Bahl-Cocke-Jelinek-Raviv (BCJR) decoding. With the evolution from 4G LTE to 5G NR, there is a demand to design a unified channel decoder that supports both LTE Turbo codes and NR low-density parity-check (LDPC) codes. One solution is to employ belief propagation (BP) decoding on the bipartite Tanner graph for both codes. However, although MacKay pointed out that Turbo codes have a sparse parity-check matrix, the existence of 4-cycles in such a matrix severely deteriorates the performance of BP decoding. In this paper, we propose two polynomial-based methods to optimize the parity-check matrix of Turbo codes by improving the sparsity while also removing 4-cycles and even 6-cycles compared to the original matrix. Simulation results show that the improved BP decoding for Turbo codes halves the error-correction performance gap between the original BP decoding and BCJR decoding, which is a promising step towards the unified channel decoder design based on the BP algorithm. Yifei Shen 0003, Yuqing Ren, Andreas Toftegaard Kristensen, Xiaohu You 0001, Chuan Zhang 0001, Andreas Peter Burg |
ICASSP | 5 |
| 2023 | Ultra-wideband fiber-THz-fiber seamless integration communication system toward 6G: architecture, key techniques, and testbed implementation
Jiao Zhang 0005, Bingchang Hua, Mingzheng Lei, Yuancheng Cai, Dongming Wang 0002, Wei Xu 0001, Chuan Zhang 0001, Yongming Huang 0001, Jianjun Yu, Xiaohu You 0001 |
Sci. China Inf. Sci. | 9 |
| 2023 | OSSP-PTA: An Online Stochastic Stepping Policy for PTA on Reinforcement LearningabstractThe dc analysis is essential and still quite challenging in large-scale nonlinear circuit simulation. Pseudo transient analysis (PTA) is a widely used and has great potential solver in the industry. However, the PTA convergence and simulation efficiency is still seriously affected by its stepping policy. This article proposes an online stochastic stepping policy (OSSP) for PTA based on deep reinforcement learning (DRL). To achieve better policy evaluation and stronger stepping exploration ability, the dual soft Actor–Critic agents work with the proposed valuation splitting and online momental scaling, enabling our OSSP to intelligently encode PTA iteration status and online further adjust forward and backward time-step size for unseen test circuits without human intervention and domain knowledge, trained solely by reinforcement learning from self-search. Our public sample buffer and priority sampling are also introduced to overcome the sparsity and imbalance of sample data. Numerical examples demonstrate that the proposed OSSP achieves a significant efficiency speedup (up to$47.0\times $less Newton–Raphson iterations) and convergence enhancement on unseen test circuits compared with the previous iter-based and switched evolution/relaxation-based stepping methods, in just one stepping iteration. Dan Niu, Yichao Dong, Zhou Jin 0001, Chuan Zhang 0001, Changyin Sun 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | A DSP-Purposed REconfigurable Acceleration Machine (DREAM) for High Energy Efficiency MIMO Signal ProcessingabstractThe wireless baseband processing algorithms are still developing and show a great diversity. The development of ASIC implementations cannot quickly adapt to the evolution of algorithms and standards. Meanwhile, the general-purpose processors cannot meet the real-time requirements in some scenarios. This paper proposes a DSP-purposed REconfigurable Acceleration Machine (DREAM) core for wireless baseband digital signal processing, which has a good trade-off between flexibility and performance. First, we abstract a set of shared operators with a moderate granularity from a variety of wireless MIMO signal processing algorithms. Then, we propose a two-step configuration process to reduce the size of the required reconfiguration bits. Besides, we design a conflict-free address generator to transfer data between the on-chip scratchpad memory and reconfiguration processing elements with high efficiency and high throughput. Finally, the prototype DREAM core has been implemented in TSMC CMOS 28 nm, and its area and power consumption have been analyzed. The chip has great flexibility in supporting a variety of wireless MIMO processing algorithms and a wide range of MIMO scales. The proposed DREAM core can achieve the normalized area efficiency and the normalized energy efficiency of$0.67~Gbps/MGE$and$15.05~Gbps/W$, which are$1.56\times $and$4.18\times $those of state-of-the-art reconfigurable implementations when running the WeJi-based MIMO detection algorithm. Kai Chen 0034, Wenqing Song, Guoqiang He, Sirui Shen, Huizheng Wang, Chuan Zhang 0001, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | An Efficient Approximate Expectation Propagation Detector With Block-Diagonal Neumann-SeriesabstractExpectation propagation (EP) achieves near-optimal performance for large-scale multiple-input multiple-output (L-MIMO) detection, however, at the expense of unaffordable matrix inversions. To tackle the issue, several low-complexity EP detectors have been proposed. However, they all fail to exploit the properties of channel matrices, thus resulting in unsatisfactory performance in non-ideal scenarios. To this end, in this paper, a block-diagonal Neumann-series-based expectation propagation approximation (BD-NS-EPA) algorithm is proposed, which is applicable for both ideal uncorrelated channels and the correlated channels with multiple-antenna user equipment system. First, a block-diagonal-based Neumann iteration is employed, which skillfully exerts the main information of the channels while reducing computational cost. An adjustable sorting message updating scheme then is introduced to reduce the update of redundant nodes during iterations. Numerical results show that, for$128\times 32$MIMO with the non-ideal channel, the proposed algorithm exhibits 0.3 dB away from the original EP when bit error-rate (BER)$=10^{-3}$, at the cost of mere 3% normalized complexity. The implementation results on SMIC 65-nm CMOS technology suggest that the proposed detector can achieve 1.252 Gbps/W and 0.275 Mbps/kGE hardware efficiency, further demonstrating that the proposed detectors can achieve a good trade-off between error-rate performance and hardware efficiency. Huizheng Wang, Bingyang Cheng, Xiaosi Tan, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Closed-Form Approximation for Performance Bound of Finite Blocklength Massive MIMO TransmissionabstractIt is supposed that ultra-reliable low latency communication (uRLLC) would continue to evolve in the future sixth generation (6G) network, to provide enhanced capability towards extreme connectivity, with the aid of well established multiple-input multiple-output (MIMO) technology. Since the latency constraint can be represented equivalently by the blocklength of a codeword, channel coding theory at a finite blocklength plays an important role in theoretic analysis of uRLLC. Based on Polyanskiy’s and Yang’s asymptotic results on maximal achievable rate, we first derive the proximate closed-form expressions for the expectation and variance of channel dispersion. Then, the upper bound of average maximal achievable rate is obtained for massive MIMO systems under ideal independent and identically distributed fading channels. Since almost all the fundamental parameters, including the spatial degree-of-freedom (DoF), are considered, this expression can be viewed as a performance bound of the spatiotemporal two-dimension channel coding to some extent. Moreover, it is shown by simulation and analysis, as the DoF goes to infinity, MIMO systems reveal a nature of deterministic transmission, since the average maximal achievable coding rate per antenna can be achieved at each transmission. In this case, the inversely proportional law observed therein implies that the blocklength in the time domain can be further shortened at the expense of spatial DoF. This exchangeability of space and time, to support a given coding rate, paves a solid and feasible road for us to further reduce latency in 6G uRLLC. Xiaohu You 0001, Bin Sheng 0003, Yongming Huang 0001, Wei Xu 0001, Chuan Zhang 0001, Dongming Wang 0002, Pengcheng Zhu 0001 |
IEEE Trans. Commun. | 5 |
| 2023 | Belief-Selective Propagation Detection for MIMO SystemsabstractCompared to the linear MIMO detectors, the Belief Propagation (BP) detector has shown greater capabilities in achieving near-optimal performance and better nature to iteratively cooperate with channel decoders. Aiming at real applications, recent works mainly fall into the category of reducing the complexity by simplified calculations, at the expense of performance sacrifice. However, the complexity is still unsatisfactory with exponentially increasing complexity or required exponentiation operations. Furthermore, the state-of-the-art (SOA) BP detectors persistently encounter error floor in high signal-to-noise ratio (SNR) region, which becomes even worse with calculation approximation. This work aims at a revised BP detector, named Belief-selective Propagation (BsP) detector by selectively utilizing the trusted incoming messages with sufficiently large a priori probabilities for updates. Two proposed strategies: symbol-based truncation (ST) and edge-based simplification (ES) squeeze the complexity (orders lower than the BP detector), while greatly relieving the error floor issue over a wide range of antenna and modulation combinations. For the 256-QAM$128 \times 64$uplink massive multiuser MIMO (MU-MIMO) system, the$\mathcal {B}(1,1)$BsP detector achieves more than 1dB performance gain (@$\text {BER}=10^{-4}$) with lower complexity than the state-of-the-art (SOA) BP detector. Trade-off between performance and complexity towards different application requirements can be conveniently obtained by tuning the parameters of the ST and ES strategies. Wenyue Zhou, Yifei Shen 0003, Liping Li 0001, Yongming Huang 0001, Chuan Zhang 0001, Xiaohu You 0001 |
IEEE Trans. Commun. | 5 |
| 2022 | Fast Sequence Repetition Node-Based Successive Cancellation List Decoding for Polar CodesabstractCompared with the bit-wise successive cancellation list (SCL) decoding of polar codes, the node-based Fast SCL decoding significantly reduces the decoding latency by identifying special constituent codes and decoding these in parallel. To further reduce the latency of current Fast SCL decoders, we first propose a fast sequence repetition (SR) node-based SCL (Fast SR-SCL) decoding algorithm, which only involves one type of node in the SCL decoding tree. Furthermore, we employ the adaptive path splitting (APS) strategy to terminate the path splitting in the SR node early, without degrading the error-correcting performance. Numerical results show that for 5G uplink codes with a length of 1024 and rates of 1/4, 1/2, and 3/4, our decoder can deliver the same decoding performance while reducing the average latency by 34.5%, 38.0%, and 39.6% compared with the state-of-the-art Fast SCL decoder for a list size L = 8. Yifei Shen 0003, Yuqing Ren, Andreas Toftegaard Kristensen, Alexios Balatsoukas-Stimming, Xiaohu You 0001, Chuan Zhang 0001, Andreas Peter Burg |
ICC | 6 |
| 2022 | Efficient polar coding scheme and implementation with shared information bits
Wenyue Zhou, Yifei Shen 0003, Liping Li 0001, Chuan Zhang 0001 |
Sci. China Inf. Sci. | 6 |
| 2022 | Joint Channel Estimation and Data Detection in Cell-Free Massive MU-MIMO SystemsabstractWe propose a joint channel estimation and data detection (JED) algorithm for densely-populated cell-free massive multiuser (MU) multiple-input multiple-output (MIMO) systems, which reduces the channel training overhead caused by the presence of hundreds of simultaneously transmitting user equipments (UEs). Our algorithm iteratively solves a relaxed version of a maximum a-posteriori JED problem and simultaneously exploits the sparsity of cell-free massive MU-MIMO channels as well as the boundedness of QAM constellations. In order to improve the performance and convergence of the algorithm, we propose methods that permute the access point and UE indices to form so-called virtual cells, which leads to better initial solutions. We assess the performance of our algorithm in terms of root-mean-squared-symbol error, bit error rate, and mutual information, and we demonstrate that JED significantly reduces the pilot overhead compared to orthogonal training, which enables reliable communication with short packets to a large number of UEs. Haochuan Song, Tom Goldstein, Xiaohu You 0001, Chuan Zhang 0001, Olav Tirkkonen, Christoph Studer |
IEEE Trans. Wirel. Commun. | 4 |
| 2021 | Adaptive Successive Cancellation Priority Decoder for 5G Polar CodesabstractAs two common successive cancellation (SC)-based decoding algorithms of polar codes, the SC list (SCL) and SC stack (SCS) decoder can achieve satisfactory error correction performance, especially with increased list size or stack depth. Nevertheless, a large list size or stack depth will lead to high computational complexities and hardware resources. To this end, successive cancellation priority (SCP) decoding with priority- first searching strategy and trellis-like storage is proposed to offer one solution. In this paper, an efficient SCP decoder is first proposed to verify its advantages over SCL and SCS decoders. Furthermore, an adaptive node-inserting scheme is proposed to reduce the number of bits insert into the priority queue. Numerical results have shown that for the polar code with transmission length 1024 and rate 1/2, the proposed adaptive SCP (ASCP) decoder can achieve significant time complexity reduction on average compared with the standard SCL decoder. The hardware architecture of SCP decoding is implemented using 65-nm CMOS technology and the results show better throughput compared with the SCS decoder. Wenqing Song, Yifei Shen 0003, Chuan Zhang 0001, Li Li 0003 |
ISCAS | 4 |
| 2021 | Live Demonstration: A Cloud-Based Cell-Free Distributed Massive MIMO SystemabstractThis is the demonstration description for a cloud- based cell-free distributed massive MIMO system, which is based on our already published work. Distributed massive MIMO antennas will result in design and implementation challenges regarding synchronization, calibration, real-time baseband processing, and so on. The contributors propose a cloud-based cellfree distributed massive MIMO system, which cannot only meet 5G NR requirements but also can be easily extended to different application scales. For this demostration, a 128 × 128 distributed MU-MIMO system with frequency 100 [email protected] GHz is given. Test results show that 10.185 Gbps throughput and more than 100 bps/Hz spectrum utilization can be obtained. On-site applications such as HD video transmission and virtual reality (VR) are offered for visitor experiences and interacts. Dongming Wang 0002, Chuan Zhang 0001, Zhenhao Ji, Yongqiang Du, Ming Jiang 0012, Xiaohu You 0001 |
ISCAS | 2 |
| 2021 | Efficient Fast-SCAN Flip Decoder for Polar CodesabstractSoft-output decoder is of great importance to be applied in iterative receivers, of which belief propagation (BP) algorithm has been widely studied for 5G low-density parity- check (LDPC) and polar codes. However, for polar codes, BP decoding suffers from high computational complexity and unsatisfactory convergence. To this end, soft cancellation (SCAN) polar decoder has recently drawn attention from academia and can be further improved by using the bit-flipping strategy. Limited by the serial nature of message propagation, the SCAN flip (SCANF) decoder cannot meet a high throughput. In this paper, we accelerate the decoding speed by the fast processing mechanism, conducting Fast-SCANF decoder. The corresponding hardware architecture is designed with memory optimization and implemented by TSMC 40nm technology, delivering a 2.1 Gbps throughput and 65 pJ/b energy. To the knowledge of authors, this is the first SCANF hardware decoder. Leyu Zhang, Yutai Sun, Yifei Shen 0003, Wenqing Song, Xiaohu You 0001, Chuan Zhang 0001 |
ISCAS | 6 |
| 2021 | Soft-Output Joint Channel Estimation and Data Detection using Deep UnfoldingabstractWe propose a novel soft-output joint channel estimation and data detection (JED) algorithm for multiuser (MU) multiple-input multiple-output (MIMO) wireless communication systems. Our algorithm approximately solves a maximum a-posteriori JED optimization problem using deep unfolding and generates soft-output information for the transmitted bits in every iteration. The parameters of the unfolded algorithm are computed by a hyper-network that is trained with a binary cross entropy (BCE) loss. We evaluate the performance of our algorithm in a coded MU-MIMO system with 8 basestation antennas and 4 user equipments and compare it to state-of-the-art algorithms separate channel estimation from soft-output data detection. Our results demonstrate that our JED algorithm outperforms such data detectors with as few as 10 iterations. Haochuan Song, Xiaohu You 0001, Chuan Zhang 0001, Christoph Studer |
ITW | 3 |
| 2021 | Implementation of a concentration-controlled chemical clock
Chongzhou Fang, Lulu Ge, Xiaosi Tan, Ziyuan Shen, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
Sci. China Inf. Sci. | 7 |
| 2021 | Orbital angular momentum multiplexing communication system over atmospheric turbulence with K-best detection
Yingmeng Ge, Liang Wu 0001, Chuan Zhang 0001, Zaichen Zhang |
Sci. China Inf. Sci. | 3 |
| 2021 | A correlation-breaking interleaving of polar codes in concatenated systems
Ya Meng, Liping Li 0001, Chuan Zhang 0001 |
Sci. China Inf. Sci. | 3 |
| 2021 | Towards 6G wireless communication networks: vision, enabling technologies, and new paradigm shiftsabstractAbstract The fifth generation (5G) wireless communication networks are being deployed worldwide from 2020 and more capabilities are in the process of being standardized, such as mass connectivity, ultra-reliability, and guaranteed low latency. However, 5G will not meet all requirements of the future in 2030 and beyond, and sixth generation (6G) wireless communication networks are expected to provide global coverage, enhanced spectral/energy/cost efficiency, better intelligence level and security, etc. To meet these requirements, 6G networks will rely on new enabling technologies, i.e., air interface and transmission technologies and novel network architecture, such as waveform design, multiple access, channel coding schemes, multi-antenna technologies, network slicing, cell-free architecture, and cloud/fog/edge computing. Our vision on 6G is that it will have four new paradigm shifts. First, to satisfy the requirement of global coverage, 6G will not be limited to terrestrial communication networks, which will need to be complemented with non-terrestrial networks such as satellite and unmanned aerial vehicle (UAV) communication networks, thus achieving a space-air-ground-sea integrated communication network. Second, all spectra will be fully explored to further increase data rates and connection density, including the sub-6 GHz, millimeter wave (mmWave), terahertz (THz), and optical frequency bands. Third, facing the big datasets generated by the use of extremely heterogeneous networks, diverse communication scenarios, large numbers of antennas, wide bandwidths, and new service requirements, 6G networks will enable a new range of smart applications with the aid of artificial intelligence (AI) and big data technologies. Fourth, network security will have to be strengthened when developing 6G networks. This article provides a comprehensive survey of recent advances and future trends in these four aspects. Clearly, 6G with additional technical requirements beyond those of 5G will enable faster and further communications to the extent that the boundary between physical and cyber worlds disappears. Xiaohu You 0001, Cheng-Xiang Wang 0001, Jie Huang 0004, Xiqi Gao 0001, Zaichen Zhang, Michael Mao Wang, Yongming Huang 0001, Chuan Zhang 0001, Yanxiang Jiang, Jiaheng Wang 0001, Bin Sheng 0003, Dongming Wang 0002, Zhiwen Pan, Pengcheng Zhu 0001, Yang Yang 0001, Zening Liu, Ping Zhang 0003, Xiaofeng Tao 0001, Shaoqian Li, Zhi Chen 0002, Xinying Ma, Chih-Lin I, Shuangfeng Han, Chengkang Pan, Zhiming Zheng 0001, Lajos Hanzo, Xuemin Shen, Y. Jay Guo, Zhiguo Ding 0001, Harald Haas, Wen Tong, Peiying Zhu, Ganghua Yang, Jue Wang 0006, Erik G. Larsson, Hien Quoc Ngo, Wei Hong 0002, Haiming Wang 0001, Debin Hou, Jixin Chen, Zhe Chen 0021, Zhangcheng Hao, Geoffrey Ye Li, Rahim Tafazolli, Yue Gao 0001, H. Vincent Poor, Gerhard P. Fettweis, Ying-Chang Liang |
Sci. China Inf. Sci. | 8 |
| 2021 | Optimizing Vertical Link Placement and Congestion Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCsabstractThe fully connected 3D-NoCs in which all routers are vertically connected with their neighbors above and below need a lot of Through-Silicon-Vias (TSVs), and they will occupy a large silicon area and reduce the fabrication yield. Thus, the idea of partially connected 3D-NoCs has emerged. The optimal number and placement of the vertical links (elevators) must be determined at the chip design stage, which is a multiobjective optimization problem of the performance and the cost. However, optimizing the static elevator placement needs a great amount of calculation and we can not examine all possible solutions at design time. Therefore, we propose a hybrid heuristic strategy for the static elevator placement and assignment, in which the genetic algorithm and the tabu search are combined. The dynamic assignment method is essential for the partially connected 3D-NoCs, and it leads to different traffic distributions and therefore has a huge impact on performance. Many previous static assignment methods can not dynamically change the elevator assignment according to the real-time states of the network, thus it may lead to network congestion. A congestion-aware dynamic assignment (CDA) scheme is proposed in this article, which considers the impact of the distance factor and the congestion factor on the network performance. Experiments show that the proposed CDA method can improve the network performance by 67%-86% compared with the random selection algorithm and can improve the reliability of the partially connected 3D-NoC as well. The key component for the CDA method, the path selection module (PSM), is implemented in FPGA, and the results show that its area cost is negligible compared with a router. Chuan Zhang 0001, Wenqing Song, Qinyu Chen, Hui Chen 0015, Li Li 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Hardware Implementation for Belief Propagation Flip Decoding of Polar CodesabstractBelief propagation (BP) decoding has natural advantages in throughput for polar codes to meet high-speed and low-latency requirements. The soft outputs of BP decoding can be utilized further for joint detection and decoding in the baseband communication system. However, its error-correction performance is not comparable with the successive cancellation list (SCL) decoding. Belief propagation flip (BPF) decoding is recently proposed to improve the error-correction performance of BP decoding and indicates the potential to compete with SCL decoding. In this paper, we propose an advanced BPF (A-BPF) scheme that reduces the decoding latency with the help of one critical bit and improves the error-correction performance by the proposed joint detection criterion. To improve area efficiency in the hardware level, an optimized sorting network is proposed and applied for the A-BPF decoder. The decoder is implemented on 65 nm CMOS technology for length-1024 and rate-1/2 polar codes, and the results show that the proposed decoder can achieve a close frame error rate performance to the SCL decoder with four lists and deliver a throughput of 5.17 Gb/s at Eb/N0= 4.0 dB. Houren Ji, Yifei Shen 0003, Wenqing Song, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Efficient Row-Layered Decoder for Sparse Code Multiple AccessabstractSparse code multiple access (SCMA) is a promising technology for the development of wireless communication, which supports a large number of overloading users and enjoys high spectral efficiency. However, conventional SCMA decoders suffer very high complexity in implementations. Changing the updating scheme is a superior approach to reduce complexity, which guarantees the updated information immediately join in the following message propagating of the current iteration and accelerates the decoding convergence. In this paper, a row-layered message passing algorithm (MPA) is proposed, which offers a good trade-off between the hardware complexity and the bit error rate (BER) performance. Simulation results show that the proposed decoder saves 66.7% computation complexity compared with the original MPA with the similar BER performance. Pipelining and folding technology are adopted in VLSI implementations. The synthesis results with 45-nm CMOS technology show that the proposed decoder can achieve higher hardware efficiency and throughput under a high frequency than the existing decoders, achieving 1777.78 Mb/s throughput with 1.112 mm2area consumption. Xu Pang, Wenqing Song, Yifei Shen 0003, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | Efficient Soft-Output Gauss-Seidel Data Detector for Massive MIMO SystemsabstractFor massive multiple-input multiple-output (MIMO) systems, linear minimum mean-square error (MMSE) detection has been shown to achieve near-optimal performance but suffers from excessively high complexity due to the large-scale matrix inversion. Being matrix inversion free, detection algorithms based on theGauss–Seidel(GS) method have been proved more efficient than conventionalNeumannseries expansion-based ones. In this paper, an efficient GS-based soft-output data detector for massive MIMO and a corresponding VLSI architecture are proposed. To accelerate the convergence of the GS method, a new initial solution is proposed. Several optimizations on the VLSI architecture level are proposed to further reduce the processing latency and area. Our reference implementation results on a Xilinx Virtex-7 XC7VX690T FPGA for a 128 base-station antenna and eight user massive MIMO system show that our GS-based data detector achieves a throughput of 732 Mb/s with close-to-MMSE error-rate performance. Our implementation results demonstrate that the proposed solution has advantages over the existing designs in terms of complexity and efficiency, especially under challenging propagation conditions. Chuan Zhang 0001, Zhizhen Wu, Christoph Studer, Zaichen Zhang, Xiaohu You 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Bipartite Belief Propagation Polar Decoding With Bit-FlippingabstractFor the scenarios with high throughput requirements, the belief propagation (BP) decoding is one of the most promising decoding strategies for polar codes. By pruning the redundant variable nodes (VNs) and check nodes (CNs) in the original factor graph, the graph is condensed to a sparse bipartite graph which is similar to the graph for low-density parity-check (LDPC) codes. In this paper, we introduce the bit-flipping scheme into the LDPC-like BP (L-BP) decoding and propose two methods to identify the error-prone VNs. By additional decoding attempts, the L-BP flip (L-BPF) decoding improves the error-rate performance with a similar average complexity for high Eb=N0values. The simulation results show that the L-BPF decoding achieves 0:25 dB gain compared with the L-BP decoding. Zihao Gong, Yifei Shen 0003, Houren Ji, Wenqing Song, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ICASSP | 7 |
| 2020 | Efficient stochastic successive cancellation list decoder for polar codes
Xiao Liang 0005, Huizheng Wang, Yifei Shen 0003, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
Sci. China Inf. Sci. | 6 |
| 2020 | Molecular computing for Markov chains
Chuan Zhang 0001, Ziyuan Shen, Zaichen Zhang, Xiaohu You 0001 |
Nat. Comput. | 1 |
| 2020 | Mathematical Modeling Analysis of Strong Physical Unclonable FunctionsabstractPhysical unclonable function (PUF) is a technique to produce secret keys or complete authentication in integrated circuits (ICs) by exploiting the uncontrollable randomness due to manufacturing process variations. For better PUF applications, efficient analysis of different designs is important. In this article, a mathematical model to analyze the performance of typical strong PUF designs is proposed and applied to arbiter PUF, ring oscillator (RO) PUF, and duty cycle (DC) PUF. For better reliability, a new PUF design, DC multiplexer (DC MUX) PUF proposed in our previous work is analyzed. The proposed model indicates that DC MUX PUF achieves 2% higher reliability than arbiter PUF under environment influences. It also shows that DC PUF achieves 10% higher reliability than RO PUF. For verification, the aforementioned four PUF designs are testified using HSPICE. As our model analysis indicates, for reliability DC MUX PUF outperforms arbiter PUF, and DC PUF outperforms RO PUF. For randomness, DC MUX PUF and DC PUF outperform arbiter PUF and RO PUF, respectively. For security, LR attacks on DC MUX PUF and arbiter PUF are performed. The training time for DC MUX PUF is 40 000 times of arbiter PUF. Yunhao Xu, Yingjie Lao, Weiqiang Liu 0001, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | Reconfigurable and Low-Complexity Accelerator for Convolutional and Generative Networks Over Finite FieldsabstractConvolutional neural networks (CNNs) have gained great success in various fields, such as computer vision and natural language processing. Besides, with the breakthrough in unsupervised learning, generative adversarial network (GAN) is recently utilized to generate virtual data from limited data sets. The generative model of GAN has impressive applications, such as style transfer and image super-resolution. However, the promising performance of CNN and GAN comes at the cost of prohibitive computation complexity. The convolution (CONV) in CNN and the transposed CONV (TCONV) in GAN are the two operations that dominant the overall complexity. The prior works exploit the fast algorithms, Winograd and fast Fourier transform (FFT), to reduce the complexity of spatial CONV. However, Winograd only supports fixed filter size while FFT has high transform overhead. Moreover, very few works apply fast algorithms to accelerate GAN models. In this article, a reconfigurable and low-complexity accelerator on ASIC for both CNN and GAN is proposed to address these problems. First, by exploiting Fermat number transform (FNT), we propose two FNT-based fast algorithms to reduce the complexity of CONV and TCONV computations, respectively. Then the architectures of the FNT-based accelerator are presented to implement the proposed fast algorithms. The methodology to determine the design parameters and optimize the dataflow is also described for obtaining maximum performance and optimal efficiency. Moreover, we implement the proposed accelerator on 65 nm 1P9M technology and evaluate it on various CNN and GAN models. The post-layout results show that our design achieves a throughput of 288.0 GOP/s on VGG-16 with 25.11 GOP/s/mm2area efficiency, which is superior to the state-of-the-art CNN accelerators. Furthermore, at least $1.7\times $ speedup over the existing accelerators is obtained on GAN. The resulting energy efficiency is $275.3\times $ and $12.5\times $ of CPU and GPU. Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Improved Belief Propagation Polar Decoders With Bit-Flipping AlgorithmsabstractSince the inherent serial nature of successive cancellation list (SCL) decoding results in a long latency, belief propagation (BP) decoding for polar codes has drawn attention for high-throughput applications. However, its error correction performance is inferior to that of SCL decoding. Therefore, the bit-flipping strategy has been recently applied to BP decoding, which can approach the SCL decoding performance through multiple additional decoding attempts. The original BP flip (BPF) decoding suffers from an inaccurate identification of erroneous bits by a fixed flip set (FS), which has been improved by the generalized BPF (GBPF) decoding. In this article, the GBPF decoding is extended to support multiple bits being flipped in one decoding attempt. In addition, for two types of decoding errors: detected errors and undetected errors, we propose two novel methods to more effectively identify erroneous bits. For detected errors, the concept of loop sets is defined and a loopbased identification method is introduced based on the study of error patterns of BP decoding. On the other hand, a method to generate a more accurate fixed FS is proposed for undetected errors, which considers the bit error distribution under BP decoding. Combining the two methods, the GBPF with merged sets (GBPF-MS) decoding can achieve the SCL-8 performance and outperforms the state-of-the-art BPF, BP list, and SC flip (SCF) decoding, for polar codes with length 1024 and information rate 1/2. Implemented by 40nm CMOS technology, the proposed GBPF-MS decoder with ten flips exhibits an average throughput of 4.19 Gbps at 2.5 dB, which is 1.6× and 1.72× faster than the state-of-the-art SCL-4 and SCF decoders, respectively. Yifei Shen 0003, Wenqing Song, Houren Ji, Yuqing Ren, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Commun. | 7 |
| 2020 | Autogeneration of Pipelined Belief Propagation Polar DecodersabstractThough belief propagation (BP) polar decoders can achieve higher throughput than successive-cancellation (SC)-based decoders, and how to efficiently generate different belief propagation decoders (BPDs) which can meet various design specifications remains challenging. To this end, an autogeneration, which can translate the generation formula of BPDs to efficient hardware implementations, has been proposed in this article. For different requirements, two BPD architectures have been given: 1) low-cost decoder (Type-I) and 2) high-throughput decoder (Type-II). The autogeneration of them can support different code rates, code lengths, and parallelisms. Synthesis results show that Type-I and Type-II provide higher throughput and hardware efficiency than the state-of-the-art (SOA) SC decoders. Moreover, compared to the SOA BPDs, both Type-I and Type-II achieve similar even better energy- and area-efficiency with a comparable throughput, for fully parallel configuration. With the autogeneration, we are able to obtain the design space regarding different design metrics, such as area efficiency, energy efficiency, and power density, within which the design optimization under given design constraints can be conducted. Yifei Shen 0003, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2020 | A General Construction and Encoder Implementation of Polar CodesabstractPuncturing and shortening are two general ways to obtain an arbitrary code length and code rate for polar codes. When some of the coded bits are punctured or shortened, it is equivalent to a situation in which the underlying channels of polar codes are different. This fact calls for a general polar code construction, which is not yet available. In this article, a general construction of polar codes is studied in two aspects: 1) the theoretical foundation of the general construction and 2) the hardware implementation of general polar codes encoders. In contrast to the original identical and independent binary-input, memoryless, symmetric (BMS) channels, these underlying BMS channels can be different. The proposed general construction of polar codes is based on the existing Tal-Vardy's procedure. The symmetric property and the degradation relationship are shown to be preserved under the general setting, rendering the possibility of a modification of Tal-Vardy's procedure. Simulation results clearly show improved error performance with reordering using the proposed new procedures. Also, a novel encoding hardware architecture is proposed, which supports puncturing and shortening modes. Implementation results show the proposed encoder achieves approximately 30% throughput improvement when one quarter of bits are punctured/shortened. Yifei Shen 0003, Liping Li 0001, Kai Niu 0001, Chuan Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2019 | Efficient Belief Propagation Detection Based on Channel Hardening for Massive MIMOabstractFor massive multiple-input multiple-output (MIMO) detection, belief propagation (BP) based on graphical models has become a popular detection algorithm since it provides a good tradeoff between performance and complexity. To further lower the complexity of BP detection, an efficient BP detection based on channel hardening (BP-CH) is proposed. In this paper, the comparison in terms of both performance and complexity between proposed BP-CH and general BP is firstly investigated exhaustively. Simulation results have shown that the proposed BP-CH achieves similar performance behavior as general BP while keeping lower computational complexity. Additionally, an folded hardware architecture for proposed BPCH detector is designed to improve the implementation efficiency. Meanwhile, VLSI implementation results have verified the great advantage of BP-CH regarding hardware overhead, especially for scenarios with large system loading factor. Shusen Jing, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ICASSP | 5 |
| 2019 | Smilodon: An Efficient Accelerator for Low Bit-Width CNNs with Task PartitioningabstractConvolutional Neural Networks (CNNs) have been widely applied in various fields such as image and video recognition, recommender systems, and natural language processing. However, the massive size and intensive computation loads prevent its feasible deployment in practice, especially on the embedded systems. As a highly competitive candidate, low bit-width CNNs are proposed to enable efficient implementation. In this paper, we propose Smilodon, a scalable, efficient accelerator for low bit-width CNNs based on a parallel streaming architecture, optimized with a task partitioning strategy. We also present the 3D systolic-like computing arrays fitting for convolutional layers. Our design is implemented on Zynq XC7Z020 FPGA, which can satisfy the needs of real-time with a frame rate of 1, 622 FPS throughput, while consuming 2 1 Watt. To the best of our knowledge, our accelerator is superior to the state-of-the-art works in the tradeoff among throughput, power efficiency, and area efficiency. Qinyu Chen, Kaifeng Cheng, Wenqing Song, Zhonghai Lu, Li Li 0003, Chuan Zhang 0001 |
ISCAS | 7 |
| 2019 | Congestion-Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCsabstractThe combination of Network-on-Chips (NoCs) and 3D IC technology, 3D NoCs, has been proven to be able to achieve a great improvement in both network performance and power consumption compared to 2D NoCs. In the traditional 3D NoC, all routers are vertically connected. Due to the large overhead of Through-Silicon-Via (TSV, e.g., low fabrication yield and the occupied silicon area), the partially connected 3D NoC has emerged. The assignment method determines the traffic loads of the vertical links (elevators), thus has a great impact on 3D-NoCs' performance. In this paper, we propose a congestion-aware dynamic elevator assignment (CDA) scheme, which takes both the distance factors and network congestion information into account. Experiments show that the performance of the proposed CDA scheme is improved by 67% to 87% compared to the random selection scheme, 8% to 25% compared to SelByDis-1, and 13% to 18% compared to SelByDis-2. Qinyu Chen, Guoqiang He, Kai Chen 0034, Zhonghai Lu, Chuan Zhang 0001, Li Li 0003 |
ISCAS | 6 |
| 2019 | Multi-Incentive Delay-Based (MID) PUFabstractThis paper proposes a new PUF, namely Multi-incentive Delay-based PUF (MID PUF), which utilizes the fast carry logic (FCL) of Field Programmable Gate Arrays (FPGAs). The proposed MID PUF is completely and efficiently implemented in XOR gates of FCLs. Compared to other single signal excited PUF designs, e.g. Arbiter PUF, multiple excitations are applied on the same delay line to produce multiple outputs. To the authors' best knowledge, this is the first strong PUF based on only FCLs. The proposed MID PUF is implemented on Xilinx Spartan-6 XC6SLX9 FPGAs and a reliability experiment is carried out under the operating temperature in a range of 0°C~70° C. The experimental results show that the proposed MID PUF has a high uniqueness and reliability performance, as well as low hardware consumption. Due to its advantages in both hardware efficiency and PUF metrics, the proposed MID PUF is promising for low-cost security applications on FPGAs. Zhengran Zhang, Chongyan Gu, Yijun Cui, Chuan Zhang 0001, Máire O'Neill, Weiqiang Liu 0001 |
ISCAS | 4 |
| 2019 | AI for 5G: research directions and paradigms
Xiaohu You 0001, Chuan Zhang 0001, Xiaosi Tan, Shi Jin 0002, Hequan Wu |
Sci. China Inf. Sci. | 2 |
| 2019 | DNA computing for combinational logic
Chuan Zhang 0001, Lulu Ge, Yuchen Zhuang, Ziyuan Shen, Zaichen Zhang, Xiaohu You 0001 |
Sci. China Inf. Sci. | 1 |
| 2019 | Thermal Sensor Placement and Thermal Reconstruction Under Gaussian and Non-Gaussian Sensor Noises for 3-D NoCabstractOn-chip thermal sensors are essential for temperature management in 3-D network-on-chip (NoC) systems. However, due to the physical (area and power) or economical constraints, the number of sensors is limited. Therefore, the two critical issues we face are: 1) how to figure out an efficient thermal sensor placement with the limited number of sensors and 2) how to reconstruct the entire thermal profile based on sensor observations. Another major issue for the thermal reconstruction is the sensor measurement accuracy. Thus, online accurate full-chip thermal reconstruction under Gaussian and non-Gaussian noises is another great challenge. In this paper, a greedy thermal sensor placement algorithm maximizing the rank of the observability Gramian is proposed. A good placement algorithm always relies on a specific reconstruction method. The proposed placement algorithm is designed for the state-space-based thermal model, thus the combination of the proposed placement algorithm and the Kalman filter-based reconstruction method provides a high reconstruction accuracy under Gaussian noise. For accurate temperature reconstruction under non-Gaussian noise, the Gaussian-Sum filter is applied to 3-D NoC. Compared with the Kalman filter, the Gaussian-Sum filter can reduce the root-mean-squared-error and the max error by 29.27%–35% and 33.26%–40.6%, respectively. A reusable architecture for the Kalman filter and the Gaussian-Sum filter has been proposed. Its hardware implementation details are presented in this paper. Besides, the performance and the area are evaluated as well. Li Li 0003, Hongbing Pan, Kun Wang 0005, Qinyu Chen, Chuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2019 | Efficient Successive Cancellation Stack Decoder for Polar CodesabstractAs an improved version of successive cancellation (SC) polar decoder, an SC stack (SCS) decoder has been proposed for performance improvement. However, the existing SCS polar decoder suffers a lot from high time complexity at low signal-to-noise ratio (SNR) region and space complexity compared with the SC decoder. To this end, two improved decoders are proposed to reduce time and space complexity in both low and high SNR regions. The first one is the segmented cyclic redundancy check (CRC)-aided SCS (SCA-SCS) decoder, which is based on segmented parity checkers. The second one is the adaptive SCS (ASCS) decoder, which has the flexibility of stack depth and searching width. Furthermore, a channel condition estimator is proposed to select appropriate decision criteria for different SNR scenarios. Results have shown that for the polar code of length 1024 and rate 1/2, two improved SCS decoders can perform better than the traditional SCS decoder. The proposed SCA-SCS decoder and the ASCS decoder can achieve 10.8% and 11.42% time complexity reduction and 31.68% and 60.85% space complexity reduction on average over binary-input additive white Gaussian noise channels (BI-AWGNCs), respectively. Efficient parallel hardware architecture of the SCS polar decoder is first proposed and implemented with 90- and 65-nm technologies. Results have verified its advantages over the state of the art (SOA). Wenqing Song, Huayi Zhou 0002, Kai Niu 0001, Zaichen Zhang, Li Li 0003, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2018 | Joint List Polar Decoder with Successive Cancellation and Sphere DecodingabstractFor polar codes, both successive cancellation list (SCL) decoding and list sphere decoding (LSD) aim to balance performance and complexity. The same list structure but different decoding schedules of SCL and LSD can lead to a combination of both schemes. In this paper, an efficient joint list decoder with SCL and LSD (JLSCD) is proposed to reduce time complexity. We apply SCL and LSD schemes simultaneously but independently, then merge them at the middle point of the decoding. Numerical results have demonstrated JLSCD scheme's advantage in complexity. FPGA implementation of JLSCD decoder is also given in this paper. Xiao Liang 0005, Huayi Zhou 0002, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ICASSP | 5 |
| 2018 | Approximate Belief Propagation Decoder for Polar CodesabstractPolar code is increasing its popularity recently for its capacity-achieving property for B-DMCs. However, when designing decoders for polar code, it has always been an inevitable concern for us to balance the decoding performance and the hardware consumption. In this paper, we propose an approximate belief propagation (BP) decoder for polar code for the first time. By introducing the approximate computation schemes, we reduced the critical path delay (CPD) and the hardware consumption of the conventional BP decoders. Simulation results show that the proposed approximate BP decoder achieves nearly the same decoding performance as the conventional one. Advantages of the proposed decoder has been verified by FPGA implementation. Menghui Xu, Shusen Jing, Jun Lin 0001, Weikang Qian, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ICASSP | 7 |
| 2018 | Efficient Deep Convolutional Neural Networks Accelerator without Multiplication and RetrainingabstractRecently, low-precision weight method has been considered as a promising scheme to efficiently implement inference of deep convolutional neural networks (DCNN). But it suffers from expensive retraining cost and accuracy degradation. In this paper, a low-bit and retraining-free quantization method, which enables DCNNs to deal inference with only shift and add operations, is proposed. The efficiency is demonstrated in terms of power consumption and chip area. Huffman coding is adopted for further compression. Then by exploring two-level systolic, an efficient hardware accelerator is introduced with respect to the given quantization strategy. Experiment results show that our method achieves higher accuracy than other low-precision networks without retraining process on ImageNet. 5× to 8× compression is obtained on popular models compared to full-precision counterparts. Furthermore, hardware implementation indicates good reduction of slices whereas maintaining throughput. Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ICASSP | 4 |
| 2018 | Efficient Circulant Matrix Construction and Implementation for Compressed SensingabstractThe design of measurement matrices is an important part in compressed sensing (CS). Random matrices superior to incoherence are considered to be optimal measurement matrices to achieve successful recovery. However, they are deficient in memory cost. Structure matrices like circulant matrices are preferred for low-memory cost. Nevertheless, their recovery performance is greatly damaged because of element coherence. In this paper, a new method called different-spaced selection & different-spaced flipping (DSS & DSF) is proposed to modify structure matrices. Based on circulant matrices, regular extraction and symbol flipping imposed on columns of measurement matrices can increase randomness to a large scale. As a result, not only near optimal recovery but also much less memory cost can be achieved. Compared with Gaussian random matrices, the memory cost can be reduced to 4% when measurement matrices based on circulant matrices are in 128 × 512 dimensions. An efficient hardware design and VLSI implementation are also presented at the end of this paper. Feng Yi, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ICASSP | 4 |
| 2018 | Basic Arithmetics Based on Analog Signal with Molecular ReactionsabstractThis paper presents a design methodology to implement basic arithmetics, including addition, subtraction, multiplication and division, with molecular reactions using analog signals. Simpler than conventional digital logic, our design can still implement the same functionality. Two kinds of designs are contained, one is based on synchronous sequential logic, and the other basis is an asynchronous one. The feasibility of our method is validated via simulations of chemical kinetics. Muhao Li, Lulu Ge, Xiaohu You 0001, Chuan Zhang 0001 |
ICC | 4 |
| 2018 | Implementation of Sinusoids and Pulse Width Modulation with Chemical ReactionsabstractThis paper offers an implementation of a sinusoid and its pulse-width modulation (PWM) based on chemical reaction networks (CRNs). The generation of a sinusoid is given under the guidance of ordinary differential equations (ODEs) combined with the non- negative concentration of chemical molecules. The accuracy of the sinusoid is verified by listing the ODEs. Comparing the concentrations of the sampled sinusoid and sawtooth wave, the target PWM could be finally synthesized. This process requires a sampler and a comparator. Based on the obtained PWM, nearly any arbitrary analog signal could be converted to a digital one. Lulu Ge, Xiaohu You 0001, Chuan Zhang 0001 |
ICC | 4 |
| 2018 | Synthesizing LDPC Belief Propagation Decoding with Molecular ReactionsabstractThis paper proposes a CRN-based implementation approach for low-density parity-check (LDPC) decoding based on belief propagation (BP). Since the belief (probability) can be naturally mapped to molecule concentration, LDPC decoding can be realized with CRNs instead of silicon based hardware. Theoretical analysis and numerical simulations have demonstrated the feasibility of the proposed approach. Note that, we do not try to substitute the silicon-based LDPC decoder with CRN- based one for high-speed applications. We show that this method can be generalized for other BP-based algorithms and is suitable for large-scale, bio- interface, and latency-insensitive applications. Xingchi Zhang, Lulu Ge, Xiaohu You 0001, Chuan Zhang 0001 |
ICC | 4 |
| 2018 | Implementation of Mealy Machine with Molecular ReactionsabstractChemical reaction networks(CRNs) have been used as a formal language for constructing and analysing molecular systems. Theoretical analysis and experiments have demonstrated that CRNs could be physically implemented with DNA strand displacement reactions. This paper proposes a method of synthesizing Mealy machine with molecular reactions, which focuses on deriving CRN from state diagram of any Mealy machine. All components of the proposed design could be experimentally implemented by DNA reactions. Therefore, the proposed CRN-based Mealy machine has potential applications in \(vitro\) and in \(vivo\) biotechnology. Zhen Li 0038, Lulu Ge, Xiaohu You 0001, Chuan Zhang 0001 |
ICC | 5 |
| 2018 | Reconfigurable Decoder for LDPC and Polar CodesabstractWith low-density parity-check (LDPC) code and polar code selected as the standard codes for 5G eMBB scenario, one challenge is how to improve the hardware efficiency when both decoders are required by one system. Since LDPC and polar codes can be decoded with belief propagation (BP) algorithms, this similarity allows us to design a reconfigurable decoder, which can decode both codes at the cost of only one decoder. Numerical and implementation results are also given in this paper to show that the proposed decoder achieves higher hardware efficiency than stand-alone LDPC or polar decoder, without harming the error performance. Ningyuan Yang, Shusen Jing, Anlan Yu, Xiao Liang 0005, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ISCAS | 7 |
| 2018 | A Channel-Blind Detection for SCMA Based on Image Processing TechniquesabstractSparse-code multiple-access (SCMA) is an effective non-orthogonal multiple-access (NOMA) technique. Existing detectors such as deterministic message passing algorithm (DMPA) are one-dimensional and require precise channel estimation. This paper proposes a blind detector from a two-dimensional perspective. Main work involves pattern construction and pre-filtering with different image techniques. In this paper, a 4 × 4 Sudoku template is applied for the pattern construction of one-dimensional SCMA signals. Total variation based on first order differential operator is adopted for global pre-filtering of DMPA. Image training is adopted in DMPA to further reduce the environment noise. The output signal of both pre-filtering methods are detect through DMPA with constant noise density N0. Numerical results show that two-dimensional blind detection can well compensate the performance when channel estimation of DMPA is not perfect. A general hardware architecture of the detecting method is also proposed in this paper. Chao Yang 0027, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ISCAS | 5 |
| 2018 | Polar-coded forward error correction for MLC NAND flash memory
Haochuan Song, Jen-Chien Fu, Shih-Jia Zeng, Jin Sha 0001, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
Sci. China Inf. Sci. | 7 |
| 2017 | Joint Detection and Decoding for Polar Coded MIMO SystemsabstractGenerally, separate detection and decoding (SDD) scheme is usually adopted by multiple-input and multiple-output (MIMO) systems. In this paper, a novel approach which combines detection and decoding jointly using K-best detection and polar codes is proposed for the first time. Since the generation matrix of polar codes is triangular, polar codes could well adapt to the structure of K-best searching tree. Moreover, the property of polarization could reduce the latency of the proposed joint detection and decoding (JDD) scheme. Based on the joint optimization, the system model is given. For successive cancellation list (SCL) polar decoding, numerical results show that the performance of the proposed JDD is superior to the state-of-the-art SDD. At the frame error rate (FER) of 10-4, JDD outperforms SDD by approximately 2.5 dB for (256,128) polar coded 4×4 16-QAM MIMO system. Furthermore, for half rate polar codes, the proposed JDD could reduce 50% complexity compared to SDD. Results indicate that JDD shows superiorities in both performance and complexity. In addition, the corresponding hardware architectures are also given to demonstrate JDD's advantages and implementation feasibilities. Yifei Shen 0003, Junmei Yang, Xiaohu You 0001, Chuan Zhang 0001 |
GLOBECOM | 5 |
| 2017 | A DNA strand displacement reaction implementation-friendly clock designabstractTo address the inherent limits of silicon-based technologies, the research on synthesizing various logic functions with chemical reaction networks (CRNs) has emerged in large numbers. However, in order to properly synthesize a given sequential logic, the difficulties lie in constructing a clock signal with an arbitrary duty cycle of M/N. Therefore, this paper is dedicated to putting forward a CRN-based design methodology, which can generate clock signals with an arbitrary duty cycle. Only unimolecular or bimolecular reactions are employed, which makes the real DNA strand displacement reactions successfully compiled from formal CRNs. First, a clock signal with 1/2 duty cycle is constructed. Then, the systematic design flow to construct an arbitrary M/N duty cycle clock is explained in details. Conditions are different when N is odd or even. All the proposed design methods come along with a theoretical basis and a numerical validation. Donglin Wen, Lulu Ge, Chuan Zhang 0001, Xiaohu You 0001 |
ICC | 4 |
| 2017 | Algorithm and architecture for joint detection and decoding for MIMO with LDPC codesabstractWith better spectral efficiency, multiple-input and multiple-output (MIMO) systems have drawn increasing attentions. Due to its near-optimal performance, K-best algorithm has been widely adopted for MIMO detection. To the best knowledge of the authors, this paper first proposes a joint detection and decoding (JDD) method for MIMO with low-density parity-check (LDPC) codes. By pruning the searching tree of K-best detection with LDPC coding constraint, the proposed JDD scheme benefits from both reduced tree-search complexity and improved performance compared to its uncoded MIMO counterpart. Numerical results of 16-QAM MIMO with (8, 2) LDPC code and 64-QAM MIMO with (18, 6) LDPC code have shown that, the proposed JDD scheme's performance is evidently superior over separated detection and decoding (SDD) scheme. More specifically, for the latter case with 12 antennas, JDD shows nearly 10 dB performance improvement than SDD when BER = 10-3. Hardware architecture and complexity analysis are also given in this paper to demonstrate JDD's advantages. Shusen Jing, Junmei Yang, Zhongfeng Wang 0001, Xiaohu You 0001, Chuan Zhang 0001 |
ISCAS | 5 |
| 2017 | Efficient metric sorting schemes for successive cancellation list decoding of polar codesabstractPath metric sorting unit of successive cancellation list (SCL) decoders for polar codes is the main concern in this paper. After reviewing existing sorting units in SCL decoders, we propose 2 new sorting schemes namely quick select (QS) based selection algorithm and simplified bitonic sorter (SBT), which exploit the special data dependency of path metrics in log-likelihood ratio based SCL decoding. Theoretical analysis shows that for the list size of L ≤ 8, QS-based selection algorithm has lower delay than existing schemes. FPGA implementation based on Artix7 Family shows that for the list size of L ≥ 16, SBT has the same delay while the hardware reduction is over 40%. Haochuan Song, Shunqing Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ISCAS | 4 |
| 2017 | Kalman Predictor-Based Proactive Dynamic Thermal Management for 3-D NoC Systems With Noisy Thermal SensorsabstractThermal sensor noise has a great impact on the efficiency and effectiveness of a dynamic thermal management (DTM) strategy. To address the problem of forecasting temperature based on noisy thermal sensors, we first propose a Kalman-based runtime thermal prediction scheme. To obtain accurate temperature predictions, a multivariate linear power model and a physically-based state space thermal model for 3-D network-on-chip are also proposed. Simulation results show that it reduces the standard deviations of the prediction error by 46%–53% compared with the auto-regressive based one under sensor noise with${\sigma =2}$. Conventional reactive DTM techniques suffer from significant performance degradation due to their pessimistic reaction, thus, based on the proposed prediction scheme, we further propose a proactive DTM strategy that primarily consists of a thermal-aware routing algorithm and a proactive throttling scheme: 1) to take into account both thermal and congestion issues, we propose a proactive congestion and thermal aware routing algorithm. Simulation results demonstrate that it can achieve better throughput as well as approach better thermal balance. Specifically, under uniform traffic, the proposed scheme reduces the maximum chip temperature by about 3.9 °C and achieves 78.3% higher throughput compared with the competing thermal optimization approach based on dynamic programming network and 2) when the temperature exceeds the threshold, existing coarse-grained reactive throttling schemes cool down the overheated nodes at the penalty of significant performance loss. In this paper, a proactive quota-based throttling scheme is proposed. Simulation results show that it improves the throughput up to 11.1% compared with the reactive throttling schemes. Li Li 0003, Kun Wang 0005, Chuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Hardware Efficient and Low-Latency CA-SCL Decoder Based on Distributed SortingabstractFor polar codes, cyclic redundancy check (CRC)aided successive cancellation list (CA-SCL) decoder has attracted increasing attention from both academia and industry. In this paper, a hardware efficient and low-latency CA-SCL polar decoder based on distributed sorting is first proposed. For path metric (PM) sorting of each level, a distributed sorting (DS) algorithm is proposed to reduce the comparison complexity from (L2) to (L) (L denotes list size), together with the latency from kL2to kL (k is a coefficient independent of L). Employing folding technique, the N-bit folding polar decoder can be implemented based on the basic √N-bit polar decoder. In addition, pipelining technique is employed to refine the timing issue resulting from folding. The CRC is performed for 2L candidate paths serially to reduce hardware cost. According to demo of (1024, 512) code on Altera Stratix V FPGA, the proposed CA-SCL decoders with L = 2 and adjustable L = 2, 4 consume 9% and 50% board resources, respectively. Decoding latencies (in terms of clock cycles) are 2, 528 and 4, 064, respectively. For L = 2 and 4, we can achieve the frame error rate (FER) of 10-2at the signal noise ratio (SNR) of 2.36 dB and 2.06 dB, respectively. Compared with the floating point results, the performance degradation is negligible. Thus, the proposed design is suitable and adjustable for different real-life scenarios. Xiao Liang 0005, Junmei Yang, Chuan Zhang 0001, Wenqing Song, Xiaohu You 0001 |
GLOBECOM | 3 |
| 2016 | Successive Cancellation Heap Polar DecodingabstractIn this paper, the successive cancellation (SC) heap polar decoding scheme is firstly proposed to reduce the complexity. Unlike SC list decoder which keepsLsame length paths, SC heap decoding stores different length paths in a heap and always decodes the global optimal path in the root. It has been strictly proved that SC heap decoding is superior to SC stack decoding because the time complexity of inserting new paths is only related to the height of the heap. Proposed SC heap decoding is a dynamic decoding scheme with robustness in various scenarios. Numerical results with binary-input additive white Gaussian noise channel (BI-AWGNC) show that SC heap decoding reduces 63.53% decoding complexity compared with SC list decoding on the same performance at SNR of 2.5 dB. A low-complexity hardware architecture for proposed SC heap decoder is also designed. Huayi Zhou 0002, Xiao Liang 0005, Chuan Zhang 0001, Shunqing Zhang, Xiaohu You 0001 |
GLOBECOM | 3 |
| 2016 | Efficient stochastic detector for large-scale MIMOabstractIn this paper, a low-complexity stochastic belief propagation (BP) detector for large-scale MIMO is first proposed. Its efficient hardware architecture, with parallel pipeline, is presented in detail. Thanks to the stochastic approach, all arithmetic operations of the detector are implemented with simple logic structures. Several approaches which can potentially improve the detection performance are exploited. Simulation results have demonstrated that the stochastic BP detector can achieve similar detection performance compared with deterministic one for 32 × 32 MIMO system with 4-quadrature amplitude modulation (4-QAM). With the increase of antenna number, the detection performance improves at the linear expense of complexity and latency. Therefore, the proposed stochastic BP detector is suitable for large-scale MIMO system applications with good balance of detection performance and implementation complexity. Junmei Yang, Chuan Zhang 0001, Shugong Xu, Xiaohu You 0001 |
ICASSP | 2 |
| 2016 | Design space exploration for hardware-efficient stochastic computing: A case study on discrete cosine transformationabstractIn recent years stochastic computing (SC) is re-gaining increasing attention for its unique advantages on low hardware cost and strong error resilience that are the key metrics for nanoscale CMOS era. However, the potential deployment of SC in practical applications is impeded by the long latency of sequential bit-stream and large complexity of pseudo random number generator (PRNG). Aiming to mitigate these challenges, this paper exploits the design space for hardware-efficient stochastic computing with a case study on 4-point discrete cosine transformation (DCT). First, an efficient compensation mechanism is proposed to solve the scaling problem of SC system. Then, two approaches, namely Splitting-Shuffling (SS) and PRNG sharing techniques are proposed to reduce the overall area and processing latency, respectively. Analysis results show that, sustaining the same computing accuracy, the joint use of the proposed approaches leads to 44% reduction in area and 49% reduction on latency than conventional SC design, respectively. Bo Yuan 0001, Chuan Zhang 0001, Zhongfeng Wang 0001 |
ICASSP | 2 |
| 2016 | Efficient architecture for soft-output massive MIMO detection with Gauss-Seidel methodabstractIn massive multiple-input multiple-output (MIMO) uplink, the minimum mean square error (MMSE) algorithm is near-optimal and linear, but still suffers from high-complexity of matrix inversion. Based on Gauss-Seidel (GS) method, an efficient architecture for massive MIMO soft-output detection is proposed in this paper. To further accelerate the convergence rate of the conventional GS method with acceptable overhead complexity, a truncated Neumann series of the first 2 terms, is employed for initialization. The architecture can meet various application requirements by flexibly adjusting the number of iterations. FPGA implementation for a 128 × 8 MIMO demonstrates its advantages in both hardware efficiency and flexibility. Zhizheng Wu 0003, Chuan Zhang 0001, Ye Xue, Shugong Xu, Xiaohu You 0001 |
ISCAS | 2 |
| 2016 | Joint detection and decoding for MIMO systems with polar codesabstractAs well known, the near-optimal K-best detection is popular in multiple-input and multiple-output (MIMO) systems. In this paper, we first propose the joint approaches of K-best detection and polar decoding. For joint detection and decoding (JDD) approach, both hard and soft decisions are considered. The simplified successive cancellation (SSC) decoding is exploited for hard decision, and the successive cancellation list (SCL) decoding is used as soft decision. The system setup for JDD is als o introduced, in which the modulation points across several channels are considered together. Simulation results have demonstrated the performance advantage of the JDD algorithms over the separated ones. For 1/2-rate polar codes, JDD schemes show 50% complexity reduction compared to the separated ones. Furthermore, by employing SSC hard decoding, the JDD algorithm is promising for high-throughput and low-complexity application s. Junmei Yang, Chuan Zhang 0001, Wenqing Song, Shugong Xu, Xiaohu You 0001 |
ISCAS | 2 |
| 2016 | Pipelined belief propagation polar decodersabstractDue to its inherent higher parallelism over successive cancellation (SC) polar decoder, belief propagation (BP) polar decoder becomes more favorable for high throughput applications. However, most existing BP decoders suffer from low utilization. In this paper, a new updating scheme, in which both left-to-right and right-to-left messages are considered identical, is first proposed for memory reduction. By revealing the similarity between BP polar decoder and fast Fourier transform (FFT) processor, both feed-forward and feed-back pipelined BP polar decoders are proposed along with detailed processing schedules. Implementation results have shown that both proposed pipelined BP decoders achieve more than 99.8% arithmetic logical units (ALUTs) reduction and 3.40% registers & block memory reduction, together with more than 7.45% speed-up, compared to the conventional fully parallel one. The proposed design approaches can be generalized with folding technique to achieve the required balance between area and speed flexibly. Junmei Yang, Chuan Zhang 0001, Huayi Zhou 0002, Xiaohu You 0001 |
ISCAS | 2 |
| 2016 | Segmented CRC-Aided SC List Polar DecodingabstractBecause of the existence of channel noise, channel coding serves as an indispensable part of mobile communication system and the essential guarantee for the reliable, accurate, and effective transmission of information. As one of the most competitive channel code candidates for the 5th generation (5G) mobile communication, polar codes are the first codes which can provably achieve the symmetric capacity of binary-input discrete memoryless channels (B-DMCs). In this paper, the segmented CRC- aided successive cancellation list (SCA-SCL) polar decoding scheme is proposed for better tradeoff of performance and complexity. Numerical results on binary-input additive white Gaussian noise channel (BI-AWGNC) have shown that, at SNR of 0.5 dB, this approach successfully provides as high as 41.65% complexity reduction and similar decoding performance compared to state-of-the-art ones. Huayi Zhou 0002, Chuan Zhang 0001, Wenqing Song, Shugong Xu, Xiaohu You 0001 |
VTC Spring | 2 |
| 2015 | Pipelined implementations of polar encoder and feed-back part for SC polar decoderabstractIn this paper, we first reveal the similarity of polar encoder and fast Fourier transform (FFT) processor. Based on this, both feed-forward and feed-back pipelined implementations of polar encoder are proposed. It is pointed out that the feedback part of SC polar decoder is nothing but a simplified version of polar encoder and therefore can be pipelined implemented also. Moreover, a general approach which uniformly constructs most pipelined polar encoders via folding transformation is proposed. Implementation results have shown that both proposed pipelined polar encoder architectures achieve more than 98.3% complexity reduction and more than 9.86% speed-up compared to the conventional implementation. Chuan Zhang 0001, Junmei Yang, Xiaohu You 0001, Shugong Xu |
ISCAS | 1 |
| 2014 | Interleaved successive cancellation polar decodersabstractPolar codes are among the most promising error correction codes due to their ability to achieve the symmetric capacities of the binary-input discrete memoryless channels (B-DMCs). However, how to design successive cancellation (SC) decoders which can maximize the hardware utilization efficiency is still challenging due to the inherent serial nature of SC decoding algorithm. To this end, in this paper, formal design approaches for designing both the time-constrained and resource-constrained interleaved SC decoders are proposed. Compared with the state-of-the-art design, the proposed interleaved decoders can achieve more than 50% reduction in term of area-time product. Chuan Zhang 0001, Keshab K. Parhi |
ISCAS | 1 |
| 2014 | Hardware architecture for list successive cancellation polar decoderabstractThis paper aims at designing an efficient hardware architecture for list successive cancellation (SC) polar decoder. Previous literatures have shown that, compared to conventional SC decoder, list SC decoder has the ability to approach the performance of maximum likelihood (ML) decoder. However, the efficient implementation of list SC decoder has not been proposed yet. To tackle this issue, first we propose a sub-optimal version of list SC decoding. Then different selections of list size L are evaluated. By introducing the pre-computation technique, the hardware architecture for a list SC decoder with L = 2 is proposed. Comparison results have shown that, for a rate-½ (1024, 512) polar code, the proposed decoder can achieve near-optimal decoding performance with less hardware cost and latency than the decoder with conventional design approach. We believe that the design approach presented in this paper will facilitate practical applications of list SC polar decoder. Chuan Zhang 0001, Xiaohu You 0001, Jin Sha 0001 |
ISCAS | 1 |
| 2014 | Efficient column-layered decoders for single block-row quasi-cyclic LDPC codesabstractThe recently proposed single block-row quasi-cyclic low-density parity-check (QC-LDPC) codes are favorable for high-speed applications. However, conventional decoder design methods are not suitable for this kind of codes. To tackle this issue, this paper aims at designing efficient column-layered single block-row QC-LDPC decoder architecture without affecting the decoding performance. Moreover, the simplified version which only requires single minimum value is also proposed for further hardware reduction. Results show that, for the rate-0.9006 (1640, 1477) single block-row QC-LDPC code, the proposed two designs achieves significant advantages in both hardware and latency over their row-layered counterpart. Chuan Zhang 0001, Xiaohu You 0001, Zhongfeng Wang 0001 |
ISCAS | 1 |
| 2014 | Efficient symbol reliability based decoding for QCNB-LDPC codesabstractAs an extension of binary low-density parity-check (LDPC) codes, non-binary LDPC (NB-LDPC) codes show significantly better performance when the code length is moderate or small. Recently, enhanced iterative hard reliability based (EIHRB) decoding algorithm is proposed to reduce the computation complexity. However, the EIHRB algorithm suffers a lot from significant performance degradation when the column weight is small. In this paper, a symbol reliability based (SRB) decoding algorithm, which also performs well when the column weight is low, is proposed for NB-LDPC decoding to improve the decoding performance. With the same maximum iteration number, around 0.38 dB extra coding gain is achieved. Furthermore, the corresponding efficient decoder architecture is proposed. Comparison results have shown that the proposed SRB algorithm can not only achieve good coding gain, but the cost for hardware implementation is reasonable. Leixin Zhou, Jin Sha 0001, Yun Chen 0001, Chuan Zhang 0001, Zhongfeng Wang 0001 |
ISCAS | 4 |
| 2012 | Reduced-latency SC polar decoder architecturesabstractPolar codes have become one of the most favorable capacity achieving error correction codes (ECC) along with their simple encoding method. However, among the very few prior successive cancellation (SC) polar decoder designs, the required long code length makes the decoding latency high. In this paper, conventional decoding algorithm is transformed with look-ahead techniques. This reduces the decoding latency by 50%. With pipelining and parallel processing schemes, a parallel SC polar decoder is proposed. Sub-structure sharing approach is employed to design the merged processing element (PE). Moreover, inspired by the real FFT architecture, this paper presents a novel input generating circuit (ICG) block that can generate additional input signals for merged PEs on-the-fly. Gate-level analysis has demonstrated that the proposed design shows advantages of 50% decoding latency and twice throughput over the conventional one with similar hardware cost. Chuan Zhang 0001, Bo Yuan 0001, Keshab K. Parhi |
ICC | 1 |
| 2012 | Efficient network for non-binary QC-LDPC decoderabstractThis paper presents approaches to develop efficient network for non-binary quasi-cyclic LDPC (QC-LDPC) decoders. By exploiting the intrinsic shifting and symmetry properties of the check matrices, significant reduction of memory size and routing complexity can be achieved. Two different efficient network architectures for Class-I and Class-II non-binary QC-LDPC decoders have been proposed, respectively. Comparison results have shown that for the code of the 64-ary (1260, 630) rate-0.5 Class-I code, the proposed scheme can save more than 70.6% hardware required by shuffle network than the state-of-the-art designs. The proposed decoder example for the 32-ary (992, 496) rate-0.5 Class-II code can achieve a 93.8% shuffle network reduction compared with the conventional ones. Meanwhile, based on the similarity of Class-I and Class-II codes, similar shuffle network is further developed to incorporate both classes of codes at a very low cost. Chuan Zhang 0001, Jin Sha 0001 |
ISCAS | 1 |
| 2009 | High-throughput GCM VLSI Architecture for IEEE 802.1ae ApplicationsabstractThis paper presents a high-throughput GCM VLSI architecture fully compliant to IEEE 802.1ae applications, which can be operated in all modes specified in the standard. Unlike previous works, with the modified parallel GHASH module, the design implements encryption efficiently without knowing the total number of data blocks in advance. Furthermore, a fully subpipelined version of loop-free key expansion architecture is employed to support constant key changes in each clock cycle. An encryptor design example with 2-parallel modified GHASH module is implemented and fabricated in Fujitsu 0.13 mum 1.2 V 1P8M CMOS technology. The ASIC implementation results demonstrate that the maximum operating frequency can reach 764.5 MHz and our design can obtain 97.9 Gb/s throughput with 547 k gates. Chuan Zhang 0001, Li Li 0003, Jun Xu 0013, Zhongfeng Wang 0001 |
ISCAS | 1 |