VLDB 2026 Research / reviewers in the wild / expert
Huayi Zhou 0002
dblp:182/4224-2
· DBLP profile ↗
12ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-1745-4135ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Area-Efficient Routing Solution for Automorphism Ensemble Decoding of Polar Codes
Jiajie Li 0001, Huayi Zhou 0002, Ryan Seah, Marwan Jalaleddine, Warren J. Gross |
IEEE Trans. Commun. | 2 |
| 2026 | MrcPunc: Modulation and Rate Compatible Puncturing for Few-Iteration 5G LDPC DecodingabstractTo meet the stringent high-throughput demands of enhanced mobile broadband (eMBB), low-density parity-check (LDPC) decoding is limited to few iterations. This constraint usually degrades error-correcting performance, rousing a critical need for puncturing towards algorithmic enhancements. However, existing puncturing methods lack modulation and rate compatibility, converging to local optima due to greedy search strategies. To this end, this paper proposes a modulation and rate compatible puncturing (MrcPunc) for 5G LDPC codes. It incorporates three key techniques: 1)a revised protograph-based extrinsic information transfer (PEXIT) analysisfor accurate performance prediction, 2)a progressive multi-path search (PMPS) algorithmto avoid local optima, and 3)a compatibility optimizationsupporting various modulations and code rates. The MrcPunc yields up to 0.6 dB gain over the 5G standard puncturing under few-iteration condition without extra decoding complexity. It is noted that MrcPunc guarantees robustness against both larger iteration counts and mismatched modulations, alongside reducing the average number of decoding iterations. Qiushi Xu, Huayi Zhou 0002, Mingyang Zhu, Ming Jiang 0012, Chuan Zhang 0001 |
IEEE Trans. Commun. | 2 |
| 2025 | Reduced-Complexity Projection-Aggregation List Decoder for Reed-Muller CodesabstractProjection-aggregation decoders have been used in conjunction with a list structure to achieve near maximum-likelihood decoding for short-length and low-rate Reed-Muller (RM) codes but suffer from high computational complexity. We reduce the worst-case computational complexity of projection-aggregation (PA) decoders by more than 50% using a scheduling scheme compared to PA decoders without the scheduling scheme, and propose a redesigned syndrome check pattern to avoid repeated syndrome computations in the decoder. A latency model based on the existing hardware architecture is proposed. Input distribution aware (IDA) decoding is adopted as a pre-possessing tool, and the average list size when using IDA decoding is analytically derived under additive white Gaussian noise and uncorrelated normalized Rayleigh fading channels. Using IDA, the average list size is reduced by 30% with less than 0.1 dB loss. The proposed list decoders require a smaller computational complexity than the state-of-the-art iterative decoder, automorphism ensemble decoding with the belief propagation constituent decoder (AED-BP) for decoding RM(7, 3) and RM(8, 3) codes. Based on the developed latency models, the PA list decoder has a smaller latency than the AED-BP and the successive cancellation list decoder to reach near maximum-likelihood decoding performance. Jiajie Li 0001, Huayi Zhou 0002, Marwan Jalaleddine, Warren J. Gross |
IEEE Trans. Commun. | 2 |
| 2025 | Low-Complexity Breadth-First Search Detection for Large-Scale MIMO SystemsabstractThanks to its near-optimal performance, breadth-first search detection (BFSD) finds widespread application in small-scale MIMO systems. However, existing BFSD methods struggle to effectively configure the width (number of candidate nodes) for each layer, resulting in prohibitive complexity in large-scale MIMO systems. To address this, we propose two width optimization schemes for BFSD. We introduce a layer-by-layer optimization framework to reduce the design space of width configurations, and a Monte Carlo-assisted method to link width configurations to detection performance. Using this linking scheme in the reduced design space, we formulate the first width optimization scheme given specific performance constraints. Then, we present another scheme that employs a theoretical linking method as an alternative to the Monte Carlo approach. Although slightly less effective, the second scheme has negligible complexity for width optimization, making it well-suited for communication scenarios with time-varying characteristics. In 128×128 MIMO systems, numerical results demonstrate that the optimized BFSD using our first and second schemes can reduce complexity by up to 82% and 65%, respectively, while achieving superior detection performance compared to state-of-the-art BFSD. Jian Zheng 0003, Yutai Sun, Huayi Zhou 0002, Wenyue Zhou, Yongming Huang 0001, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Commun. | 3 |
| 2024 | Decoding of Polar Codes Using Quadratic Unconstrained Binary OptimizationabstractPolar codes encounter challenges in decoder complexity while preserving good error-correction properties. Instead of conventional decoders, a quantum annealer (QA) decoder has been proposed to explore untapped possibilities. For future QA applications, a crucial prerequisite is transforming the optimization problem into quadratic unconstrained binary optimization (QUBO) form. However, existing QUBO forms for polar decoding result in suboptimal frame error rate (FER) performance for codes exceeding 8 bits. This paper redesigns the QUBO form for polar decoding. We first introduce a novel receiver constraint modeled by the binary cross-entropy (BCE) function. Utilizing a simulated annealing (SA) solver with the proposed QUBO form with BCE (QUBO-BCE) achieves maximum-likelihood (ML) performance for a code length of 32 bits. Next, to reduce the number of variables, we remove the frozen variables and introduce a simplified QUBO-BCE form (SQUBO-BCE). Additionally, CRC polynomials are modelled into constraints in QUBO form, resulting in a CRC-aided SQUBO-BCE (CA-SQUBO-BCE) form for polar decoding to further enhance the FER. Numerical results demonstrate that SQUBO-BCE achieves ML performance and reduces up to 61.5% of variables compared to QUBO-BCE. Furthermore, the proposed CA-SQUBO-BCE achieves near CRC-aided ML performance. The proposed SQUBO-BCE requires the lowest number of SA processes to reach a specific FER. Huayi Zhou 0002, Ryan Seah, Marwan Jalaleddine, Warren J. Gross |
IEEE J. Sel. Areas Commun. | 1 |
| 2023 | Hybrid GRAND Sphere Decoding: Accelerated GRAND for Low-Rate CodesabstractGuessing random additive noise decoding (GRAND) and sphere decoding (SD) are two algorithms that can achieve maximum likelihood decoding. In this paper, a hybrid GRAND-SD (HGRAND) scheme is proposed to extend GRAND to low-rate codes. An accelerated GRAND decoder, assisted by a sphere decoder running in parallel and giving hints to it to allow skipping of certain candidates allows HGRAND to achieve a latency below the minimum latency of the individual component decoders while guaranteeing error-correction performance. Huayi Zhou 0002, Warren J. Gross |
ISCAS | 1 |
| 2023 | Partial Ordered Statistics Decoding with Enhanced Error PatternsabstractGuessing Random Additive Noise Decoding (GRAND) excels at decoding high-rate codes but struggles to decode low-rate codes with reasonable complexity. Ordered Statistics Decoding (OSD) specifically excels in decoding short codes irrespective of rates; however, OSD necessitates the use of Gaussian elimination which introduces additional time, space and computational complexity. Partial Ordered Statistics Decoding (POSD) was proposed to reduce the time, space, and computational complexity of OSD; however, the current partition-based POSD has poor decoding performance since it does not generate test error patterns across partitions. In this paper, we propose to improve the decoding performance of POSD by incorporating test error patterns inspired by GRAND methods. This work offers a trade-off between performance and complexity compared to existing decoders such as GRAND and OSD. We enhance POSD by optimizing the scheduling of Test Error Patterns (TEPs) and show that our technique can be applied to any code in a standard form. At a target BER 10−4with eBCH (128,64) the enhanced error patterns achieve more than 0.6 dB gain in performance compared to the POSD with partition-based error patterns. Moreover, at a target frame error rate of 10−5, POSD uses 10× less binary operations compared to GRAND when decoding eBCH (128,64) and RLC(128,64) codes. With BCH (127,29) and RLC(128,32), at a target frame error rate of 10−2, POSD with enhanced error patterns with a maximum number of queries (MQ) of 104achieves up to a 2 dB gain to its GRAND equivalent which is using 107maximum number of queries. Marwan Jalaleddine, Huayi Zhou 0002, Jiajie Li 0001, Warren J. Gross |
ISIT | 2 |
| 2019 | Efficient Successive Cancellation Stack Decoder for Polar CodesabstractAs an improved version of successive cancellation (SC) polar decoder, an SC stack (SCS) decoder has been proposed for performance improvement. However, the existing SCS polar decoder suffers a lot from high time complexity at low signal-to-noise ratio (SNR) region and space complexity compared with the SC decoder. To this end, two improved decoders are proposed to reduce time and space complexity in both low and high SNR regions. The first one is the segmented cyclic redundancy check (CRC)-aided SCS (SCA-SCS) decoder, which is based on segmented parity checkers. The second one is the adaptive SCS (ASCS) decoder, which has the flexibility of stack depth and searching width. Furthermore, a channel condition estimator is proposed to select appropriate decision criteria for different SNR scenarios. Results have shown that for the polar code of length 1024 and rate 1/2, two improved SCS decoders can perform better than the traditional SCS decoder. The proposed SCA-SCS decoder and the ASCS decoder can achieve 10.8% and 11.42% time complexity reduction and 31.68% and 60.85% space complexity reduction on average over binary-input additive white Gaussian noise channels (BI-AWGNCs), respectively. Efficient parallel hardware architecture of the SCS polar decoder is first proposed and implemented with 90- and 65-nm technologies. Results have verified its advantages over the state of the art (SOA). Wenqing Song, Huayi Zhou 0002, Kai Niu 0001, Zaichen Zhang, Li Li 0003, Xiaohu You 0001, Chuan Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Joint List Polar Decoder with Successive Cancellation and Sphere DecodingabstractFor polar codes, both successive cancellation list (SCL) decoding and list sphere decoding (LSD) aim to balance performance and complexity. The same list structure but different decoding schedules of SCL and LSD can lead to a combination of both schemes. In this paper, an efficient joint list decoder with SCL and LSD (JLSCD) is proposed to reduce time complexity. We apply SCL and LSD schemes simultaneously but independently, then merge them at the middle point of the decoding. Numerical results have demonstrated JLSCD scheme's advantage in complexity. FPGA implementation of JLSCD decoder is also given in this paper. Xiao Liang 0005, Huayi Zhou 0002, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
ICASSP | 2 |
| 2016 | Successive Cancellation Heap Polar DecodingabstractIn this paper, the successive cancellation (SC) heap polar decoding scheme is firstly proposed to reduce the complexity. Unlike SC list decoder which keepsLsame length paths, SC heap decoding stores different length paths in a heap and always decodes the global optimal path in the root. It has been strictly proved that SC heap decoding is superior to SC stack decoding because the time complexity of inserting new paths is only related to the height of the heap. Proposed SC heap decoding is a dynamic decoding scheme with robustness in various scenarios. Numerical results with binary-input additive white Gaussian noise channel (BI-AWGNC) show that SC heap decoding reduces 63.53% decoding complexity compared with SC list decoding on the same performance at SNR of 2.5 dB. A low-complexity hardware architecture for proposed SC heap decoder is also designed. Huayi Zhou 0002, Xiao Liang 0005, Chuan Zhang 0001, Shunqing Zhang, Xiaohu You 0001 |
GLOBECOM | 1 |
| 2016 | Pipelined belief propagation polar decodersabstractDue to its inherent higher parallelism over successive cancellation (SC) polar decoder, belief propagation (BP) polar decoder becomes more favorable for high throughput applications. However, most existing BP decoders suffer from low utilization. In this paper, a new updating scheme, in which both left-to-right and right-to-left messages are considered identical, is first proposed for memory reduction. By revealing the similarity between BP polar decoder and fast Fourier transform (FFT) processor, both feed-forward and feed-back pipelined BP polar decoders are proposed along with detailed processing schedules. Implementation results have shown that both proposed pipelined BP decoders achieve more than 99.8% arithmetic logical units (ALUTs) reduction and 3.40% registers & block memory reduction, together with more than 7.45% speed-up, compared to the conventional fully parallel one. The proposed design approaches can be generalized with folding technique to achieve the required balance between area and speed flexibly. Junmei Yang, Chuan Zhang 0001, Huayi Zhou 0002, Xiaohu You 0001 |
ISCAS | 3 |
| 2016 | Segmented CRC-Aided SC List Polar DecodingabstractBecause of the existence of channel noise, channel coding serves as an indispensable part of mobile communication system and the essential guarantee for the reliable, accurate, and effective transmission of information. As one of the most competitive channel code candidates for the 5th generation (5G) mobile communication, polar codes are the first codes which can provably achieve the symmetric capacity of binary-input discrete memoryless channels (B-DMCs). In this paper, the segmented CRC- aided successive cancellation list (SCA-SCL) polar decoding scheme is proposed for better tradeoff of performance and complexity. Numerical results on binary-input additive white Gaussian noise channel (BI-AWGNC) have shown that, at SNR of 0.5 dB, this approach successfully provides as high as 41.65% complexity reduction and similar decoding performance compared to state-of-the-art ones. Huayi Zhou 0002, Chuan Zhang 0001, Wenqing Song, Shugong Xu, Xiaohu You 0001 |
VTC Spring | 1 |