EDBT 2026 Demo / reviewers in the wild / expert
Wu Guan
dblp:127/4988
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-4321-7288ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temperature-Controlled Quantization-Aware CSI Feedback via a Multi-Value Transformer
Bolin Yan, Wu Guan, Liping Liang 0001 |
ICC | 3 |
| 2026 | HMNet: A Hybrid Multi-Resolution Network with Local Uniform Quantization for Low-Complexity CSI Feedback
Bolin Yan, Mingwei Yan, Wu Guan, Liping Liang 0001 |
WCNC | 4 |
| 2026 | Mutual Constrained Min-Sum Check-Belief Propagation for 5G LDPC Codes: Algorithm and ImplementationabstractState-of-the-art (SOA) designs for 5G New Radio (NR) low-density parity-check (LDPC) face a dilemma: enduring the high complexity of cumulative calculations or accepting performance degradation from unreliable messages for degree-1 variable nodes (VNs). This paper introduces the mutual constrained min-sum check-belief propagation (MCMS-CBP) decoding algorithm to resolve this dilemma. By establishing the mutual constrained cycle in well-designed message updating, while incorporating the constraint compensation strategy, the non-cumulative MCMS-CBP substantially improves the reliability of check-belief to variable-node messages for degree-1 VNs. Numerical results show that fixed-point MCMS-CBP maintains a gap of less than 0.1 dB compared to floating-point flooding belief propagation decoding algorithm across various 5G NR LDPC codes and offers inherently low complexity for hardware implementation. In addition, this paper presents an area-throughput efficient 5G NR LDPC decoder which is based on fixed-point MCMS-CBP and supports all 5G NR LDPC codes. Key techniques for optimizing the area and latency of decoder include: the sign-bit storage, reducing memory overhead by 87.5%; the intra-layer first block schedule, eliminating pipeline stalls;the partially aligned element arrangement, halving the cyclic shifter cost while eliminating memory pre-writeback rotation. Implementation result on SMIC 28nm technology demonstrates that the proposed decoder has a cell area of 1.610 mm$\boldsymbol {^{2}}$and a peak throughput of 58.63 Gbps, yielding a 36.42 Gbps/mm$\boldsymbol {^{2}}$maximum area efficiency that is$\boldsymbol {2.72\times }$higher than SOA LDPC decoders. Ziqin Yan, Peihao Fan, Wu Guan, Liping Liang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | A High-Performance ML-KEM Architecture With Task-Level ParallelismabstractAs the standardization of postquantum cryptography progresses, module-lattice-based key-encapsulation mechanism (ML-KEM), the standardized successor to CRYSTALS-Kyber, has become a primary target for hardware deployment. Owing to the tight interactions among ML-KEM compute tasks, achieving end-to-end high performance remains challenging. To address this problem, the proposed task-level parallelism (TLP)-centric system scheduler organizes a highly parallel task flow for ML-KEM, enabling conflict-free overlap among major tasks and substantially reducing the overall cycle count. Furthermore, by codesigning a MUX-steered dual-RAM 2-in/2-out access mechanism and a parallel polynomial compute unit, the memory system and compute fabric jointly sustain this high-parallelism task flow, resulting in a high-performance ML-KEM hardware architecture. Implemented on a Xilinx Artix-7 FPGA at 210 MHz, the proposed architecture supports all three ML-KEM security levels, uses 13.8 K LUTs, 9.6 K FFs, and 3.66 K slices, and completes in 18.9/28.7/$41.8~\mu $s for ML-KEM-512/768/1024, respectively. Compared with the latest designs, the proposed architecture achieves46%–47%shorter runtime,while simultaneously maintaining balanced resource usage and improving area-time product by 21%–23%. Peihao Fan, Ziqin Yan, Wu Guan, Liping Liang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | SpeedLLM: An FPGA Co-design of Large Language Model Inference AcceleratorabstractThis paper introduces SpeedLLM, a neural network accelerator designed on the Xilinx Alevo U280 platform and optimized for the Tinyllama framework to enhance edge computing performance. Key innovations include data stream parallelism, a memory reuse strategy, and Llama2 operator fusion, which collectively reduce latency and energy consumption. SpeedLLM's data pipeline architecture optimizes the read-compute-write cycle, while the memory strategy minimizes FPGA resource demands. The operator fusion boosts computational density and throughput. Results show SpeedLLM outperforms traditional Tinyllama implementations, achieving up to 4.8× faster performance and 1.18× lower energy consumption, offering improvements in edge devices. Wu Guan, Liping Liang 0001, Hanqing Luo |
HPDC | 2 |
| 2025 | Lightweight Spatial-Temporal Resolution Network for Massive MIMO CSI FeedbackabstractThe growing complexity of AI-based channel state information (CSI) feedback within Massive Multi-Input Multi-Output (MIMO) systems poses a substantial challenge, particularly for low-capability devices such as small cells and edge computing systems. To address this challenge, this work proposes a novel lightweight network named LSTCNet. LSTCNet employs a combined parallel-recurrent mechanism to extract spatial and temporal features from CSI. By substituting fundamental matrix computations for complex linear operations, LSTCNet is specifically tailored to meet the rigorous requirements of CSI feedback tasks. Additionally, the incorporation of a configurable quantizer ensures alignment with prevailing industry standards. The experimental results demonstrate that LSTCNet achieves high accuracy and exhibits strong generalization capabilities, all while maintaining a low computational complexity. The normalized mean square error (NMSE) of LSTCNet is comparable to that of Transformer-based networks, while its complexity is less than 0.56 times that of CsiNet; however, the complexity of TransNet exceeds 6.34 times that of CsiNet. Furthermore, the squared generalized cosine similarity (SGCS) of LSTCNet is comparable to that of EVCsiNet, while its complexity is only 0.05 times that of EVCsiNet. Bolin Yan, Wu Guan, Liping Liang 0001 |
IJCNN | 3 |
| 2024 | Adaptive Granularity Progressive LDPC Decoding for NAND Flash MemoryabstractProgressive low-density parity check (LDPC) code decoding has been widely used to correct increasing raw bit errors in NAND Flash memory. Once the decoding of a single logical page fails, the read-retry operation will reprocess at an increased read level with more accurate initial log-likelihood ratio (LLR) messages. However, the traditional progressive LDPC decodings with inappropriate read-level-increase granularities of read-retry operations introduce unnecessary flash read latency. By taking advantage of globally coupled LDPC (GC-LDPC) codes, an improved adaptive granularity progressive LDPC decoding (IAGPD) is proposed. This method can estimate the number of uncorrectable bit errors before each read-retry operation by detecting the unsatisfied local parity checks and general syndrome in the decoding failure. Then, it adaptively selects the optimal read-level-increase granularities for read-retry operations in the progressive LDPC decoding. Compared with the existing decoding methods, only by an extra 0.098% of the decoder area and two clock cycles, our method can reduce the flash read latency by up to 43%. And the solid-state drive (SSD) read response time on MQsim can be reduced by up to 32%. Binhao Bao, Qianhui Li, Wu Guan, Liping Liang 0001, Xin Qiu 0008 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Check-Belief Propagation Decoding of LDPC CodesabstractVariant belief propagation (BP) algorithms are applied to low-density parity-check (LDPC) codes. However, conventional decoders suffer from a large resource consumption due to gathering messages from all the neighbour variable-nodes and/or check-nodes through cumulative calculations. In this paper, a check-belief propagation (CBP) decoding algorithm is proposed. Check-belief is used as the probability that the corresponding parity-check is satisfied. All check-beliefs are iteratively enlarged in a sequential recursive order, and successful decoding will be achieved after the check-beliefs are all big enough. Compared to previous algorithms employing a large number of cumulative calculations to gather all the neighbour messages, CBP decoding can renew each check-belief by propagating it from one check-node to another through only one variable-node, resulting in a low complexity decoding with no cumulative calculations. The simulation results and analyses show that the CBP algorithm provides little error-rate performance loss in contrast with the previous BP algorithms, but consumes much fewer calculations and memories than them. It earns a big benefit in terms of complexity. Wu Guan, Liping Liang 0001 |
IEEE Trans. Commun. | 1 |