EDBT 2026 Demo / reviewers in the wild / expert
Liping Liang 0001
dblp:47/5734-1
· DBLP profile ↗
14ranked-venue papers
0as first author
14since 2021 · last 2026
0009-0007-6809-2537ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 10 since 2021Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temperature-Controlled Quantization-Aware CSI Feedback via a Multi-Value Transformer
Bolin Yan, Wu Guan, Liping Liang 0001 |
ICC | 4 |
| 2026 | HMNet: A Hybrid Multi-Resolution Network with Local Uniform Quantization for Low-Complexity CSI Feedback
Bolin Yan, Mingwei Yan, Wu Guan, Liping Liang 0001 |
WCNC | 5 |
| 2026 | A Markov-Chain-Based PUF Using Chain-Block-Obfuscation Mechanism Resisting Machine Learning Attacks With High Uniformity RobustnessabstractPhysical Unclonable Functions (PUFs) are lightweight hardware security primitives suitable for resource-constrained Internet of Things(IoT) devices. The Arbiter PUF (APUF), as a classic strong PUF structure, is well-suited for lightweight device authentication. Unfortunately, due to certain structural characteristics, the classic APUF can be successfully modeled by various machine learning (ML) models with only a small number of CRP (Challenge-Response Pair) samples. To address this security issue, In this paper, we build upon the classic APUF structure by introducing a Chain Block Obfuscation(CBO) mechanism and a Markov obfuscation mechanism at the input and output stages, respectively. This approach demonstrates strong resistance against four machine learning models—logistic regression (LR), support vector machine (SVM), covariance matrix adaptation evolutionary strategies (CMA-ES), and artificial neural networks (ANN)—while optimizing hardware resource usage by at least 53.4% compared to previous attack-resistant structures. Even with up to 2M CRPs for training, the prediction accuracy remains around 50%. Additionally, XOR, a widely adopted output obfuscation mechanism by many researchers, has limitations in output uniformity when the number of intra-chip parallel APUFs is small. Studies have shown that the uniqueness between intra-chip APUFs may not reach the ideal value of 50%. In such cases, especially when there are only two parallel APUFs in the circuit, the uniformity of the XOR output is affected by insufficient uniqueness. To address this issue, this paper proposes a MUX-based Markov output obfuscation mechanism. Experimental data on the Xilinx Artix-7 field-programmable gate array (FPGA) platform demonstrate that this mechanism exhibits strong robustness in uniformity compared to the traditional XOR output obfuscation, particularly when the number of intra-chip parallel APUFs is small. Junhong Gan, Hanqing Luo, Yuanfeng Xie, Liping Liang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | Mutual Constrained Min-Sum Check-Belief Propagation for 5G LDPC Codes: Algorithm and ImplementationabstractState-of-the-art (SOA) designs for 5G New Radio (NR) low-density parity-check (LDPC) face a dilemma: enduring the high complexity of cumulative calculations or accepting performance degradation from unreliable messages for degree-1 variable nodes (VNs). This paper introduces the mutual constrained min-sum check-belief propagation (MCMS-CBP) decoding algorithm to resolve this dilemma. By establishing the mutual constrained cycle in well-designed message updating, while incorporating the constraint compensation strategy, the non-cumulative MCMS-CBP substantially improves the reliability of check-belief to variable-node messages for degree-1 VNs. Numerical results show that fixed-point MCMS-CBP maintains a gap of less than 0.1 dB compared to floating-point flooding belief propagation decoding algorithm across various 5G NR LDPC codes and offers inherently low complexity for hardware implementation. In addition, this paper presents an area-throughput efficient 5G NR LDPC decoder which is based on fixed-point MCMS-CBP and supports all 5G NR LDPC codes. Key techniques for optimizing the area and latency of decoder include: the sign-bit storage, reducing memory overhead by 87.5%; the intra-layer first block schedule, eliminating pipeline stalls;the partially aligned element arrangement, halving the cyclic shifter cost while eliminating memory pre-writeback rotation. Implementation result on SMIC 28nm technology demonstrates that the proposed decoder has a cell area of 1.610 mm$\boldsymbol {^{2}}$and a peak throughput of 58.63 Gbps, yielding a 36.42 Gbps/mm$\boldsymbol {^{2}}$maximum area efficiency that is$\boldsymbol {2.72\times }$higher than SOA LDPC decoders. Ziqin Yan, Peihao Fan, Wu Guan, Liping Liang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | A High-Performance ML-KEM Architecture With Task-Level ParallelismabstractAs the standardization of postquantum cryptography progresses, module-lattice-based key-encapsulation mechanism (ML-KEM), the standardized successor to CRYSTALS-Kyber, has become a primary target for hardware deployment. Owing to the tight interactions among ML-KEM compute tasks, achieving end-to-end high performance remains challenging. To address this problem, the proposed task-level parallelism (TLP)-centric system scheduler organizes a highly parallel task flow for ML-KEM, enabling conflict-free overlap among major tasks and substantially reducing the overall cycle count. Furthermore, by codesigning a MUX-steered dual-RAM 2-in/2-out access mechanism and a parallel polynomial compute unit, the memory system and compute fabric jointly sustain this high-parallelism task flow, resulting in a high-performance ML-KEM hardware architecture. Implemented on a Xilinx Artix-7 FPGA at 210 MHz, the proposed architecture supports all three ML-KEM security levels, uses 13.8 K LUTs, 9.6 K FFs, and 3.66 K slices, and completes in 18.9/28.7/$41.8~\mu $s for ML-KEM-512/768/1024, respectively. Compared with the latest designs, the proposed architecture achieves46%–47%shorter runtime,while simultaneously maintaining balanced resource usage and improving area-time product by 21%–23%. Peihao Fan, Ziqin Yan, Wu Guan, Liping Liang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | SpeedLLM: An FPGA Co-design of Large Language Model Inference AcceleratorabstractThis paper introduces SpeedLLM, a neural network accelerator designed on the Xilinx Alevo U280 platform and optimized for the Tinyllama framework to enhance edge computing performance. Key innovations include data stream parallelism, a memory reuse strategy, and Llama2 operator fusion, which collectively reduce latency and energy consumption. SpeedLLM's data pipeline architecture optimizes the read-compute-write cycle, while the memory strategy minimizes FPGA resource demands. The operator fusion boosts computational density and throughput. Results show SpeedLLM outperforms traditional Tinyllama implementations, achieving up to 4.8× faster performance and 1.18× lower energy consumption, offering improvements in edge devices. Wu Guan, Liping Liang 0001, Hanqing Luo |
HPDC | 3 |
| 2025 | ISO 26262-Aligned Functional Safety Verification Framework with Explainable Graph Neural NetworkabstractThe growing complexity and integration of automotive electronic systems, driven by advancements in intelligent vehicles and autonomous driving, make functional safety (FuSa) verification critical for ensuring system reliability. Traditional fault injection (FI) methods face inefficiencies, scalability limitations, and interpretability gaps, hard to meet the stringent requirements of ISO 26262 standards for safety-critical automotive systems. This paper proposes an explainable Graph Neural Network(GNN)-based framework for FuSa verification in automotive electronics through three core contributions: graph neural networks for modeling circuit structures to identify critical fault nodes, gradient-driven feature importance analysis to optimize selective hardening with minimal resource overhead and GNNExplainer to visualize critical nodes and connections driving fault-criticality predictions through subgraph analysis. Validated across diverse circuits the framework achieves up to 99.6% precision and 99.8% F1-score in fault detection while significantly reducing the simulation time. Notably, the framework improves accuracy by approximately 5% compared to state-of-the-art (SOTA) methods while requiring only half the fault injection data. Integrated explainable artificial intelligence techniques provide transparent decision traces, ensuring compliance with ISO 26262 traceability requirements. Through feature selection, the framework achieves comparable accuracy with a minimal feature set, significantly reducing computational overhead while maintaining performance. By bridging AI-driven automation with rigorous safety certification, this work establishes a scalable, efficient, and interpretable solution for FuSa verification in automotive SoCs. Yutao Sun, Jiehua Huang, Xiangping Liao, Liping Liang 0001 |
ICCAD | 5 |
| 2025 | Lightweight Spatial-Temporal Resolution Network for Massive MIMO CSI FeedbackabstractThe growing complexity of AI-based channel state information (CSI) feedback within Massive Multi-Input Multi-Output (MIMO) systems poses a substantial challenge, particularly for low-capability devices such as small cells and edge computing systems. To address this challenge, this work proposes a novel lightweight network named LSTCNet. LSTCNet employs a combined parallel-recurrent mechanism to extract spatial and temporal features from CSI. By substituting fundamental matrix computations for complex linear operations, LSTCNet is specifically tailored to meet the rigorous requirements of CSI feedback tasks. Additionally, the incorporation of a configurable quantizer ensures alignment with prevailing industry standards. The experimental results demonstrate that LSTCNet achieves high accuracy and exhibits strong generalization capabilities, all while maintaining a low computational complexity. The normalized mean square error (NMSE) of LSTCNet is comparable to that of Transformer-based networks, while its complexity is less than 0.56 times that of CsiNet; however, the complexity of TransNet exceeds 6.34 times that of CsiNet. Furthermore, the squared generalized cosine similarity (SGCS) of LSTCNet is comparable to that of EVCsiNet, while its complexity is only 0.05 times that of EVCsiNet. Bolin Yan, Wu Guan, Liping Liang 0001 |
IJCNN | 5 |
| 2025 | Graph Attention Networks Based Fault Prediction Framework for Functional Safety VerificationabstractAs automotive chips grow in complexity, the cost of functional safety(FuSa) verification rises sharply. This paper proposes a Graph Attention Network (GAT) based framework that extracts fault propagation features and iteratively identifies critical nodes prone to single-point failures under ISO 26262 guidance. Tested on five open-source circuits, the framework achieves 97.89% accuracy and 98.42% F1score. Leave-one-out cross-validation yields 97.12% accuracy and 96.86% F1-score, demonstrating strong generalization from small to large-scale circuits. Compared to traditional RTL fault injection and neural network methods, it reduces fault simulation time and training data usage by about half. The framework outperforms existing techniques in both accuracy and efficiency, offering a practical solution for scalable automotive circuit safety verification. Yutao Sun, Jiehua Huang, Xiangping Liao, Liping Liang 0001 |
ITC | 5 |
| 2025 | A Novel High-Throughput FFT Processor With a Block-Level Pipeline for 5G MIMO OFDM SystemsabstractIn fifth-generation (5G) communication systems, multiple input multiple output (MIMO) and orthogonal frequency-division multiplexing (OFDM) are two critical technologies. Fast Fourier transform (FFT), as the core processing steps of OFDM, directly affects the overall system performance. In this brief, we proposed a novel block-level pipelined architecture, which divides the FFT processor into three pipeline blocks: input, radix, and output. Each pipeline block can run in a different FFT simultaneously to achieve higher throughput. Specifically, to reduce the OFDM system-level latency of 5G applications, the FFT processor supports weighted overlap and add (WOLA) on the cyclic prefix and suffix of OFDM symbols. This architecture is implemented using TSMC 12-nm technology, with a processor die area of 0.89 mm2and a power consumption of 568 mW at 1 GHz. The FFT processor can achieve a system-level throughput up to 2.66 GS/s. Hanqing Luo, Shengnan Lin, Liping Liang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | An FPGA-Based Emulation Platform for Functional Safety Verification in Automotive SoC SystemsabstractThe increasing complexity of automotive SoC systems, particularly in the context of autonomous vehicles, demands rigorous functional safety verification methods. This paper proposes a novel Functional Safety FPGA Fault Injection Tool (FSF-FIT) which builds an FPGA-based emulation platform by using FPGA fault injection techniques. The core innovation of FSFFIT lies in its comprehensive FPGA functional safety evaluation process, targeting specific components within the Design Under Test (DUT), down to individual LUTs or registers. This detailed approach allows for accurate reliability level (RL) and diagnostic coverage (DC) assessments of safety mechanisms. The speed and accuracy of FSFFIT were validated through experiments conducted on the XuanTie C906 RISC-V processor as well as on a multiplier with different security mechanisms. The results demonstrated that FSFFIT’s performance is consistent with that of SSIM, a certified ISO26262-compliant functional safety tool, while also achieving faster execution times. Additionally, by comparing the results with some state-of-the-art (SOTA) FPGA fault injection tools, our method is superior in terms of injection speed. Yutao Sun, Zean Huang, Liping Liang 0001 |
ATS | 4 |
| 2024 | Adaptive Granularity Progressive LDPC Decoding for NAND Flash MemoryabstractProgressive low-density parity check (LDPC) code decoding has been widely used to correct increasing raw bit errors in NAND Flash memory. Once the decoding of a single logical page fails, the read-retry operation will reprocess at an increased read level with more accurate initial log-likelihood ratio (LLR) messages. However, the traditional progressive LDPC decodings with inappropriate read-level-increase granularities of read-retry operations introduce unnecessary flash read latency. By taking advantage of globally coupled LDPC (GC-LDPC) codes, an improved adaptive granularity progressive LDPC decoding (IAGPD) is proposed. This method can estimate the number of uncorrectable bit errors before each read-retry operation by detecting the unsatisfied local parity checks and general syndrome in the decoding failure. Then, it adaptively selects the optimal read-level-increase granularities for read-retry operations in the progressive LDPC decoding. Compared with the existing decoding methods, only by an extra 0.098% of the decoder area and two clock cycles, our method can reduce the flash read latency by up to 43%. And the solid-state drive (SSD) read response time on MQsim can be reduced by up to 32%. Binhao Bao, Qianhui Li, Wu Guan, Liping Liang 0001, Xin Qiu 0008 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Check-Belief Propagation Decoding of LDPC CodesabstractVariant belief propagation (BP) algorithms are applied to low-density parity-check (LDPC) codes. However, conventional decoders suffer from a large resource consumption due to gathering messages from all the neighbour variable-nodes and/or check-nodes through cumulative calculations. In this paper, a check-belief propagation (CBP) decoding algorithm is proposed. Check-belief is used as the probability that the corresponding parity-check is satisfied. All check-beliefs are iteratively enlarged in a sequential recursive order, and successful decoding will be achieved after the check-beliefs are all big enough. Compared to previous algorithms employing a large number of cumulative calculations to gather all the neighbour messages, CBP decoding can renew each check-belief by propagating it from one check-node to another through only one variable-node, resulting in a low complexity decoding with no cumulative calculations. The simulation results and analyses show that the CBP algorithm provides little error-rate performance loss in contrast with the previous BP algorithms, but consumes much fewer calculations and memories than them. It earns a big benefit in terms of complexity. Wu Guan, Liping Liang 0001 |
IEEE Trans. Commun. | 2 |
| 2023 | A 60-Mode High-Throughput Parallel-Processing FFT Processor for 5G/4G ApplicationsabstractThis article presents a 60-mode high-throughput parallel-processing memory-based fast Fourier transform (FFT) processor for fifth-generation (5G)/4G applications. The proposed architecture adopts a multi-FFT parallel-processing scheme to significantly reduce the idle computation cycles of small point FFT due to the deep pipeline. The proposed scheme can narrow the throughput gap between different FFT sizes. In conjunction with the parallel processing scheme, a configurable 16-parallelism butterfly unit is proposed to support a maximum of five radix-3, four radix-4, three radix-5, two radix-8, or one radix-9 operation in one cycle. Furthermore, this article demonstrates conflict-free memory access with a parallel method based on fusion shift combined with a hopping-based method to simplify the data arrangement at different radix stages. According to the orthogonal frequency division multiplexing (OFDM) application characteristics, our processor is extended to support many additional commonly used functions, such as zero padding, cyclic prefix insertion, and 7.5-kHz frequency shift, which can simplify the system call in 5G/4G applications. The FFT processor has been successfully integrated into a small-cell base station baseband system-on-chip (SoC) and taped out at the TSMC 12-nm technology. The implementation result reveals that the die area of the proposed processor is 0.374 mm2 with a power consumption of 235.5 mW at 1 GHz, and the processor supports a balanced throughput up to 3.92 GS/s at all 60-mode, which is better than that of the state-of-the-art (SOTA) designs. Qinzhi Hong, Hanqing Luo, Xin Qiu 0008, Liping Liang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |