EDBT 2026 Demo / reviewers in the wild / expert
Changxu Liu
dblp:377/5072
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0000-0132-1461ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Fault-Aware Architecture for Reliable Sparse Matrix Multiplication
Yuxuan Qiao, Changxu Liu, Junjie Zuo, Baoyu Fan, Fan Yang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | TL-CSE: Microarchitecture-Compiler Co-design Space Exploration via Transfer LearningabstractThe system design of domain-specific processors is a challenging task due to the vast architecture-compiler search space and time-consuming simulation processes. Nowadays, architecture space exploration and compiler optimization are conducted in silo. However, due to the interdependence between microarchitecture and compiler, this segmented approach shrinks the hardware-software design space, leading to sub-optimal global outcomes. To solve problems in co-optimization, this paper introduces TL-CSE, a microarchitecture compiler co-design space exploration framework. It utilizes a bi-level optimization framework with transfer learning techniques to efficiently explore the hardware-software system design space. Results demonstrate that TL-CSE improves the quality of the Pareto optimal system by 31.5% and speeds up the exploration time by 6.1 times compared to previous co-design frameworks. Jinyi Shen, Changxu Liu, Li Shang 0001, Fan Yang 0001 |
ASP-DAC | 4 |
| 2025 | AcclMT: A Highly Resource-Efficient and Flexible Poseidon Hash-Based Merkle Tree ArchitectureabstractMerkle Tree is a fundamental cryptographic primitive in Zero-Knowledge Proof (ZKP) protocols, sharing significant computational workloads with the Number Theoretic Transform (NTT) in zkSTARK schemes. Merkle Tree is a tree structure where nodes are primarily generated through hash computations. Among them, Poseidon Hash, as a ZK-friendly hash function, has emerged as one of the most widely adopted choices. Therefore, hardware acceleration of building Merkle Tree based on Poseidon Hash can significantly enhance the performance of ZKP protocols. We propose AcclMT, a highly resourceefficient and flexible Poseidon Hash-based Merkle Tree architecture. Our design employs hardware-software co-design and optimizes the hashing data flow, resulting in an area-efficient Poseidon Hash engine that improves modular multiplication resource utilization. Furthermore, AcclMT uses these engines alongside hierarchical on-chip cache and optimized task scheduling for building large Merkle Trees. It also supports flexible parameter configurations for various requirements. Experimental results show that our proposed Poseidon Hash engine achieves a $14.3 \times$ speedup compared to the latest FPGA-based work. By improving resource utilization, it also reduces area usage by 14.8% compared to unoptimized design. AcclMT achieves up to $1665 \times$ speedup over software implementations in building Merkle tree, with average utilization of 95.9% and 99.2% for the two hash engines. Changxu Liu, Hao Zhou 0015, Zhuoyuan Yang, Yinlong Li, Shiyong Wu, Fan Yang 0001 |
DAC | 1 |
| 2025 | Myosotis: An Efficiently Pipelined and Parameterized Multiscalar Multiplication Architecture via Data SharingabstractZero-knowledge proof (ZKP) is a widely used privacy-preserving technology, where multiscalar multiplication (MSM) accounts for over 70% of the computational workload. The acceleration of MSM can enhance the overall performance of ZKP, making it a focal point of community attention. However, in practical applications involving the deployment of multiple MSM accelerators, existing designs often overlook strategies for optimizing bandwidth and area efficiency. To address this, we propose Myosotis, an efficiently pipelined and parameterized MSM architecture. By sharing input data and allocating cache effectively, it mitigates average transmission bandwidth in runtime. Myosotis also supports the use of multiple point addition (PADD) units to achieve performance gains, balancing area overhead and latency for improved area efficiency. Different parameter selection enables a tradeoff between the performance, area, and bandwidth of the MSM accelerator. When benchmarking with MSM degrees between$2^{18}$and$2^{26}$, our proposed baseline design achieves up to$3.32\times $and$6.72\times $speedups over state-of-the-art FPGA and ASIC designs. Compared to the baseline, Myosotis with two window MSMs and one PADD unit reduces bandwidth demand by 43% while maintaining similar area and latency. On the other hand, Myosotis with three window MSMs and two PADD units decreases latency by 43% and bandwidth by 17%, with only a 9% area increase. Changxu Liu, Hao Zhou 0015, Patrick Dai, Yinlong Li, Shiyong Wu, Fan Yang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | ReZK: A Highly Reconfigurable Accelerator for Zero-Knowledge ProofabstractZero-knowledge proof (ZKP) plays a significant role in privacy protection technology. However, the proof generation phase requires considerable time and hardware resources. In this phase, Number Theoretic Transform or Inverse Number Theoretic Transform (NTT/INTT) in polynomial computation, as well as Multiple Scalar Multiplication (MSM), are bottlenecks that dominate the execution time. In this paper, we propose a highly reconfigurable accelerator ReZK to accelerate ZKP proof generation phase, focusing on NTT/INTT and MSM. According to the configurations, ReZK can be configured as NTT, INTT, and MSM with variable sizes and bit-widths by adjusting the data path between on-chip memories and arithmetic cores. As the basic unit of arithmetic cores, the reconfigurable processing element (PE) in ReZK is composed of pipelined modular multipliers and modular adders that support variable bit-widths. It can perform butterfly or arithmetic operations. Based on the reconfigurable PEs, the ReZK core can implement NTT/INTT with different sizes and bit-widths, or a fully pipelined point adder (PADD). Additionally, we propose a modularized MSM scheduling architecture to support various bit-widths. The on-chip memories are also well organized for reuse. In NTT/INTT mode, 4-way 256-bit or 2-way 384-bit NTT/INTT can be computed in parallel. In MSM mode, for different elliptic curves, ReZK is capable of processing 4-way 256-bit or 2-way 384-bit MSM in parallel. Hao Zhou 0015, Changxu Liu, Li Shang 0001, Fan Yang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Gypsophila: A Scalable and Bandwidth-Optimized Multi-Scalar Multiplication ArchitectureabstractMulti-Scalar Multiplication (MSM) is a fundamental cryptographic primitive, which plays a crucial role in Zero-knowledge proof systems. In this paper, we optimize the single MSM Process Element (PE) utilizing buckets with fewer conflicts, enhanced by Greedy-based scheduling, to achieve higher efficiency. The evaluation results show our optimized single MSM PE achieving a speedup of over two times on average, peaking at 3.63 times compared to previous works. Furthermore, we introduce Gypsophila, a scalable and bandwidth-optimized architecture for implementing multiple MSM PEs. Leveraging the characteristics of the bucket method, we optimize the data flow by balancing the throughput of bucket classification, bucket aggregation, and result aggregation in MSM. Simultaneously, multiple PEs with different data access patterns share a universal point input channel and post-processing unit, which improves the module utilization and mitigates the bandwidth pressure. Gypsophila with 16 PEs, accomplishes 16 MSM tasks in a mere 1.01% additional time, showcasing an approximate 7.8% reduction in area, with only about 116 of the bandwidth requirement, compared with 16 PEs without input channel and post-process unit sharing. Changxu Liu, Hao Zhou 0015, Jiamin Xu, Patrick Dai, Fan Yang 0001 |
DAC | 1 |
| 2024 | HMNTT: A Highly Efficient MDC-NTT Architecture for Privacy-preserving ApplicationsabstractIn privacy-preserving applications like Post-Quantum Cryptography (PQC) and Fully Homomorphic Encryption (FHE), polynomial multiplication is common, and the Number Theoretic Transform (NTT) is a key algorithm for reducing its complexity. In this paper, we present HMNTT, a highly efficient MDC-NTT architecture. Utilizing the four-step NTT algorithm and a pipelined transpose module, HMNTT offers a highly efficient and scalable architecture for handling NTT with large degrees. We optimize the processing element (PE) to alleviate backpressure and data conflicts in data flow. Leveraging FPGA characteristics, we construct a modular multiplication module to reduce resource usage and improve operating frequency. Evaluation results indicate that HMNTT achieves an average of 2.34 × and 1.26 × reduction in Area-Time Product compared to the latest pipelined NTT architectures. Changxu Liu, Danqing Tang, Hao Zhou 0015, Shoumeng Yan, Fan Yang 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | A Fully Pipelined Reconfigurable Montgomery Modular Multiplier Supporting Variable Bit-WidthsabstractRecently, there has been increased emphasis on privacy-preserving computation technologies, such as homomorphic encryption (HE) and zero-knowledge proof (ZKP). Modular multiplication is a critical component for both HE and ZKP. Variable bit-width is a must for many applications of privacy-preserving computation, due to variable bit-width requirements for different cryptography schemes. However, the majority of modular multipliers that support variable bit-width configurations exhibit relatively low throughput. This work presents a fully pipelined Montgomery modular multiplier with variable bit-width support. Truncated multipliers are introduced to reduce the resources of modular multipliers in our approach. In order to meet different bit-width requirements, the proposed modular multiplier can be dynamically reconfigured. The proposed design can support widely used bit-width configurations, specifically, 384-bit, 256-bit, and 128-bit. 256-bit and 128-bit modes support parallel computation of 2 and 6 sets of operands, respectively. Compared with existing variable bit-width modular multipliers, the proposed reconfigurable modular multiplier significantly improves the throughputs with even lower resources. Hao Zhou 0015, Changxu Liu, Li Shang 0001, Fan Yang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | PriorMSM: An Efficient Acceleration Architecture for Multi-Scalar MultiplicationabstractMulti-Scalar Multiplication (MSM) is a computationally intensive task that operates on elliptic curves based on GF(P) . It is commonly used in zero-knowledge proof (ZKP), where it accounts for a significant portion of the computation time required for proof generation. In this article, we present PriorMSM, an efficient acceleration architecture for MSM. We propose a Priority-Based Scheduling Mechanism (PBSM) based on a multi-FIFO and multi-bank architecture to accelerate the implementation of MSM. By increasing the pairing success rate of internal points, PBSM reduces the number of bubbles in the pipeline of point addition (PADD), consequently improving the data throughput of the pipeline. We also introduce an advanced parallel bucket aggregation algorithm, leveraging PADD’s fully pipelined characteristics to significantly accelerate the implementation of bucket aggregation. We perform a sensitivity analysis on the crucial parameter of window size in MSM. The results indicate that the window size of the MSM significantly impacts its latency. Area-Time Product (ATP) metric is introduced to guide the selection of the optimal window size, balancing the performance and cost for practical applications of subsequent MSM implementations. PriorMSM is evaluated using the TSMC 28 nm process. It achieves a maximum speedup of 10.9× compared to the previous custom hardware implementations and a maximum speedup of 3.9× compared to the GPU implementations. Changxu Liu, Hao Zhou 0015, Patrick Dai, Li Shang 0001, Fan Yang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |