Weiliang Ma

dblp:260/5851 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HieraNTT: A Memory Hierarchy-Aware Data Access Architecture for Efficient Number Theoretic Transform on GPU
Qian Xiong, Weiliang Ma, Ligang He, Yufan Bai, Yao Chen 0008, Hai Jin 0001, Xuanhua Shi
APPT2
2025 gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography
abstract
Elliptic Curve Cryptography (ECC) is an encryption method that provides security comparable to traditional techniques like Rivest–Shamir–Adleman (RSA) but with lower computational complexity and smaller key sizes, making it a competitive option for applications such as blockchain, secure multi-party computation, and database security. However, the throughput of ECC is still hindered by the significant performance overhead associated with elliptic curve (EC) operations, which can affect their efficiency in real-world scenarios. This article presents gECC , a versatile framework for ECC optimized for GPU architectures, specifically engineered to achieve high-throughput performance in EC operations. To maximize throughput, gECC incorporates batch-based execution of EC operations and microarchitecture-level optimization of modular arithmetic. It employs Montgomery’s trick [ 40 ] to enable batch EC computation and incorporates novel computation parallelization and memory management techniques to maximize the computation parallelism and minimize the access overhead of GPU global memory. Furthermore, we analyze the primary bottleneck in modular multiplication by investigating how the user codes of modular multiplication are compiled into hardware instructions and what these instructions’ issuance rates are. We identify that the efficiency of modular multiplication is highly dependent on the number of Integer Multiply-Add (IMAD) instructions. To eliminate this bottleneck, we propose novel techniques to minimize the number of IMAD instructions by leveraging predicate registers to pass the carry information and using addition and subtraction instructions (IADD3) to replace IMAD instructions. Our experimental results show that, for ECDSA and ECDH, the two commonly used ECC algorithms, gECC can achieve performance improvements of 5.56 × and 4.94 ×, respectively, compared to the state-of-the-art GPU-based system. In a real-world blockchain application, we can achieve performance improvements of 1.56 ×, compared to the state-of-the-art CPU-based system. gECC is completely and freely available at https://github.com/CGCL-codes/gECC .
Qian Xiong, Weiliang Ma, Xuanhua Shi, Yongluan Zhou, Hai Jin 0001, Haozhou Wang, Zhengru Wang
ACM Trans. Archit. Code Optim.2
2025 Corrigendum: gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography
abstract
This is a corrigendum for the article “gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography” published in ACM Trans. Arch. Code Optim. 22, 3, Article 84 (September 2025), 27 pages.
Qian Xiong, Weiliang Ma, Xuanhua Shi, Yongluan Zhou, Hai Jin 0001, Haozhou Wang, Zhengru Wang
ACM Trans. Archit. Code Optim.2
2023 GZKP: A GPU Accelerated Zero-Knowledge Proof System
abstract
Zero-knowledge proof (ZKP) is a cryptographic protocol that allows one party to prove the correctness of a statement to another party without revealing any information beyond the correctness of the statement itself. It guarantees computation integrity and confidentiality, and is therefore increasingly adopted in industry for a variety of privacy-preserving applications, such as verifiable outsource computing and digital currency.
Weiliang Ma, Qian Xiong, Xuanhua Shi, Xiaosong Ma, Hai Jin 0001, Haozhao Kuang, Mingyu Gao 0001, Ye Zhang 0042, Haichen Shen, Weifang Hu
ASPLOS (2)1
2020 Capuchin: Tensor-based GPU Memory Management for Deep Learning
abstract
In recent years, deep learning has gained unprecedented success in various domains, the key of the success is the larger and deeper deep neural networks (DNNs) that achieved very high accuracy. On the other side, since GPU global memory is a scarce resource, large models also pose a significant challenge due to memory requirement in the training process. This restriction limits the DNN architecture exploration flexibility.
Xuanhua Shi, Hulin Dai, Hai Jin 0001, Weiliang Ma, Qian Xiong, Fan Yang 0024, Xuehai Qian
ASPLOS5