Qian Xiong

dblp:34/7608 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HieraNTT: A Memory Hierarchy-Aware Data Access Architecture for Efficient Number Theoretic Transform on GPU
Qian Xiong, Weiliang Ma, Ligang He, Yufan Bai, Yao Chen 0008, Hai Jin 0001, Xuanhua Shi
APPT1
2025 gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography
abstract
Elliptic Curve Cryptography (ECC) is an encryption method that provides security comparable to traditional techniques like Rivest–Shamir–Adleman (RSA) but with lower computational complexity and smaller key sizes, making it a competitive option for applications such as blockchain, secure multi-party computation, and database security. However, the throughput of ECC is still hindered by the significant performance overhead associated with elliptic curve (EC) operations, which can affect their efficiency in real-world scenarios. This article presents gECC , a versatile framework for ECC optimized for GPU architectures, specifically engineered to achieve high-throughput performance in EC operations. To maximize throughput, gECC incorporates batch-based execution of EC operations and microarchitecture-level optimization of modular arithmetic. It employs Montgomery’s trick [ 40 ] to enable batch EC computation and incorporates novel computation parallelization and memory management techniques to maximize the computation parallelism and minimize the access overhead of GPU global memory. Furthermore, we analyze the primary bottleneck in modular multiplication by investigating how the user codes of modular multiplication are compiled into hardware instructions and what these instructions’ issuance rates are. We identify that the efficiency of modular multiplication is highly dependent on the number of Integer Multiply-Add (IMAD) instructions. To eliminate this bottleneck, we propose novel techniques to minimize the number of IMAD instructions by leveraging predicate registers to pass the carry information and using addition and subtraction instructions (IADD3) to replace IMAD instructions. Our experimental results show that, for ECDSA and ECDH, the two commonly used ECC algorithms, gECC can achieve performance improvements of 5.56 × and 4.94 ×, respectively, compared to the state-of-the-art GPU-based system. In a real-world blockchain application, we can achieve performance improvements of 1.56 ×, compared to the state-of-the-art CPU-based system. gECC is completely and freely available at https://github.com/CGCL-codes/gECC .
Qian Xiong, Weiliang Ma, Xuanhua Shi, Yongluan Zhou, Hai Jin 0001, Haozhou Wang, Zhengru Wang
ACM Trans. Archit. Code Optim.1
2025 Corrigendum: gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography
abstract
This is a corrigendum for the article “gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography” published in ACM Trans. Arch. Code Optim. 22, 3, Article 84 (September 2025), 27 pages.
Qian Xiong, Weiliang Ma, Xuanhua Shi, Yongluan Zhou, Hai Jin 0001, Haozhou Wang, Zhengru Wang
ACM Trans. Archit. Code Optim.1
2024 MixDehazeNet: Mix Structure Block For Image Dehazing Network
abstract
Image dehazing is a typical task in the low-level vision field. Previous studies verified the effectiveness of vanilla convolution kernel, transformer, and attention mechanism in dehazing. However, there are two main drawbacks in those methods: vanilla convolution and transformer have the short-comings of the insufficient receptive field and a large number of parameters respectively, and the previous design of the attention mechanism does not sufficiently consider an uneven hazy distribution. In this paper, a novel framework named Mix Structure Image Dehazing Network (MixDehazeNet) is proposed to solve the two issues mentioned above. Specifically, it mainly consists of two parts: the multi-scale parallel large convolution kernel module and the enhanced parallel attention module. Compared with a single vanilla kernel or transformer, parallel large kernels with multi-scale have a large receptive field and a relatively smaller amount of parameters, and the multi-scale characteristics of the image. It can restore a single pixel based on a large range of surrounding pixels and simultaneously recover texture details while capturing large hazed areas. In addition, an enhanced parallel attention module is designed according to atmospheric scattering models, which can extract shared global information and location-dependent local information of the original feature in parallel. It performs better at uneven hazy distribution. Extensive experiments on five benchmarks demonstrate the amazing effectiveness of our proposed methods. We achieved or approached state-of-the-art performance in five standard datasets. The code is released in https://github.com/AmeryXiong/MixDehazeNet.
Qian Xiong, BingRong Xu, Duanfeng Chu
IJCNN2
2023 GZKP: A GPU Accelerated Zero-Knowledge Proof System
abstract
Zero-knowledge proof (ZKP) is a cryptographic protocol that allows one party to prove the correctness of a statement to another party without revealing any information beyond the correctness of the statement itself. It guarantees computation integrity and confidentiality, and is therefore increasingly adopted in industry for a variety of privacy-preserving applications, such as verifiable outsource computing and digital currency.
Weiliang Ma, Qian Xiong, Xuanhua Shi, Xiaosong Ma, Hai Jin 0001, Haozhao Kuang, Mingyu Gao 0001, Ye Zhang 0042, Haichen Shen, Weifang Hu
ASPLOS (2)2
2022 Reveal training performance mystery between TensorFlow and PyTorch in the single GPU environment
Hulin Dai, Xuanhua Shi, Ligang He, Qian Xiong, Hai Jin 0001
Sci. China Inf. Sci.5
2020 Capuchin: Tensor-based GPU Memory Management for Deep Learning
abstract
In recent years, deep learning has gained unprecedented success in various domains, the key of the success is the larger and deeper deep neural networks (DNNs) that achieved very high accuracy. On the other side, since GPU global memory is a scarce resource, large models also pose a significant challenge due to memory requirement in the training process. This restriction limits the DNN architecture exploration flexibility.
Xuanhua Shi, Hulin Dai, Hai Jin 0001, Weiliang Ma, Qian Xiong, Fan Yang 0024, Xuehai Qian
ASPLOS6