EDBT 2026 Demo / reviewers in the wild / expert
Zhuoran Ji
dblp:230/8033
· DBLP profile ↗
18ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0001-9767-2767ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 10 first-author · 17 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Framework for Developing and Optimizing Fully Homomorphic Encryption Programs on GPUsabstractIn sensitive domains such as healthcare and finance, machine learning increasingly employs Fully Homomorphic Encryption (FHE) to secure both user data and models. Although FHE's intrinsic parallelism naturally aligns with GPU architectures, optimizing GPU kernels alone remains insufficient for efficient end-to-end FHE application development. The inherent complexity of FHE schemes and intricate GPU-specific details impede developers from focusing on high-level program logic. Additionally, FHE's high memory requirements, fine-grained memory operations, and redundant computations introduce further optimization challenges, resulting in inefficiencies even when GPU kernels are individually optimized. This paper introduces EasyFHE, a framework designed to simplify the development and optimization of GPU-accelerated FHE applications. Similar to PyTorch, EasyFHE provides high-level interfaces for defining computational logic while automatically handling low-level tasks, such as implementation selection and memory management. Furthermore, it incorporates an optimization framework that systematically addresses performance bottlenecks by applying tailored optimization passes during the lowering from high-level FHE programs to GPU kernels. Compared to state-of-the-art open-source GPU FHE libraries, EasyFHE uniquely supports FHE programs with memory requirements exceeding typical GPU capacities, achieving an average speedup of 2.88× with a peak of 4.39×. Jianyu Zhao 0004, Xueyu Wu 0001, Guang Fan 0001, Mingzhe Zhang 0005, Shoumeng Yan, Lei Ju 0001, Zhuoran Ji |
ASPLOS (2) | 7 |
| 2026 | Pipelonk: Accelerating End-to-End Zero-Knowledge Proof Generation on GPUs for PLONK-Based ProtocolsabstractZero-knowledge proofs (ZKPs) are cryptographic protocols that allow verification of statements without disclosing the underlying information. Among them, PLONK-based ZKPs are particularly notable for offering succinct, non-interactive proofs of knowledge with a universal trusted setup, leading to widespread adoption in blockchain and cryptocurrency applications. Nonetheless, their broader deployment is hindered by long proof-generation times and substantial memory demands. While GPUs can accelerate these computations, their limited memory capacity introduces significant challenges for efficient end-to-end proof generation. Zhiyuan Zhang 0008, Yanxin Cai, Wenhao Yin, Xueyu Wu 0001, Yi Wang 0003, Lei Ju 0001, Zhuoran Ji |
PPoPP | 7 |
| 2025 | Accelerating Number Theoretic Transform with Multi-GPU Systems for Efficient Zero Knowledge ProofabstractZero-knowledge proofs validate statements without revealing any information, pivotal for applications such as verifiable outsourcing and digital currencies. However, their broad adoption is limited by the prolonged proof generation times, mainly due to two operations: Multi-Scalar Multiplication (MSM) and Number Theoretic Transform (NTT). While MSM has been efficiently accelerated using multi-GPU systems, NTT has not, due to the high inter-GPU communication overhead incurred by its permutation data access pattern. Zhuoran Ji, Jianyu Zhao 0004, Peimin Gao, Xiangkai Yin, Lei Ju 0001 |
ASPLOS (1) | 1 |
| 2025 | TensorNTT: Architecture-Aware Optimizations for Number-Theoretic Transform on Tensor Core Unit
Xiangkai Yin, Shuoyu Wang, Zimeng Zhou, Lei Ju 0001, Zhuoran Ji |
IEEE Big Data | 6 |
| 2025 | Cube-fx: Mapping Taylor Expansion Onto Matrix Multiplier-Accumulators of Huawei Ascend AI ProcessorsabstractTaylor expansion, a mature method for function evaluations used in Artificial Intelligence (AI) applications, approximates functions with polynomials. In addition to the function evaluations, AI applications require massive matrix multiplications, inspiring manufacturers to propose AI processors with matrix multiplier-accumulators (MACs). However, compared with the powerful Matrix MACs, the vectorized units of the AI processors cannot efficiently carry the existing Taylor expansion implementation of Single Instruction Multiple Data (SIMD) parallelism. Leveraging the Matrix MACs for Taylor expansion becomes an ideal direction. In previous studies, migrating optimized algorithms to the Matrix MACs requires matrix generation during the runtime. The generation is expensive and even cancels the accelerations brought by the Matrix MACs on the AI processors, which Taylor expansion also suffers. This article presents Cube-fx, a mapping algorithm of Taylor expansion for multiple functions onto Matrix MACs. Cube-fx expresses the building and computation in matrix multiplications without inefficient dynamic matrix generation. On Huawei Ascend processors, Cube-fx averagely achieves 1.64× speedups compared with vectorized Horner's Method with 56.38$\%$vectorized operations reduced. Yifeng Tang, Huaman Zhou, Zhuoran Ji, Cho-Li Wang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | FedEFsz: Fair Cross-Silo Federated Learning System With Error-Bounded Lossy CompressionabstractCross-Silo federated learning systems have been identified as an efficient approach to scaling DNN training across geographically-distributed data silos to preserve the privacy of the training data. Communication efficiency and fairness are two major issues that need to be both satisfied when federated learning systems are deployed in practice. Simultaneously guaranteeing both of them, however, is exceptionally difficult because simply combining communication reduction and fairness optimization approaches often causes non-converged training or drastic accuracy degradation. To bridge this gap, we proposeFedEFsz. On the one hand, it integrates the state-of-the-art error-bounded lossy compressor SZ3 into cross-silo federated learning systems to significantly reduce communication traffic during the training. On the other hand, it achieves a high fairness (i.e., rather consistent model accuracy and performance across different clients) through a carefully designed heuristic algorithm that can tune the error-bound of SZ3 for different clients during the training. Extensive experimental results based on a GPU cluster with 65 GPU cards show thatFedEFszimproves the fairness across different benchmarks by up to$60.88\%$and meanwhile reduces the communication traffic by up to$315\times$. Sheng Di, Benben Liu, Zhuoran Ji, Guanpeng Li, Xiaoyi Lu 0001, Amelie Chi Zhou, Khalid Ayedh Alharthi, Jiannong Cao 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | FedCSpc: A Cross-Silo Federated Learning System With Error-Bounded Lossy Parameter CompressionabstractCross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. To bridge this gap, we proposeFedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designedFedCSpcproposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show thatFedCSpccan achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4Gb size model,FedCSpcsignificantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality). Sheng Di, Kai Zhao 0008, Sian Jin, Dingwen Tao, Zhuoran Ji, Benben Liu, Khalid Ayedh Alharthi, Jiannong Cao 0001, Franck Cappello |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2024 | Accelerating Multi-Scalar Multiplication for Efficient Zero Knowledge Proofs with Multi-GPU SystemsabstractZero-knowledge proof is a cryptographic primitive that allows for the validation of statements without disclosing any sensitive information, foundational in applications like verifiable outsourcing and digital currency. However, the extensive proof generation time limits its widespread adoption. Even with GPU acceleration, proof generation can still take minutes, with Multi-Scalar Multiplication (MSM) accounting for about 78.2% of the workload. To address this, we present DistMSM, a novel MSM algorithm tailored for distributed multi-GPU systems. At the algorithmic level, DistMSM adapts Pippenger's algorithm for multi-GPU setups, effectively identifying and addressing bottlenecks that emerge during scaling. At the GPU kernel level, DistMSM introduces an elliptic curve arithmetic kernel tailored for contemporary GPU architectures. It optimizes register pressure with two innovative techniques and leverages tensor cores for specific big integer multiplications. Compared to state-of-the-art MSM implementations, DistMSM offers an average 6.39× speedup across various elliptic curves and GPU counts. An MSM task that previously took seconds on a single GPU can now be completed in mere tens of milliseconds. It showcases the substantial potential and efficiency of distributed multi-GPU systems in ZKP acceleration. Zhuoran Ji, Zhiyuan Zhang 0008, Jiming Xu, Lei Ju 0001 |
ASPLOS (3) | 1 |
| 2024 | A Compiler-Like Framework for Optimizing Cryptographic Big Integer Multiplication on GPUsabstractWith the growth of digital data and rising security concerns, techniques for privacy-preserving computation have become increasingly essential. Big integer multiplication, pivotal for these applications, is compute-intensive but poses challenges for GPU acceleration due to its complexity and the need for application-specific tailored implementations. This paper presents IMCompiler, a compiler-like framework that automatically gen-erates optimized GPU kernels for integer multiplications used in cryptosystems. It features a frontend-IR-backend structure, where the Intermediate Representation (IR) employs a segmented integer multiplication algorithm to decouple architecture-specific optimizations from high-level parameters. The frontend can then easily translate integer multiplication with various high-level parameters into the IR, while the backend focuses on fine-tuning a single GPU kernel for each device, enabling automatic code generation. Moreover, we introduce a computation diagram to facilitate the analysis of parallelization strategies, inspiring many optimizations, including two-dimensional parallelization, tailored caching strategy, index transposing, and lazy carrying. Experiments show that IMCompiler achieves a 4.47× speedup compared to the widely used baseline and 1.42 × over Nvidia's official library. The speedup will be even higher for larger integers and higher-capacity GPUs. Zhuoran Ji, Jianyu Zhao 0004, Jiming Xu, Shoumeng Yan, Lei Ju 0001 |
MICRO | 1 |
| 2024 | POSTER: Accelerating High-Precision Integer Multiplication used in Cryptosystems with GPUsabstractHigh-precision integer multiplication is crucial in privacy-preserving computational techniques but poses acceleration challenges on GPUs due to its complexity and the diverse bit lengths in cryptosystems. This paper introduces GIM, an efficient high-precision integer multiplication algorithm accelerated with GPUs. It employs a novel segmented integer multiplication algorithm that separates implementation details from bit length, facilitating code optimizations. We also present a computation diagram to analyze parallelization strategies, leading to a series of enhancements. Experiments demonstrate that this approach achieves a 4.47× speedup over the commonly used baseline. Zhuoran Ji, Jiming Xu, Lei Ju 0001 |
PPoPP | 1 |
| 2023 | Embedding Communication for Federated Graph Neural Networks with Privacy GuaranteesabstractGraph Neural Networks (GNNs) have been widely used in many Machine Learning (ML) tasks as they show remarkable performance in modeling graph structure data. While several distributed GNN frameworks have been proposed to tackle the training of huge graphs, uploading the local graph data to the central server for model training is impractical in real-world scenarios due to privacy concerns. Federated Learning (FL) is introduced as an effective technology to address the privacy issue, allowing edge clients to collaboratively train the ML models locally. However, GNNs follow a recursive neighborhood aggregation scheme. Computing the representation vector, also known as embedding, of one node requires aggregating feature vectors of its neighbors. Training the model based on local subgraphs would suffer from information loss and result in accuracy degradation. This paper presents EmbC-FGNN, an efficient Federated _Graph Neural Network framework that enables Node Embedding Communication among training clients in a privacy-preserving way. EmbC-FGNN first proposes an Embedding Server (ES) to maintain and synchronize the shared embeddings among edge workers. It allows training devices to expand the local subgraphs with exchanged embeddings to improve the model accuracy without revealing local node features and graph topology. To minimize the communication costs of the ES, we introduce a periodic embedding synchronization strategy to reduce the communication frequency. Furthermore, we apply asynchronous training to accelerate the convergence speed. Experimental results on several graph neural networks and datasets demonstrate that EmbC-FGNN can improve the overall accuracy (more than 10% for Reddit dataset) and achieve good round-to-accuracy performance. Xueyu Wu 0001, Zhuoran Ji, Cho-Li Wang |
ICDCS | 2 |
| 2022 | Optimizing Aggregate Computation of Graph Neural Networks with on-GPU Interpreter-Style ProgrammingabstractGraph Neural Networks (GNNs) generalize deep learning to graph-structured data and show great success in many tasks. However, their irregular aggregation kernels make them inefficient on GPUs. The unpredictable control flow and memory references of irregular kernels prohibit most optimizations designed for regular ones. For example, even if the nodes have overlapped neighbors, reusing them via shared memory is non-trivial, as the neighborhoods used are runtime information. This paper presents regGNN, an aggregation implementation that can benefit from the optimizations designed for regular kernels. It proposes a concept named "semi-regular" to describe the aggregate computation: the irregularity only comes from the neighborhood traversal; aggregating the high-dimensional vectors, which dominates the computation, is data-independent and thus incurs no irregularity. regGNN encodes the aggregate computation steps of each thread block into an aggregate script, which replaces the graph as an input of the GPU kernel. The GPU kernel is like an interpreter, and the aggregate script can be regarded as written in a simple GPU scripting language. The optimizations designed for regular kernels can then be applied to the aggregate script, as it is static and regular. regGNN demonstrates three optimizations: (1) intelligently scheduling nodes and customizing shared memory replacement to maximize data reuse, (2) reassigning nodes among warps for load balancing, and (3) aligning the aggregate script to improve memory latency hiding. Compared with the state-of-the-art GNN frameworks, regGNN achieves 2.81× throughput on average for moderate-scale GNNs. The speedup increases to 5.21× for GNNs with small hidden sizes and 100s × for deep GNNs. Zhuoran Ji, Cho-Li Wang |
PACT | 1 |
| 2022 | Efficient exact K-nearest neighbor graph construction for billion-scale datasets using GPUs with tensor coresabstractApproximate nearest neighbor search plays a fundamental role in many areas, and the k-nearest neighbor graph (KNNG) becomes a promising solution, especially in high-dimensional space. The advantages of KNNG come at the expense of high construction time, which is in quadratic time complexity in the number of points. Many GPUs have adopted specialized hardware units for matrix multiplication, providing an even higher arithmetic throughput. This paper presents flyKNNG, a GPU KNNG construction algorithm for billion-scale datasets. It deploys the distance matrix calculation to matrix multiplication units and adopts on-the-fly top-k selection to avoid transferring the exa-scale distance matrix to/from device memory. flyKNNG co-designs the two key algorithms to optimize the overall performance: the distance matrix calculation algorithm considers the data communication costs and pruning strategy of top-k selection; the top-k selection algorithm is also specially designed for on-the-fly usage, which impacts the data reuse and instruction-level parallelism of the distance matrix calculation as little as possible. Moreover, our top-k selection algorithm is optimized for the special data layout adopted by most matrix multiplication units. Experiments show that flyKNNG achieves 4.67X speedup compared with CUML/FAISS, one of the state-of-the-art approaches. Zhuoran Ji, Cho-Li Wang |
ICS | 1 |
| 2022 | Compiler-Directed Incremental Checkpointing for Low Latency GPU PreemptionabstractGPUs are widely used in data centers to accelerate data-parallel applications. The multiuser and multitasking environment provides a strong incentive for preemptive GPU multitasking, especially for latency-sensitive jobs. Due to the large contexts of GPU kernels, preemptive GPU context switching is costly. Many novel GPU preemption techniques are proposed. Among them, checkpoint-based GPU preemption enables low latency GPU preemption but incurs a high runtime overhead. Prior studies propose to exclude dead registers from the checkpoint file to reduce the runtime overhead. It works well for CPUs, but it is not rare that a live register is not updated between two checkpoints for GPU kernels. This paper presents TripleC, a compiler-directed incremental checkpointing technique specially designed for GPU preemption. It further excludes the registers, which have not been overwritten since the last time they were spilled, from the checkpoint file with data flow analysis. The checkpoint placement algorithm of TripleC can properly estimate a checkpoint's cost under incremental checkpointing. It also considers the interaction among checkpoints so that the overall cost is minimized. Moreover, TripleC relaxes the conventional checkpointing constraint that the whole register context must be spilled before passing the checkpoint. Because of the diverse control flow, placing a register spilling instruction at different points incurs different costs. TripleC minimizes the cost with a two-phase algorithm that schedules these register spilling instructions at compilation time. Evaluations show that TripleC reduces the runtime overhead by 12.9 % on average compared with the state-of-the-art non-incremental checkpointing approach. Zhuoran Ji, Cho-Li Wang |
IPDPS | 1 |
| 2022 | Momentum-driven adaptive synchronization model for distributed DNN training on HPC clusters
Zhuoran Ji, Cho-Li Wang |
J. Parallel Distributed Comput. | 2 |
| 2021 | Collaborative GPU Preemption via Spatial Multitasking for Efficient GPU Sharing
Zhuoran Ji, Cho-Li Wang |
Euro-Par | 1 |
| 2021 | Accelerating DBSCAN Algorithm with AI Chips for Large DatasetsabstractDBSCAN is a popular clustering algorithm, which shows great success in many real-world applications. Its advantages come at the expense of massive computation, especially for computing the distance matrix. Driven by deep learning, many Artificial Intelligence (AI) chips have been developed. With efficient matrix multiplication units, AI chips can significantly accelerate the distance calculation. However, DBSCAN also needs to identify and count the neighbors for each point. It is challenging for most AI chips due to over-specialization. Moreover, the increasing data size and the limited device memory capacity force DBSCAN to follow a mini-batch manner. It results in a high data transfer overhead, which further hinders the performance of DBSCAN on AI chips. In this paper, we propose two novel techniques to address the challenges of accelerating the DBSCAN algorithm with AI chips: (1) new neighbor identification algorithms using bitwise operations only, while traditional solutions require the compare-and-select operations that are weakly supported in AI chips; and (2) two speculative execution strategies to reduce the data transfer overhead induced by mini-batches. Evaluations show that deploying distance matrix calculation to tensor cores achieves 2.61 × speedup on Nvidia RTX 3090. On Huawei Ascend 310, our neighbor identification algorithms achieve 17.88 × throughout of using CPUs for neighbor identification. The speculative execution strategies further reduce the execution time by 15.1% on average for normal datasets and up to 99.0% for sparse datasets. Zhuoran Ji, Cho-Li Wang |
ICPP | 1 |
| 2021 | CTXBack: Enabling Low Latency GPU Context Switching via Context FlashbackabstractEfficient GPU preemption mechanisms are critical for task prioritization in multitasking environments, especially for latency-sensitive GPU applications. However, due to the large context of GPU kernels, simply borrowing context switching mechanisms from CPU space incurs substantial latency and overhead. To address this problem, we propose CTXBack, which allows a thread block to execute context switching at a preceding instruction with a smaller context. It enables low latency GPU context switching for latency-sensitive applications on shared GPUs. Three complementary ways are proposed to make more preceding instructions into valid points to execute context switching, giving CTXBack a higher chance to find instructions with smaller contexts. Evaluations show that CTXBack reduces the context by 61.0%, which is only 1.09× of the minimum possible context size. With only 0.41% runtime overhead, the preemption latency and resuming time are reduced by 63.1% and 50.0% on average compared to the traditional approach. Zhuoran Ji, Cho-Li Wang |
IPDPS | 1 |