EDBT 2026 Demo / reviewers in the wild / expert
Xueyu Wu 0001
dblp:252/6997-1
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-3325-0682ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Framework for Developing and Optimizing Fully Homomorphic Encryption Programs on GPUsabstractIn sensitive domains such as healthcare and finance, machine learning increasingly employs Fully Homomorphic Encryption (FHE) to secure both user data and models. Although FHE's intrinsic parallelism naturally aligns with GPU architectures, optimizing GPU kernels alone remains insufficient for efficient end-to-end FHE application development. The inherent complexity of FHE schemes and intricate GPU-specific details impede developers from focusing on high-level program logic. Additionally, FHE's high memory requirements, fine-grained memory operations, and redundant computations introduce further optimization challenges, resulting in inefficiencies even when GPU kernels are individually optimized. This paper introduces EasyFHE, a framework designed to simplify the development and optimization of GPU-accelerated FHE applications. Similar to PyTorch, EasyFHE provides high-level interfaces for defining computational logic while automatically handling low-level tasks, such as implementation selection and memory management. Furthermore, it incorporates an optimization framework that systematically addresses performance bottlenecks by applying tailored optimization passes during the lowering from high-level FHE programs to GPU kernels. Compared to state-of-the-art open-source GPU FHE libraries, EasyFHE uniquely supports FHE programs with memory requirements exceeding typical GPU capacities, achieving an average speedup of 2.88× with a peak of 4.39×. Jianyu Zhao 0004, Xueyu Wu 0001, Guang Fan 0001, Mingzhe Zhang 0005, Shoumeng Yan, Lei Ju 0001, Zhuoran Ji |
ASPLOS (2) | 2 |
| 2026 | Pipelonk: Accelerating End-to-End Zero-Knowledge Proof Generation on GPUs for PLONK-Based ProtocolsabstractZero-knowledge proofs (ZKPs) are cryptographic protocols that allow verification of statements without disclosing the underlying information. Among them, PLONK-based ZKPs are particularly notable for offering succinct, non-interactive proofs of knowledge with a universal trusted setup, leading to widespread adoption in blockchain and cryptocurrency applications. Nonetheless, their broader deployment is hindered by long proof-generation times and substantial memory demands. While GPUs can accelerate these computations, their limited memory capacity introduces significant challenges for efficient end-to-end proof generation. Zhiyuan Zhang 0008, Yanxin Cai, Wenhao Yin, Xueyu Wu 0001, Yi Wang 0003, Lei Ju 0001, Zhuoran Ji |
PPoPP | 4 |
| 2023 | Embedding Communication for Federated Graph Neural Networks with Privacy GuaranteesabstractGraph Neural Networks (GNNs) have been widely used in many Machine Learning (ML) tasks as they show remarkable performance in modeling graph structure data. While several distributed GNN frameworks have been proposed to tackle the training of huge graphs, uploading the local graph data to the central server for model training is impractical in real-world scenarios due to privacy concerns. Federated Learning (FL) is introduced as an effective technology to address the privacy issue, allowing edge clients to collaboratively train the ML models locally. However, GNNs follow a recursive neighborhood aggregation scheme. Computing the representation vector, also known as embedding, of one node requires aggregating feature vectors of its neighbors. Training the model based on local subgraphs would suffer from information loss and result in accuracy degradation. This paper presents EmbC-FGNN, an efficient Federated _Graph Neural Network framework that enables Node Embedding Communication among training clients in a privacy-preserving way. EmbC-FGNN first proposes an Embedding Server (ES) to maintain and synchronize the shared embeddings among edge workers. It allows training devices to expand the local subgraphs with exchanged embeddings to improve the model accuracy without revealing local node features and graph topology. To minimize the communication costs of the ES, we introduce a periodic embedding synchronization strategy to reduce the communication frequency. Furthermore, we apply asynchronous training to accelerate the convergence speed. Experimental results on several graph neural networks and datasets demonstrate that EmbC-FGNN can improve the overall accuracy (more than 10% for Reddit dataset) and achieve good round-to-accuracy performance. Xueyu Wu 0001, Zhuoran Ji, Cho-Li Wang |
ICDCS | 1 |
| 2022 | KAFL: Achieving High Training Efficiency for Fast-K Asynchronous Federated LearningabstractFederated Averaging (FedAvg) and its variants are prevalent optimization algorithms adopted in Federated Learning (FL) as they show good model convergence. However, such optimization methods are mostly running in a synchronous flavor which is plagued by the straggler problem, especially in the real-world FL scenario. Federated learning involves a massive number of resource-weak edge devices connected to the intermittent networks, exhibiting a vastly heterogeneous training environment. The asynchronous setting is a plausible solution to fulfill the resources utilization. Yet, due to data and device heterogeneity, the training bias and model staleness dramatically downgrade the model performance. This paper presents KAFL, a fast-K Asynchronous Federated Learning framework, to improve the system and statistical efficiency. KAFL allows the global server to iteratively collect and aggregate (1) the parameters uploaded by the fastest K edge clients (K-FedAsync); or (2) the first M updated parameters sent from any clients (Mstep-FedAsync). Compared to the fully asynchronous setting, KAFL helps the server obtain a better direction toward the global optima as it collects the information from at least K clients or M parameters. To further improve the convergence speed of KAFL, we propose a new weighted aggregation method which dynamically adjusts the aggregation weights according to the weight deviation matrix and client contribution frequency. Experimental results show that KAFL achieves a significant time-to-target-accuracy speedup on both IID and Non-IID datasets. To achieve the same model accuracy, KAFL reduces more than 50% training time for five CNN and RNN models, demonstrating the high training efficiency of our proposed framework. Xueyu Wu 0001, Cho-Li Wang |
ICDCS | 1 |
| 2021 | FedSCR: Structure-Based Communication Reduction for Federated LearningabstractFederated Learning allows edge devices to collaboratively train a shared model on their local data without leaking user privacy. The non-independent-and-identically-distributed (Non-IID) property of data distribution, which leads to severe accuracy degradation, and enormous communication overhead for aggregating parameters should be tackled in federated learning. In this article, we conduct a detailed analysis of parameter updates on the Non-IID datasets and compare the difference with the IID setting. Experimental results exhibit that parameter update matrices are structure-sparse and show that more gradients could be identified as negligible updates on the Non-IID data. As a result, we propose a structure-based communication reduction algorithm, called FedSCR, that reduces the number of parameters transported through the network while maintaining the model accuracy. FedSCR aggregates the parameter updates over channels and filters, identifies and removes the redundant updates by comparing the aggregated values with a threshold. Unlike the traditional structured pruning methods, FedSCR retains the complete model that does not require to be retrained and fine-tuned. The local loss and weight divergence on each device vary a lot because of the unbalanced data distribution. We further propose an adaptive FedSCR, that dynamically changes the bounded threshold, to enhance the model robustness on the Non-IID data. Evaluation results show that our proposed strategies achieve almost 50 percent upstream communication reduction without loss of accuracy. FedSCR can be integrated into state-of-the-art federated learning algorithms to dramatically reduce the number of parameters pushed to the global server with a tolerable accuracy reduction. Xueyu Wu 0001, Xin Yao 0008, Cho-Li Wang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | FluentPS: A Parameter Server Design with Low-frequency Synchronization for Distributed Deep LearningabstractWith pursuing high accuracy on big datasets, current research prefers designing complex neural networks, which need to maximize data parallelism for short training time. Many distributed deep learning systems, such as MXNet and Petuum, widely use parameter server framework with relaxed synchronization models. Although these models could cost less on each synchronization, its frequency is still high among many workers, e.g., the soft barrier introduced by Stale Synchronous Parallel (SSP) model. In this paper, we introduce our parameter server design, namely FluentPS, which can reduce frequent synchronization and optimize communication overhead in a large-scale cluster. Different from using a single scheduler to manage all parameters' synchronization in some previous designs, our system allows each server to independently adjust schemes for synchronizing its parameter shard and overlaps the push and pull processes of different servers. We also explore two methods to improve the SSP model: (1) lazy execution of buffered pull requests to reduce the synchronization frequency and (2) a probability-based strategy to pause the fast worker at a probability under SSP condition, which avoids unnecessary waiting of fast workers. We evaluate ResNet-56 with the same large batch size at different cluster scales. While guaranteeing robust convergence, FluentPS gains up to 6× speedup and reduce 93.7% communication time costs than PS-Lite. The raw SSP model causes up to 131× delayed pull requests than our improved synchronization model, which can provide fine-tuned staleness controls and achieve higher accuracy. Xin Yao 0008, Xueyu Wu 0001, Cho-Li Wang |
CLUSTER | 2 |