Xiaodian Cheng

dblp:298/5195 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0002-3696-064XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Accelerating Privacy-Preserving Machine Learning With GeniBatch
abstract
Cross-silo privacy-preserving machine learning (PPML) adopt; Partial Homomorphic Encryption (PHE) for secure data combination and high-quality model training across multiple organizations (e.g., medical and financial). However, PHE introduces significant computation and communication overheads due to data inflation. Batch optimization is an encouraging direction to mitigate the problem by compressing multiple data into a single ciphertext. While promising, it is impractical for a large number of cross-silo PPML applications due to the limited vector operations support and severe data corruption.
Junxue Zhang 0001, Xiaodian Cheng, Hong Zhang 0025, Yilun Jin, Shuihai Hu, Han Tian, Kai Chen 0005
EuroSys3
2024 Accelerating Neural Recommendation Training with Embedding Scheduling
Chaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian, Xinchen Wan, Hao Wang 0116, Kai Chen 0005
NSDI3
2024 High-Performance Hardware Acceleration Architecture for Cross-Silo Federated Learning
abstract
Cross-silo federated learning (FL) adopts various cryptographic operations to preserve data privacy, which introduces significant performance overhead. In this paper, we identify nine widely-used cryptographic operations and design an efficient hardware architecture to accelerate them. However, directly offloading them on hardware statically leads to (1) inadequate hardware acceleration due to the limited resources allocated to each operation; (2) insufficient resource utilization, since different operations are used at different times. To address these challenges, we propose FLASH, a high-performance hardware acceleration architecture for cross-silo FL systems. At its heart, FLASH extracts two basic operators—modular exponentiation and multiplication—behind the nine cryptographic operations and implements them as highly-performant engines to achieve adequate acceleration. Furthermore, it leverages a dataflow scheduling scheme to dynamically compose different cryptographic operations based on these basic engines to obtain sufficient resource utilization. We have implemented a fully-functional FLASH prototype with Xilinx VU13P FPGA and integrated it with FATE, the most widely-adopted cross-silo FL framework. Experimental results show that, for the nine cryptographic operations, FLASH achieves up to$14.0\times$and$3.4\times$acceleration over CPU and GPU, translating to up to$6.8\times$and$2.0\times$speedup for realistic FL applications, respectively. We finally evaluate the FLASH design as an ASIC, and it achieves$23.6\times$performance improvement upon the FPGA prototype.
Junxue Zhang 0001, Xiaodian Cheng, Liu Yang 0008, Jinbin Hu 0001, Han Tian, Kai Chen 0005
IEEE Trans. Parallel Distributed Syst.2
2023 Communication Efficient Secret Sharing with Dynamic Communication-Computation Conversion
abstract
Secret Sharing (SS) is widely adopted in secure Multi-Party Computation (MPC) with its simplicity and computational efficiency. However, SS-based MPC protocol introduces significant communication overhead due to interactive operations on secret sharings over the network. For instance, training a neural network model with SS-based MPC may incur tens of thousands of communication rounds among parties, making it extremely hard for real-world deployment.To reduce the communication overhead of SS, prior works statically convert interactive operations to equivalent non-interactive operations with extra computation cost. However, we show that such static conversion misses chances for optimization, and further present SOLAR, an SS-based MPC framework that aims to reduce the communication overhead through dynamic communication-computation conversion. At its heart, SOLAR converts interactive operations that involve communication among parties to equivalent non-interactive operations within each party with extra computations and introduces a speculative strategy to perform opportunistic conversion when CPU is idle for network transmission. We have implemented and evaluated SOLAR on several popular MPC applications, and achieved 1.6-8.1 times speedup in multi-thread setting compared to the basic SS and 1.2-8.6 times speedup over static conversion.
Zhenghang Ren, Xiaodian Cheng, Mingxuan Fan, Junxue Zhang 0001, Cheng Hong 0001
INFOCOM2
2023 FLASH: Towards a High-performance Hardware Acceleration Architecture for Cross-silo Federated Learning
Junxue Zhang 0001, Xiaodian Cheng, Liu Yang 0008, Jinbin Hu 0001, Kai Chen 0005
NSDI2
2022 Herald: An Embedding Scheduler for Distributed Embedding Model Training
abstract
Given the ability to represent categorical features, embedding models have gained great success on many internet services. State-of-the-art training frameworks enable embedding cache in GPU workers to benefit from hardware acceleration while supporting massive category representations (embeddings) in the limited-capacity GPU device memory. However, based on our measurements, naively adopting a cache system in embedding model training leads to non-negligible communications overhead between caches and the global parameter server. We observe that many such communications are avoidable, given the predictability and sparsity natures of embedding cache accesses in distributed training.
Chaoliang Zeng, Xiaodian Cheng, Han Tian, Hao Wang 0116, Kai Chen 0005
APNet2