Benben Liu

dblp:120/8684 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0009-8300-9740ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A Dynamic Virtual Memory Management System for LLMs on AI Chips
Gaolin Wei, Chen Zhang 0013, Xin Yao 0008, Benben Liu
ICCD6
2025 FedEFsz: Fair Cross-Silo Federated Learning System With Error-Bounded Lossy Compression
abstract
Cross-Silo federated learning systems have been identified as an efficient approach to scaling DNN training across geographically-distributed data silos to preserve the privacy of the training data. Communication efficiency and fairness are two major issues that need to be both satisfied when federated learning systems are deployed in practice. Simultaneously guaranteeing both of them, however, is exceptionally difficult because simply combining communication reduction and fairness optimization approaches often causes non-converged training or drastic accuracy degradation. To bridge this gap, we proposeFedEFsz. On the one hand, it integrates the state-of-the-art error-bounded lossy compressor SZ3 into cross-silo federated learning systems to significantly reduce communication traffic during the training. On the other hand, it achieves a high fairness (i.e., rather consistent model accuracy and performance across different clients) through a carefully designed heuristic algorithm that can tune the error-bound of SZ3 for different clients during the training. Extensive experimental results based on a GPU cluster with 65 GPU cards show thatFedEFszimproves the fairness across different benchmarks by up to$60.88\%$and meanwhile reduces the communication traffic by up to$315\times$.
Sheng Di, Benben Liu, Zhuoran Ji, Guanpeng Li, Xiaoyi Lu 0001, Amelie Chi Zhou, Khalid Ayedh Alharthi, Jiannong Cao 0001
IEEE Trans. Parallel Distributed Syst.3
2025 FedCSpc: A Cross-Silo Federated Learning System With Error-Bounded Lossy Parameter Compression
abstract
Cross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. To bridge this gap, we proposeFedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designedFedCSpcproposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show thatFedCSpccan achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4Gb size model,FedCSpcsignificantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality).
Sheng Di, Kai Zhao 0008, Sian Jin, Dingwen Tao, Zhuoran Ji, Benben Liu, Khalid Ayedh Alharthi, Jiannong Cao 0001, Franck Cappello
IEEE Trans. Parallel Distributed Syst.7
2024 FedFa: A Fully Asynchronous Training Paradigm for Federated Learning
Sheng Di, Benben Liu, Khalid Ayedh Alharthi, Jiannong Cao 0001
IJCAI4
2014 GPU-based biclustering for microarray data analysis in neurocomputing
Benben Liu, Yao Xin, Ray C. C. Cheung, Hong Yan 0001
Neurocomputing1
2014 Design Exploration of Geometric Biclustering for Microarray Data Analysis in Data Mining
abstract
Biclustering is an important technique in data mining for searching similar patterns. Geometric biclustering (GBC) method is used to reduce the complexity of the NP-complete biclustering algorithm. This paper studies three commonly used modern platforms including multi-core CPU, GPU and FPGA to accelerate this GBC algorithm. By analyzing the parallelizing property of the GBC algorithm, we design 1) a multi-threaded software running on a server grade multi-core CPU system, 2) a CUDA program for GPU to accelerate the GBC algorithm, and 3) a novel parameterizable and scalable hardware architecture implemented on an FPGA. Genes microarray pattern analysis is employed as an example to demonstrate performance comparisons on different platforms. In particular, we compare the speed and energy efficiency of the three proposed methods. We found that 1) GPU achieves the highest average speedup of 48 × compared to single-threaded GBC program, 2) Our FPGA design can achieve higher speedup of 4 × for the computation for large microarray, and 3) FPGA consumes the least energy, which is about 3.53 × more efficient than the single-threaded GBC program.
Benben Liu, Chi Wai Yu, Doris Z. Wang, Ray C. C. Cheung, Hong Yan 0001
IEEE Trans. Parallel Distributed Syst.1
2012 GPU-Based Biclustering for Neural Information Processing
Alan W. Y. Lo, Benben Liu, Ray C. C. Cheung
ICONIP (5)2