VLDB 2026 Research / reviewers in the wild / expert
Qianpiao Ma
dblp:243/3028
· DBLP profile ↗
20ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0001-8684-3495ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 15 · 4 first-author · 14 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NAST: In-Network Aggregation with Worker Selection for Accelerating Distributed Training
Jianfeng Bao, Peng Yang 0022, Gongming Zhao, Huihui Tang, Hongli Xu 0001, Qianpiao Ma |
IWQoS | 6 |
| 2026 | Joint Optimization of Video Recommendation and Cooperative Edge Caching for Maximizing Profit
Haoqiu Luo, Youling Zeng, Yufan Shen, Yue Zeng 0002, Liying Li 0002, Qianpiao Ma, Peijin Cong, Junlong Zhou |
IWQoS | 6 |
| 2026 | FedQuad: Adaptive Layer-Wise LoRA Deployment and Activation Quantization for Federated Fine-TuningabstractFederated fine-tuning (FedFT) provides an effective paradigm for fine-tuning large language models (LLMs) in privacy-sensitive scenarios. However, practical deployment remains challenging due to the limited resources on end devices. Existing methods typically utilize parameter-efficient fine-tuning (PEFT) techniques, such as Low-Rank Adaptation (LoRA), to substantially reduce communication overhead. Nevertheless, significant memory usage for activation storage and computational demands from full backpropagation remain major barriers to efficient deployment on resource-constrained end devices. Moreover, substantial resource heterogeneity across devices results in severe synchronization bottlenecks, diminishing the overall fine-tuning efficiency. To address these issues, we propose FedQuad, a novel LoRA-based FedFT framework that adaptively adjusts the LoRA depth (the number of consecutive tunable LoRA layers from the output) according to devices' computational power, while employing activation quantization to reduce memory overhead, thereby enabling efficient deployment on resource-constrained devices. Specifically, FedQuad first identifies the feasible and efficient combinations of LoRA depth and the number of activation quantization layers based on device-specific resource constraints. Subsequently, FedQuad employs a greedy strategy to select the optimal configurations for each device, effectively accommodating system heterogeneity. Extensive experiments demonstrate that FedQuad achieves a 1.4–5.3× convergence acceleration compared to state-of-the-art baselines when reaching target accuracy, highlighting its efficiency and deployability in resource-constrained and heterogeneous end-device environments. Jianchun Liu, Rukuo Li, Hongli Xu 0001, Qianpiao Ma, Jiaming Yan, Liusheng Huang |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | DySTop: Dynamic Staleness Control and Topology Construction for Asynchronous Decentralized Federated LearningabstractFederated Learning (FL) has emerged as a potential distributed learning paradigm that enables model training on edge devices (i.e., workers) while preserving data privacy. However, its reliance on a centralized server leads to limited scalability. Decentralized federated learning (DFL) eliminates the dependency on a centralized server by enabling peer-to peer model exchange. Existing DFL mechanisms mainly employ synchronous communication, which may result in training inefficiencies under heterogeneous and dynamic edge environments. Although a few recent asynchronous DFL (ADFL) mechanisms have been proposed to address these issues, they typically yield stale model aggregation and frequent model transmission, leading to degraded training performance on non-IID data and high communication overhead. To overcome these issues, we present DySTop, an innovative mechanism that jointly optimizes dynamic staleness control and topology construction in ADFL. In each round, multiple workers are activated, and a subset of their neighbors is selected to transmit models for aggregation, followed by local training. We provide a rigorous convergence analysis for DySTop, theoretically revealing the quantitative relationships between the convergence bound and key factors such as maximum staleness, activating frequency, and data distribution among workers. From the insights of the analysis, we propose a worker activation algorithm (WAA) for staleness control and a phase-aware topology construction algorithm (PTCA) to reduce communication overhead and handle data non-IID. Extensive evaluations through both large-scale simulations and real-world testbed experiments demonstrate that our DySTop reduces completion time by 46.7% and the communication resource consumption by 48.3% compared to state-of-the-art solutions, while maintaining the same model accuracy. Yizhou Shi, Qianpiao Ma, Yan Xu 0027, Junlong Zhou, Ming Hu 0003, Yunming Liao, Hongli Xu 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Toward Communication-Efficient Decentralized Federated Graph Learning Over Non-IID DataabstractDecentralized Federated Graph Learning (DFGL) overcomes the potential bottlenecks of the parameter server in FGL. However, extensive cross-worker communication of graph node embeddings during DFGL training introduces substantial communication costs. To improve communication efficiency, constructing sparse network topologies or applying graph sampling are potential methods. In this paper, we first reveal the bidirectional coupling between network topology construction and graph sampling, underscoring the necessity of their joint optimization. Motivated by this insight, we proposeDuplex, a unified framework that co-optimizes these two components by explicitly modeling their interdependent relationship, thereby significantly reducing communication costs while enhancing training performance in DFGL.Duplexformulates the decision-making process as a coordinated configuration$\langle \mathbf {A}, \mathbf {R} \rangle$, where$\bf {A}$is the adjacency matrix of the network topology and$\bf {R}$denotes the set of graph sampling ratios for workers. However, determining proper coordinated configurations to achieve optimal communication efficiency and training performance (e.g., model accuracy and convergence rate) is challenging due to several practical issues,e.g., statistical heterogeneity and dynamic network conditions. To overcome these challenges,Duplexintroduces a novel learning-driven algorithm to adaptively determine optimal network topologies and graph sampling ratios for workers. Experimental results demonstrate thatDuplexreduces completion time by 20.1%–48.8% and communication costs by 16.7%–37.6% to achieve target accuracy, while improving accuracy by 3.3%–7.9% under identical resource budgets compared to baselines. Shilong Wang 0002, Jianchun Liu, Hongli Xu 0001, Chenxia Tang, Qianpiao Ma, Liusheng Huang |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | Asynchronous Federated Learning Over Non-IID Data via Over-the-Air ComputationabstractFederated learning (FL) enables training AI models across distributed edge devices (i.e., workers) using local data, while facing challenges including communication resource constraints, edge heterogeneity, and non-IID data. Over-the-air computation (AirComp) has emerged as a promising technique to improve communication efficiency by leveraging the superposition property of a wireless multiple access channel (MAC) for model aggregation. However, over-the-air aggregation requires strict synchronization among edge devices, which is essentially incompatible with the asynchronous FL mechanisms often used to handle edge heterogeneity. To overcome this incompatibility, we propose Air-FedGA, a grouping-based asynchronous FL mechanism via AirComp, where workers are organized into groups for synchronized over-the-air aggregation within each group, while groups asynchronously communicate with the parameter server to update the global model. This design retains the communication efficiency of AirComp while addressing training inefficiency caused by edge heterogeneity. We provide a rigorous convergence analysis for Air-FedGA, theoretically quantifying how the convergence bound depends on several key factors, such as the maximum staleness, the degree of non-IID data among groups, and the AirComp aggregation mean squared error (MSE). Guided by these theoretical insights, we propose power control and worker grouping algorithms to minimize the convergence bound by jointly optimizing the AirComp aggregation MSE and the grouping strategy. We conduct experiments on classical models and datasets, and the results demonstrate that our proposed mechanism and algorithms can accelerate the model training by 1.83-$2.22\times $compared with the state-of-the-art solutions. Qianpiao Ma, Xiaozhu Song, Junlong Zhou, Haibo Wang 0004, Yunming Liao, Jianchun Liu, Hongli Xu 0001 |
IEEE Trans. Netw. | 1 |
| 2025 | Accelerating End-Cloud Collaborative Inference via Near Bubble-Free Pipeline Optimization
Luyao Gao, Jianchun Liu, Hongli Xu 0001, Sun Xu, Qianpiao Ma, Liusheng Huang |
INFOCOM | 5 |
| 2025 | Air-FedGA: A Grouping Asynchronous Federated Learning Mechanism Exploiting Over-The-Air ComputationabstractFederated learning (FL) is a new paradigm to train AI models over distributed edge devices (i.e., workers) using their local data, while confronting various challenges including communication resource constraints, edge heterogeneity and data Non-IID. Over-the-air computation (AirComp) is a promising technique to achieve efficient utilization of communication resource for model aggregation by leveraging the superposition property of a wireless multiple access channel (MAC). However, AirComp requires strict synchronization among edge devices, which is hard to achieve in heterogeneous scenarios. In this paper, we propose an AirComp-based grouping asynchronous federated learning mechanism (Air-FedGA), which combines the advantages of AirComp and asynchronous FL to address the communication and heterogeneity challenges. Specifically, AirFedGA organizes workers into groups and performs over-theair aggregation within each group, while groups asynchronously communicate with the parameter server to update the global model. In this way, Air-FedGA accelerates the FL model training by over-the-air aggregation, while relaxing the synchronization requirement of this aggregation technology. We theoretically prove the convergence of Air-FedGA. We formulate a training time minimization problem for Air-FedGA and propose the power control and worker grouping algorithm to solve it, which jointly optimizes the power scaling factors at edge devices, the denoising factors at the parameter server, as well as the worker grouping strategy. We conduct experiments on classical models and datasets, and the results demonstrate that our proposed mechanism and algorithm can speed up FL model training by$\mathbf{29.9\% - 71.6\%}$compared with the state-of-the-art solutions. Qianpiao Ma, Junlong Zhou, Xiangpeng Hou, Jianchun Liu, Hongli Xu 0001, Jianeng Miao, Qingmin Jia |
IPDPS | 1 |
| 2025 | Dynamic task offloading and resource allocation for energy-harvesting end-edge-cloud computing systems
Xiaozhu Song, Qianpiao Ma, Gan Zheng 0002, Liying Li 0002, Peijin Cong, Junlong Zhou |
J. Syst. Archit. | 2 |
| 2025 | FRACTAL: Data-Aware Clustering and Communication Optimization for Decentralized Federated LearningabstractDecentralized federated learning (DFL) is a promising technique to enable distributed machine learning over edge nodes without relying on a centralized parameter server. However, existing DFL network topologies, such as fully connected, partially connected, or lower-tier hierarchical topology often struggle to effectively address the unique challenges presented by edge networks, including edge heterogeneity, communication resource constraint, and data Non-IID. In order to tackle these challenges, we propose a data-aware clustering algorithm, called FRACTAL, to construct a multi-tier hierarchical topology in a bottomup manner taking into consideration both data distribution and communication efficiency for DFL. We theoretically explore the quantitative relationship between the convergence bound of multi-tier FL and the data distribution among each-tier servers. To further improve communication efficiency and address edge heterogeneity, we deploy a time-sharing communication scheduling algorithm within each fractal unit (the basic structure in FRACTAL consisting of multiple nodes and an aggregator), called magic mirror method (MMM), to determine the optimal order of model distributing and uploading for nodes. We conduct extensive experiments on the classical models and datasets to evaluate the performance of FRACTAL, and the results show that FRACTAL can significantly accelerate the DFL model training by 48.6%- 72.3% compared with the state-of-the-art solutions. Qianpiao Ma, Jianchun Liu, Hongli Xu 0001, Qingmin Jia, Renchao Xie |
IEEE Trans. Big Data | 1 |
| 2025 | CADER: Cost-Efficient Cloud Application Deployment With Tenant Requirement Guarantee in Multi-CloudsabstractMotivated by the need to reduce vendor lock-in and address concerns regarding dedicated hardware availability, cloud applications have increasingly adopted a multi-cloud deployment strategy, in which cloud applications are deployed in different zones associated with various cloud service providers. When deploying cloud applications in multi-clouds, there are three crucial and coupled metrics:deployment cost,access delayandtraffic demand. Unfortunately, existing works overlook either the data transfer cost in deployment cost or the access delay and traffic demand requirements, resulting in high operating costs or low user QoS. To bridge this gap, this paper proposes theCost-EfficientApplicationDeployment Framework (CADER) with tenant requirement guarantee in multi-clouds environment. However, due to the challenges of service price heterogeneity, transfer cost diversity, and resource limitation, achieving cost-efficient cloud application deployment while satisfying all tenant requirements is not an easy task. To tackle this issue, we design an approximate algorithm based on the random rounding method and prove that its approximate ratio is$O(\log g)$, where$g$is the number of cloud zones. Results of in-depth simulations indicate that CADER can reduce the application deployment cost ranging from 16% to 38% compared to commonly used alternatives while ensuring the satisfaction of tenant requirements. Huaqing Tu, Ziqiang Hua, Qianpiao Ma, Hanguang Luo, Gongming Zhao, Hongli Xu 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2025 | FedACS: An Adaptive Client Selection Framework for Communication-Efficient Federated Graph LearningabstractFederated graph learning (FGL) has been proposed to collaboratively train the increasing graph data with graph neural networks (GNNs) in a recommendation system. Nevertheless, implementing an efficient recommendation system with FGL still faces two primary challenges, i.e., limited communication bandwidth and non-IID local graph data. Existing works typically reduce communication frequency or transmission amount, which may suffer significant performance degradation under non-IID settings. Furthermore, some researchers propose to share the underlying structure among clients, which brings massive communication cost. To this end, we propose an efficient FGL framework, named FedACS, which adaptively selects a subset of clients for model training, to alleviate communication overhead and non-IID issues simultaneously. In FedACS, the global GNN model learns significant hidden edges and the structure of graph data among selected clients, enhancing recommendation efficiency. This capability distinguishes it from the traditional FL client selection methods. To optimize the client selection process, we introduce a multi-armed bandit (MAB) based algorithm to select participating clients according to the resource budgets and the training performance (i.e., RMSE). Experimental results indicate that FedACS improves RMSE by 5.4% over baselines with the same resource budget and reduces communication costs by up to 70.7% to achieve the same RMSE performance. Hongli Xu 0001, Xianjun Gao, Jianchun Liu, Qianpiao Ma, Liusheng Huang |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Towards Communication-Efficient Federated Graph Learning: An Adaptive Client Selection PerspectiveabstractFederated graph learning (FGL) has been proposed to collaboratively train the increasing graph data with graph neural networks (GNNs) in a recommendation system, aggregating the features of graph nodes and edges among these nodes. Nevertheless, implementing an efficient recommendation system with FGL still faces two primary challenges, i.e., limited communication bandwidth and non-IID local graph data. Existing works typically reduce communication frequency or transmission amount, which may suffer significant performance degradation under non-IID settings. Furthermore, some researchers propose to share the underlying structure information among all clients, which brings massive communication cost. To this end, we propose an efficient FGL framework, named FedACS, which adaptively selects a subset of clients for model training, to alleviate communication overhead and non-IID issues simultaneously. In FedACS, the global GNN model can learn significant hidden edges and the structure of graph data among selected clients, enhancing recommendation efficiency. This capability distinguishes it from the traditional FL client selection methods. To optimize the client selection process, we introduce a multi-armed bandit (MAB) based algorithm to select participating clients according to the resource budgets and the training performance (i.e., RMSE) under different data distributions. Experimental results show that, given the same resource budget, FedACS achieves the RMSE improvement of 5.4% over the baselines. Besides, when achieving the same RMSE performance, FedACS saves up to approximately 70.7% communication cost, compared with the baselines. Xianjun Gao, Jianchun Liu, Hongli Xu 0001, Qianpiao Ma, Lun Wang 0003 |
IWQoS | 4 |
| 2024 | Dynamic Staleness Control for Asynchronous Federated Learning in Decentralized Topology
Qianpiao Ma, Jianchun Liu, Qingmin Jia, Xiaomao Zhou, Yujiao Hu, Renchao Xie |
WASA (2) | 1 |
| 2024 | FedCD: A Hybrid Federated Learning Framework for Efficient Training With IoT DevicesabstractWith billions of IoT devices producing vast data globally, privacy and efficiency challenges arise in AI applications. Federated learning (FL) has been widely adopted to train deep neural networks (DNNs) without privacy leakage. Existing centralized and decentralized FL architectures have limitations, including memory burden, huge bandwidth pressure and non-IID data issues. This paper introduces a novel hybrid FL framework, named FedCD, merging the benefits of both centralized and decentralized FL architectures. FedCD strategically distributes the model based on layer sizes and consensus distances (i.e., the deviation between the local models and the global average models), effectively relieving network bandwidth pressures and accelerating training speed even under the non-IID setting. This method significantly mitigates resource constraints and improves model accuracy, offering a promising solution to the challenges in distributed machine learning. Extensive experiment results show the high effectiveness of FedCD. The total completion time of FedCD is reduced by 16.3%-53% and the average accuracy improvement is 1.85% compared to the baselines. Jianchun Liu, Pengcheng Qu, Sun Xu, Zhi Liu 0002, Qianpiao Ma, Jinyang Huang |
IEEE Internet Things J. | 6 |
| 2024 | FedUC: A Unified Clustering Approach for Hierarchical Federated LearningabstractFederated learning (FL) is an effective approach to train models collaboratively among distributed edge nodes (i.e., workers) while facing three crucial challenges, edge heterogeneity, resource constraint, and Non-IID data. Under the parameter server (PS) architecture, a single parameter server may become the system bottleneck and cannot well deal with the edge heterogeneity, while the peer-to-peer (P2P) architecture causes significant communication consumption to achieve satisfactory training performance. To this end, hierarchical aggregation (HA) architecture is proposed to cluster workers to tackle the edge heterogeneity and reduce communication consumption for FL. However, the existing researches on HA architecture cannot provide a unified clustering approach for various inter-cluster aggregation patterns (e.g., centralized or decentralized structure, synchronous or asynchronous mode). In this paper, we explore the quantitative relationship between the convergence bounds of different inter-cluster patterns and several factors, e.g., data distribution, frequency of clusters participating in inter-cluster aggregation (for asynchronous modes), and inter-cluster topology (for decentralized structures). Based on the convergence bounds, we design a unified clustering algorithm FedUC to organize workers for different patterns. Experimental results on classical models and datasets show that FedUC can greatly accelerate the model training of different patterns by 1.79-7.39× compared with the state-of-the-art clustering methods. Qianpiao Ma, Yang Xu 0020, Hongli Xu 0001, Jianchun Liu, Liusheng Huang |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | YOGA: Adaptive Layer-Wise Model Aggregation for Decentralized Federated LearningabstractTraditional Federated Learning (FL) is a promising paradigm that enables massive edge clients to collaboratively train deep neural network (DNN) models without exposing raw data to the parameter server (PS). To avoid the bottleneck on the PS, Decentralized Federated Learning (DFL), which utilizes peer-to-peer (P2P) communication without maintaining a global model, has been proposed. Nevertheless, DFL still faces two critical challenges, i.e., limited communication bandwidth and not independent and identically distributed (non-IID) local data, thus hindering efficient model training. Existing works commonly assume full model aggregation at periodic intervals, i.e., clients periodically collect models from peers. To reduce the communication cost, these methods allow clients to collect model(s) from selected peers, but often result in a significant degradation of model accuracy when dealing with non-IID data. Alternatively, the layer-wise aggregation mechanism has been proposed to alleviate communication overhead under the PS architecture, but its potential in DFL remains rarely explored yet. To this end, we propose an efficient DFL framework YOGA that adaptively performs layer-wise model aggregation and training. Specifically, YOGA first generates the ranking of layers in the model according to the learning speed and layer-wise divergence. Combining with the layer ranking and peers’ status information (i.e., data distribution and communication capability), we propose the max-match (MM) algorithm to generate the proper layer-wise model aggregation policy for the clients. Extensive experiments on DNN models and datasets show that YOGA saves communication cost by about 45% without sacrificing the model performance compared with the baselines, and provides 1.53-$3.5\times $speedup on the physical platform. Jun Liu 0083, Jianchun Liu, Hongli Xu 0001, Yunming Liao, Zhiyuan Wang 0002, Qianpiao Ma |
IEEE/ACM Trans. Netw. | 6 |
| 2023 | FedCD: A Hybrid Centralized-Decentralized Architecture for Efficient Federated LearningabstractWith billions of IoT devices producing vast data globally, privacy and efficiency challenges arise in AI applications. Federated learning (FL) has been widely adopted to train deep neural networks (DNNs) without privacy leakage. Existing centralized and decentralized FL architectures have limitations, including memory burden, huge bandwidth pressure and non-IID data issues. This paper introduces a novel framework, named FedCD, merging the benefits of both centralized and decentralized FL architectures. FedCD strategically distributes the model based on layer sizes and consensus distances (measuring the deviation between the local models and the global average models), effectively relieving network bandwidth pressures and accelerating training speed even under the non-IID setting. This method significantly mitigates resource constraints and improves model accuracy, offering a promising solution to the challenges in distributed machine learning. Extensive experiment results show the high effectiveness of FedCD. The total completion time of FedCD is reduced by 16.3%-53% and the average accuracy improvement is 1.85% compared to the existing FL systems. Pengcheng Qu, Jianchun Liu, Zhiyuan Wang 0002, Qianpiao Ma, Jinyang Huang |
ICPADS | 4 |
| 2021 | FedSA: A Semi-Asynchronous Federated Learning Mechanism in Heterogeneous Edge ComputingabstractFederated learning (FL) involves training machine learning models over distributed edge nodes (i.e., workers) while facing three critical challenges, edge heterogeneity, Non-IID data and communication resource constraint. In the synchronous FL, the parameter server has to wait for the slowest workers, leading to significant waiting time due to edge heterogeneity. Though asynchronous FL can well tackle the edge heterogeneity, it requires frequent model transfers, resulting in massive communication resource consumption. Moreover, the different relative frequency of workers participating in asynchronous updating may seriously hurt training accuracy, especially on Non-IID data. In this paper, we propose a semi-asynchronous federated learning mechanism (FedSA), where the parameter server aggregates a certain number of local models by their arrival order in each round. We theoretically analyze the quantitative relationship between the convergence bound of FedSA and different factors,e.g., the number of participating workers in each round, the degree of data Non-IID and edge heterogeneity. Based on the convergence bound, we present an efficient algorithm to determine the number of participating workers to minimize the training completion time. To further improve the training accuracy on Non-IID data, FedSA deploys adaptive learning rates for workers by their relative participation frequency. We extend our proposed mechanism to the dynamic and multiple learning tasks scenarios. Experimental results on the testbed show that our proposed mechanism and algorithms address the three challenges more effectively than the state-of-the-art solutions. Qianpiao Ma, Yang Xu 0020, Hongli Xu 0001, Zhida Jiang, Liusheng Huang, He Huang 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2019 | Lightweight Flow Distribution for Collaborative Traffic Measurement in Software Defined NetworksabstractMany important functions in software defined networks can benefit from fine-grained traffic measurement at flow level. Because TCAM-based flow entries only provide aggregate traffic statistics, prior research has suggested to perform flow-level measurement in SRAM and balance the measurement load across the network through collaborative traffic measurement. The key problem of collaborative measurement is to provide a mechanism to distribute flows to switches such that each switch can identify its subset of flows to measure. We observe that the prior work has focused on optimizing flow distribution among switches, but overlooked their high space and per-packet processing overhead introduced to the data plane, which becomes a serious issue in large SDN systems. In this paper, we propose a new lightweight solution to the flow distribution problem. It follows the design principle of alleviating complexity of the data plane by minimizing the data-plane space and processing overhead. At the control plane, we formulate flow distribution as optimization problems under two scenarios that implement collaborative measurement by edge switches only and by edge/core switches together, respectively. Our extensive simulations demonstrate that, comparing with the best existing work, the proposed lightweight solution achieves a comparable performance in terms of load balancing, while drastically reducing both space overhead and per-packet processing overhead, making it more practical in real-world systems that are sensitive to the additional overhead introduced by flow distribution. Hongli Xu 0001, Shigang Chen, Qianpiao Ma, Liusheng Huang |
INFOCOM | 3 |