VLDB 2026 Research / reviewers in the wild / expert
Zirui Lian
dblp:305/9561
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-5104-1852ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PPFL: A Parameter Behavior-Driven Plug-in Personalization Engine for Federated LearningabstractPersonalized Federated Learning (PFL) customizes models for each client to mitigate challenges from non-IID data, wherein a dominant strategy is model decoupling that partitions models into shared and personalized parts based on architectural priors (e.g., backbone vs. head). However, we reveal a critical flaw in this strategy: it induces "intrinsic drift," a performance degradation often more severe than the well-known client drift, which limits final accuracy. We trace this drift to a steep cliff of high loss emerging from the naive stitching of shared and personalized parts. To address this, we shift from architectural partitioning to a parameter behavior-driven paradigm. We introduce PPFL, an approach that employs a novel soft-fusion strategy guided by parameter-wise behavioral perception. PPFL dynamically infers each parameter's functional role—whether it behaves more like a 'personalist' or a 'generalist' in the current context—by synthesizing its multifaceted behavior observed during local training. Extensive experiments on image, text, and multimodal classification benchmarks show that PPFL outperforms eight state-of-the-art baselines by up to 5.3%. Moreover, it can function as a plug-in module, boosting the accuracy of vanilla FedAvg with a 16.82% absolute gain. Qianyue Cao, Zongwei Zhu, Zirui Lian, Rui Zhang 0040, Boyu Li 0006, Yi Xiong 0003, Xuehai Zhou |
AAAI | 3 |
| 2026 | CSCL: Bridging the plasticity-stability gap in continuous supervised contrastive learning
Yi Xiong 0003, Liqi Xiang, Qianyue Cao, Zongwei Zhu, Zirui Lian, Xuehai Zhou |
Neural Networks | 5 |
| 2025 | HaloFL: Efficient Heterogeneity-Aware Federated Learning Through Optimal Submodel Extraction and Dynamic Sparse AdjustmentabstractFederated learning (FL) is an advanced framework that enables collaborative training of machine learning models across edge devices. An effective strategy to enhance training efficiency is to allocate the optimal submodel based on each device’s resource capabilities. However, system heterogeneity significantly increases the difficulty of allocating submodel parameter budgets appropriately for each device, leading to the straggler problem. Meanwhile, data heterogeneity complicates the selection of the optimal submodel structure for specific devices, thereby impacting training performance. Furthermore, the dynamic nature of edge environments, such as fluctuations in network communication and computational resources, exacerbates these challenges, making it even more difficult to precisely extract appropriately sized and structured submodels from the global model. To address the challenges in heterogeneous training environments, we propose an efficient FL framework, namely, HaloFL. The framework dynamically adjusts the structure and parameter budget of submodels during training by evaluating three dimensions: 1) model-wise performance; 2) layer-wise performance; and 3) unit-wise performance. First, we design a data-aware model unit importance evaluation method to determine the optimal submodel structure for different data distributions. Next, using this evaluation method, we analyze the importance of model layers and reallocate parameters from noncritical layers to critical layers within a fixed parameter budget, further optimizing the submodel structure. Finally, we introduce a resource-aware dual-UCB multiarmed bandit agent, which dynamically adjusts the total parameter budget of submodels according to changes in the training environment, allowing the framework to better adapt to the performance differences of heterogeneous devices. Experimental results demonstrate that HaloFL exhibits outstanding efficiency in various dynamic and heterogeneous scenarios, achieving up to a 14.80% improvement in accuracy and a$3.06\times $speedup compared to existing FL frameworks. Zirui Lian, Qianyue Cao, Zongwei Zhu, Cheng Ji 0002, Changlong Li 0006, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | GOFL: An Accurate and Efficient Federated Learning Framework Based on Gradient Optimization in Heterogeneous IoT SystemsabstractFederated learning (FL) is designed for training models using data distributed across multiple Internet of Things (IoT) devices or servers, reducing data transfer overhead and ensuring data security. However, the decentralization and diversity of IoT devices introduce statistical and system heterogeneity, which can lead to unstable model training and even system crashes. Although many studies attribute performance issues to client-drift caused by this heterogeneity, there is a lack of insight into how different forms of heterogeneity impact local model gradient variations and model convergence. In this article, we investigate model gradient distribution characteristics in heterogeneous training. We find that the challenge is not solely due to client-drift but is also closely linked to a high degree of model overfitting, which negatively affects local model training and equilibrium convergence. To address this challenge, we introduce an efficient framework called gradient optimization with FL (GOFL). First, GOFL incorporates the federated gradient normalization (FGN) technique to maintain gradient distribution consistency while mitigating client-drift stemming from heterogeneity. We also highlight the benefits of FGN in reducing local model overfitting and improving convergence. Second, GOFL introduces the federated device aggregation (FDA) strategy, a critical addition to FGN. It adaptively guides device selection and aggregation based on device contributions, ensuring a more balanced training approach in the face of system heterogeneity. The experimental results demonstrate that GOFL achieves state-of-the-art training accuracy while reducing the number of training rounds. In particular, it improves the accuracy of the classical FL framework FedAvg by 30.57% and reduces the number of convergence rounds by 5.17 times. Zirui Lian, Zongwei Zhu, Xuehai Zhou, Weihong Liu |
IEEE Internet Things J. | 1 |
| 2024 | NebulaFL: Self-Organizing Efficient Multilayer Federated Learning Framework With Adaptive Load Tuning in Heterogeneous Edge SystemsabstractAs a promising edge intelligence technology, federated learning (FL) enables Internet of Things (IoT) devices to train the models collaboratively while ensuring the data privacy and security. Recently, hierarchical FL (HFL) has been designed to promote distributed training in the intricate hierarchical structure of IoT. However, the coarse-grained hierarchical schemes usually fail to thoroughly adapt to the hierarchical environment, leading to high training latency. Meanwhile, highly heterogeneous communication and computation delays due to the device diversity (the system heterogeneity) and decentralized data distribution due to the decentralized device distribution (the data heterogeneity) exacerbate the above challenges. This article proposes NebulaFL, a dual heterogeneity-aware multilayer FL framework, to support efficient distributed training in IoT scenarios. NebulaFL proposes an innovative multilayer architecture organization scheme to adapt the complex hierarchical heterogeneous scenarios. Specifically, through a finer-grained division of the HFL hierarchy, hybrid synchronous-asynchronous training is implemented at both the global system and local device-layer levels. More importantly, to adaptively build a heterogeneity-aware hierarchical training architecture, NebulaFL considers the effect of dual heterogeneity in the architectural organization scheme to determine the optimal location of devices in a multilayer environment. To further improve the training efficiency during the training process, NebulaFL employs an augmented multiarmed bandit technique based on the reinforcement learning to adjust the device-layer training load by evaluating the dynamic training utility and convergence uncertainty feedback. Experiments demonstrate that NebulaFL achieves up to a$15.68\times $speed-up ratio and a 23.94% increase in the training accuracy compared to the latest or classic approaches. Zirui Lian, Qianyue Cao, Weihong Liu, Zongwei Zhu, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Ace-Sniper: Cloud-Edge Collaborative Scheduling Framework With DNN Inference Latency Modeling on Heterogeneous DevicesabstractThe cloud–edge collaborative inference requires efficient scheduling of artificial intelligence (AI) tasks to the appropriate edge intelligence devices. Gls DNN inference latency has become a vital basis for improving scheduling efficiency. However, edge devices exhibit highly heterogeneous due to the differences in hardware architectures, computing power, etc. Meanwhile, the diverse deep neural networks (DNNs) are continuing to iterate over time. The diversity of devices and DNNs introduces high computational costs for measurement methods, while invasive prediction methods face significant development efforts and application limitations. In this article, we propose and develop Ace-Sniper, a scheduling framework with DNN inference latency modeling on heterogeneous devices. First, to address the device heterogeneity, a unified hardware resource modeling (HRM) is designed by considering the platforms as black-box functions that output feature vectors. Second, neural network similarity (NNS) is introduced for feature extraction of diverse and frequently iterated DNNs. Finally, with the results of HRM and NNS as input, the performance characterization network is designed to predict the latencies of the given unseen DNNs on heterogeneous devices, which can be combined into most time-based scheduling algorithms. Experimental results show that the average relative error of DNN inference latency prediction is 11.11%, and the prediction accuracy reaches 93.2%. Compared with the nontime-aware scheduling methods, the average waiting time for tasks is reduced by 82.95%, and the platform throughput is improved by 63% on average. Weihong Liu, Jiawei Geng, Zongwei Zhu, Cheng Ji 0002, Changlong Li 0006, Zirui Lian, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | Arch2End: Two-Stage Unified System-Level Modeling for Heterogeneous Intelligent DevicesabstractThe surge in intelligent edge computing has propelled the adoption and expansion of the distributed embedded systems (DESs). Numerous scheduling strategies are introduced to improve the DES throughput, such as latency-aware and group-based hierarchical scheduling. Effective device modeling can help in modular and plug-in scheduler design. For uniformity in scheduling interfaces, an unified device performance modeling is adopted, typically involving the system-level modeling that incorporates both the hardware and software stacks, broadly divided into two categories. Fine-grained modeling methods based on the hardware architecture analysis become very difficult when dealing with a large number of heterogeneous devices, mainly because much architecture information is closed-source and costly to analyse. Coarse-grained methods are based on the limited architecture information or benchmark models, resulting in insufficient generalization in the complex inference performance of diverse deep neural networks (DNNs). Therefore, we introduce a two-stage system-level modeling method (Arch2End), combining limited architecture information with scalable benchmark models to achieve an unified performance representation. Stage one leverages public information to analyse architectures in an uniform abstraction and to design the benchmark models for exploring the device performance boundaries, ensuring uniformity. Stage two extracts critical device features from the end-to-end inference metrics of extensive simulation models, ensuring universality and enhancing characterization capacity. Compared to the state-of-the-art methods, Arch2End achieves the lowest DNN latency prediction relative errors in the NAS-Bench-201 (1.7%) and real-world DNNs (8.2%). It also showcases superior performance in intergroup balanced device grouping strategies. Weihong Liu, Zongwei Zhu, Boyu Li 0006, Yi Xiong 0003, Zirui Lian, Jiawei Geng, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Sniper: cloud-edge collaborative inference scheduling with neural network similarity modelingabstractThe cloud-edge collaborative inference demands scheduling the artificial intelligence (AI) tasks efficiently to the appropriate edge smart device. However, the continuously iterative deep neural networks (DNNs) and heterogeneous devices pose great challenges for inference tasks scheduling. In this paper, we propose a self-update cloud-edge collaborative inference scheduling system (Sniper) with time awareness. At first, considering that similar networks exhibit similar behaviors, we develop a non-invasive performance characterization network (PCN) based on neural network similarity (NNS) to accurately predict the inference time of DNNs. Moreover, PCN and time-based scheduling algorithms can be flexibly combined into the scheduling module of Sniper. Experimental results show that the average relative error of network inference time prediction is about 8.06%. Compared with the traditional method without time awareness, Sniper can reduce the waiting time by 52% on average while achieving a stable increase in throughput. Weihong Liu, Jiawei Geng, Zongwei Zhu, Zirui Lian |
DAC | 5 |
| 2022 | FedNorm: An Efficient Federated Learning Framework with Dual Heterogeneity Coexistence on Edge Intelligence SystemsabstractFederated learning (FL) is an emerging distributed learning paradigm, which aims to train machine learning models on geo-decentralized edge devices while keeping the training data stored locally. However, due to the scattered and diverse properties of edge devices, FL is often accompanied by typical heterogeneous features. One of the key challenges is statistical heterogeneity (aka non-independent identically distributed data, Non-IID), which leads to severe client-drift problem and unstable convergence. Moreover, the computational heterogeneity of devices can result in large computation time variation and thus exacerbate client-drift through inconsistent local training steps. The previous studies either ignore the client-drift problem or ignore the scatter in local gradient information, causing limited optimization effect. This paper proposes FedNorm framework to enable training Non-IID data on heterogeneous devices efficiently. First, a local model consistency update method is introduced to mitigate client-drift by allowing heterogeneous edge devices to implement different local training steps. Next, a federated gradient normalization method is introduced to reduce gradient scattering and achieves stable convergence of the model by balancing the gradient information of each edge device. We conducted extensive ablation experiments on different training tasks and training platforms with dual heterogeneity. The experimental results show that FedNorm achieves 1.52 × -3.52× speedup on convergence ratio and 7.38%-13.90% improvement in accuracy, compared to the state-of-the-art frameworks on CIFAR10. Zirui Lian, Weihong Liu, Zongwei Zhu, Xuehai Zhou |
ICCD | 1 |
| 2021 | HADFL: Heterogeneity-aware Decentralized Federated Learning FrameworkabstractFederated learning (FL) supports training models on geographically distributed devices. However, traditional FL systems adopt a centralized synchronous strategy, putting high communication pressure and model generalization challenge. Existing optimizations on FL either fail to speedup training on heterogeneous devices or suffer from poor communication efficiency. In this paper, we propose HADFL, a framework that supports decentralized asynchronous training on heterogeneous devices. The devices train model locally with heterogeneity-aware local steps using local data. In each aggregation cycle, they are selected based on probability to perform model synchronization and aggregation. Compared with the traditional FL system, HADFL can relieve the central server’s communication pressure, efficiently utilize heterogeneous computing power, and can achieve a maximum speedup of 3.15x than decentralized-FedAvg and 4.68x than Pytorch distributed training scheme, respectively, with almost no loss of convergence accuracy. Zirui Lian, Weihong Liu, Zongwei Zhu, Cheng Ji 0002 |
DAC | 2 |
| 2021 | AGQFL: Communication-efficient Federated Learning via Automatic Gradient Quantization in Edge Heterogeneous SystemsabstractWith the widespread use of artificial intelligent (AI) applications and dramatic growth in data volumes from edge devices, there are currently many works that place the training of AI models onto edge devices. The state-of-the-art edge training framework, federated learning (FL), requires to transfer of a large amount of data between edge devices and the central server, which causes heavy communication overhead. To alleviate the communication overhead, gradient compression techniques are widely used. However, the bandwidth of the edge devices is usually different, causing communication heterogeneity. Existing gradient compression techniques usually adopt a fixed compression rate and do not take the straggler problem caused by the communication heterogeneity into account. To address these issues, we propose AGQFL, an automatic gradient quantization method consisting of three modules: quantization indicator module, quantization strategy module and quantization optimizer module. The quantization indicator module automatically determines the adjustment direction of quantization precision by measuring the convergence ability of the current model. Following the indicator and the physical bandwidth of each node, the quantization strategy module adjusts the quantization precision at run-time. Furthermore, the quantization optimizer module designs a new optimizer to reduce the training bias and eliminate the instability during the training process. Experimental results show that AGQFL can greatly speed up the training process in edge AI systems while maintaining or even improving model accuracy. Zirui Lian, Yanru Zuo, Weihong Liu, Zongwei Zhu |
ICCD | 1 |