EDBT 2026 Demo / reviewers in the wild / expert
Qianyue Cao
dblp:377/1446
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0001-2861-8211ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PPFL: A Parameter Behavior-Driven Plug-in Personalization Engine for Federated LearningabstractPersonalized Federated Learning (PFL) customizes models for each client to mitigate challenges from non-IID data, wherein a dominant strategy is model decoupling that partitions models into shared and personalized parts based on architectural priors (e.g., backbone vs. head). However, we reveal a critical flaw in this strategy: it induces "intrinsic drift," a performance degradation often more severe than the well-known client drift, which limits final accuracy. We trace this drift to a steep cliff of high loss emerging from the naive stitching of shared and personalized parts. To address this, we shift from architectural partitioning to a parameter behavior-driven paradigm. We introduce PPFL, an approach that employs a novel soft-fusion strategy guided by parameter-wise behavioral perception. PPFL dynamically infers each parameter's functional role—whether it behaves more like a 'personalist' or a 'generalist' in the current context—by synthesizing its multifaceted behavior observed during local training. Extensive experiments on image, text, and multimodal classification benchmarks show that PPFL outperforms eight state-of-the-art baselines by up to 5.3%. Moreover, it can function as a plug-in module, boosting the accuracy of vanilla FedAvg with a 16.82% absolute gain. Qianyue Cao, Zongwei Zhu, Zirui Lian, Rui Zhang 0040, Boyu Li 0006, Yi Xiong 0003, Xuehai Zhou |
AAAI | 1 |
| 2026 | FedGAMA: Federated Learning on Heterogeneous and Long-Tailed Data via Group-Wise Asymmetric Masked Aggregation
Chenyue Xu, Zongwei Zhu, Qianyue Cao, Rui Zhang 0040, Xuehai Zhou |
KSEM (1) | 3 |
| 2026 | CSCL: Bridging the plasticity-stability gap in continuous supervised contrastive learning
Yi Xiong 0003, Liqi Xiang, Qianyue Cao, Zongwei Zhu, Zirui Lian, Xuehai Zhou |
Neural Networks | 3 |
| 2026 | AsyncGrid: An Intralayer and Interlayer Asynchronous Hybrid Parallelism System for Responsive Edge LLM InferenceabstractEdge deployment of large language models (LLMs) is increasingly attractive due to its advantages in privacy, customization, and availability. However, edge environments face significant challenges in reducing Time-to-First-Token (TTFT). TTFT consists of (1) queuing delay and (2) prefill latency, both of which are exacerbated by edge‑resource constraints: the substantial computational demands of LLM inference grow superlinearly with prompt length, causing high prefill latency; and limited edge resources restrict prefill throughput, preventing the timely handling of incoming requests, thereby exacerbating queuing delays. Model parallelism is a commonly used solution in cloud-based systems, but directly applying it to edge environments proves ineffective. Intra-layer parallelism (e.g., tensor/sequence parallelism) can reduce prefill latency but suffers from frequent global synchronization, which bottlenecks prefill throughput due to edge-limited interconnection bandwidth. Inter-layer parallelism (e.g., pipeline parallelism) improves prefill throughput via fully asynchronous execution but retains high prefill latency due to stage-wise serialized computation. To address this dilemma, this paper leverages the properties of the causal attention mechanism in LLMs and proposes Intra-layer Asynchronous Parallelism (IAP), which performs intra-layer parallel computations to reduce prefill latency while avoiding global synchronization to mitigate prefill throughput bottlenecks. Moreover, considering communication sensitivity in intra-layer parallelism, this paper integrate IAP with inter-layer asynchronous parallelism into a unified plan space. This hybrid parallelism adapts to diverse hardware and request loads, enabling more effective TTFT optimization. To enable the end-to-end implementation of this hybrid parallelism, this paper propose AsyncGrid, an LLM inference system tailored for responsive edge LLM inference. AsyncGrid (1) models runtime overheads through a performance profiler, (2) employs an integer programming (IP) formulation to optimize execution plan, with the objective of minimizing latency while meeting throughput requirements, and (3) implements fine-grained communication optimization during runtime. A comprehensive evaluation on an edge testbed demonstrates AsyncGrid’s significant advantages over existing methods, achieving substantial improvements in both homogeneous and heterogeneous settings. Yi Xiong 0003, Rui Zhang 0040, Yulong Zu, Weihong Liu, Zongwei Zhu, Jiawei Geng, Boyu Li 0006, Qianyue Cao, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2025 | Magnifier: A Chiplet Feature-Aware Test Case Generation Method for Deep Learning AcceleratorsabstractThe development of deep learning has led to increasing demands for computation and memory, making multi-chiplet accelerators a powerful solution. Multi-chiplet accelerators require more precise consideration of hardware configurations and mapping schemes in terms of computation, memory, and communication patterns compared to monolithic designs, in order to avoid underutilization of performance. However, there is currently a lack of performance testing methods specifically tailored for multi-chiplet accelerators. Existing testing methods primarily focus on correctness testing and do not address potential performance issues from a hardware perspective. To address these issues, this paper proposes Magnifier: a test case generation method for performance testing of multi-chiplet accelerators. Firstly, we analyze typical multi-chiplet accelerator prototype from the perspectives of computation, memory, and communication patterns, and summarize a chiplet feature-aware operator task set. Next, we define the test evaluation metric IPPstd and use a candidate operator set to construct a sampling space for model-level test cases. Finally, we build a GAN to learn the distribution of high-diversity test cases, enabling the rapid generation of high-quality test cases. We validate the proposed method on both simulated and real multi-chiplet accelerators. Experiments show that Magnifier can improve the metric of test cases by up to 3.42 times and significantly reduce generation time, providing valuable insights for optimizing the hardware and software of multi-chiplet accelerators. Boyu Li 0006, Zongwei Zhu, Weihong Liu, Qianyue Cao, Changlong Li 0006, Cheng Ji 0002, Xi Li 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | HaloFL: Efficient Heterogeneity-Aware Federated Learning Through Optimal Submodel Extraction and Dynamic Sparse AdjustmentabstractFederated learning (FL) is an advanced framework that enables collaborative training of machine learning models across edge devices. An effective strategy to enhance training efficiency is to allocate the optimal submodel based on each device’s resource capabilities. However, system heterogeneity significantly increases the difficulty of allocating submodel parameter budgets appropriately for each device, leading to the straggler problem. Meanwhile, data heterogeneity complicates the selection of the optimal submodel structure for specific devices, thereby impacting training performance. Furthermore, the dynamic nature of edge environments, such as fluctuations in network communication and computational resources, exacerbates these challenges, making it even more difficult to precisely extract appropriately sized and structured submodels from the global model. To address the challenges in heterogeneous training environments, we propose an efficient FL framework, namely, HaloFL. The framework dynamically adjusts the structure and parameter budget of submodels during training by evaluating three dimensions: 1) model-wise performance; 2) layer-wise performance; and 3) unit-wise performance. First, we design a data-aware model unit importance evaluation method to determine the optimal submodel structure for different data distributions. Next, using this evaluation method, we analyze the importance of model layers and reallocate parameters from noncritical layers to critical layers within a fixed parameter budget, further optimizing the submodel structure. Finally, we introduce a resource-aware dual-UCB multiarmed bandit agent, which dynamically adjusts the total parameter budget of submodels according to changes in the training environment, allowing the framework to better adapt to the performance differences of heterogeneous devices. Experimental results demonstrate that HaloFL exhibits outstanding efficiency in various dynamic and heterogeneous scenarios, achieving up to a 14.80% improvement in accuracy and a$3.06\times $speedup compared to existing FL frameworks. Zirui Lian, Qianyue Cao, Zongwei Zhu, Cheng Ji 0002, Changlong Li 0006, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | FedStar: Efficient Federated Learning on Heterogeneous Communication NetworksabstractThe proliferation of multi-media applications and increased computing power of mobile devices have led to the development of personalized artificial intelligent (AI) applications that utilize the massive user-information residing on them. However, the traditional centralized training paradigm is not applicable in this scenario due to potential privacy risks and high communication overhead. Federated learning (FL) provides an option to these applications. Nevertheless, the heterogeneity of computing and communication latency among devices have posed great challenges to building efficient learning frameworks. Existing optimizations on FL either fail to speed up training on heterogeneous devices or suffer from poor communication efficiency. In this paper, we propose FedStar, an efficient FL framework that supports decentralized asynchronous training on heterogeneous communication networks. Considering the heterogeneous computing power in the network, FedStar supports running heterogeneity-aware local steps on each device. What’s more, considering the heterogeneous communication latency and possibly unreachable communication path between some devices, FedStar generates a decentralized communication topology that can achieve maximal training throughput. Finally, it adopts weighted aggregation to guarantee high convergence accuracy of global model. Theoretical analysis results show the convergence behaviour of FedStar under non-convex settings. Experimental results show that FedStar can achieve a speedup of 4.81× than the state-of-the-art FL schemes with high convergence accuracy. Qianyue Cao, Yongchun Zheng, Zongwei Zhu, Cheng Ji 0002, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | NebulaFL: Self-Organizing Efficient Multilayer Federated Learning Framework With Adaptive Load Tuning in Heterogeneous Edge SystemsabstractAs a promising edge intelligence technology, federated learning (FL) enables Internet of Things (IoT) devices to train the models collaboratively while ensuring the data privacy and security. Recently, hierarchical FL (HFL) has been designed to promote distributed training in the intricate hierarchical structure of IoT. However, the coarse-grained hierarchical schemes usually fail to thoroughly adapt to the hierarchical environment, leading to high training latency. Meanwhile, highly heterogeneous communication and computation delays due to the device diversity (the system heterogeneity) and decentralized data distribution due to the decentralized device distribution (the data heterogeneity) exacerbate the above challenges. This article proposes NebulaFL, a dual heterogeneity-aware multilayer FL framework, to support efficient distributed training in IoT scenarios. NebulaFL proposes an innovative multilayer architecture organization scheme to adapt the complex hierarchical heterogeneous scenarios. Specifically, through a finer-grained division of the HFL hierarchy, hybrid synchronous-asynchronous training is implemented at both the global system and local device-layer levels. More importantly, to adaptively build a heterogeneity-aware hierarchical training architecture, NebulaFL considers the effect of dual heterogeneity in the architectural organization scheme to determine the optimal location of devices in a multilayer environment. To further improve the training efficiency during the training process, NebulaFL employs an augmented multiarmed bandit technique based on the reinforcement learning to adjust the device-layer training load by evaluating the dynamic training utility and convergence uncertainty feedback. Experiments demonstrate that NebulaFL achieves up to a$15.68\times $speed-up ratio and a 23.94% increase in the training accuracy compared to the latest or classic approaches. Zirui Lian, Qianyue Cao, Weihong Liu, Zongwei Zhu, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |