Zhaoyang Du

dblp:231/8624 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 3 first-author · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HDFL: A Hierarchical Decentralized Federated Learning Framework for Dynamic and Heterogeneous IoV Environments
abstract
Traditional federated learning (FL) approaches face significant challenges when applied to dynamic and heterogeneous Internet of Vehicles (IoV) environments, which are characterized by frequent node mobility, unstable communication links, and highly non-independent and identically distributed (Non-IID) data. In particular, decentralized network topologies exacerbate the difficulty of maintaining model consistency, thereby impairing overall learning performance. To address these challenges, we propose a new hierarchical decentralized federated learning (HDFL) framework. This framework combines the advantages of centralization and decentralization, builds a three-layer collaborative structure, and improves communication flexibility through an asynchronous model exchange mechanism between the edge and the client. Simultaneously, HDFL introduces a local fine-tuning strategy based on knowledge distillation to enhance the generalization ability and stability of the model. Experimental results using an urban traffic simulation platform show that HDFL consistently outperforms representative decentralized FL methods in terms of the achieved accuracy and convergence speed under heterogeneous IoV environments.
Celimuge Wu, Yangfei Lin, Zhaoyang Du, Jianhang Tang, Soufiene Djahel
INFOCOM4
2026 Real-Time Network Behavior Modeling for Collaborative Operations of Low-Altitude UAV Swarms
Yalong Li 0001, Celimuge Wu, Zhaoyang Du, Yangfei Lin, Soufiene Djahel, Kai Liu 0001
IWCMC3
2026 BDGraS: Bandwidth-adaptive dual-relation gravity model for efficient cooperative vehicle selection in autonomous driving
Yalong Li 0001, Yangfei Lin, Zhaoyang Du, Kai Liu 0001, Wugedele Bao, Celimuge Wu
Comput. Networks4
2025 CTXNL: A Software-Hardware Co-designed Solution for Efficient CXL-Based Transaction Processing
abstract
Transaction processing systems are the crux for modern data-center applications, yet current multi-node systems are slow due to network overheads. This paper advocates for Compute Express Link (CXL) as a network alternative, which enables low-latency and cache-coherent shared memory accesses. However, directly adopting standard CXL primitives leads to performance degradation due to the high cost of maintaining cross-node cache coherence. To address the CXL challenges, this paper introduces CTXNL, a software-hardware co-designed system that implements a novel hybrid coherence primitive tailored to the loosely coherent nature of transactional data. The core innovation of CTXNL is empowering transaction system developers with the ability to selectively achieve data coherence. Our evaluations on OLTP workloads demonstrate that CTXNL enhances performance, outperforming current network-based systems and achieves up to 2.08x greater throughput than vanilla CXL memory sharing architectures across universal transaction processing policies.
Cong Li 0008, Yijin Guan, Dimin Niu, Tianchan Guan, Zhaoyang Du, Xingda Wei, Guangyu Sun 0003
ASPLOS (2)7
2025 A Knowledge Distillation-Based Framework for Enhanced Long and Short Term Road Traffic Prediction
abstract
Accurate traffic congestion prediction is essential for optimizing urban traffic management and mitigating congestion and its consequences. However, conventional prediction models often struggle to simultaneously capture long-term periodic patterns and short-term fluctuations, leading to low prediction accuracy and computational inefficiencies. To overcome this limitation, we propose a knowledge distillation-based framework for enhanced long and short term road traffic prediction. The framework employs a teacher-student architecture, where the teacher model utilizes long-term historical data and a dynamic adjacency matrix to extract periodic traffic patterns, while the student model captures short-term variations and integrates distilled long-term knowledge to enhance responsiveness to sudden congestion changes. To resolve the dimensional mismatch between long-term and short-term feature representations, we introduce a feature alignment mechanism that reduces the dimensionality of high-dimensional intermediate outputs from the teacher model. Experimental evaluations demonstrate that our approach significantly outperforms baseline models, such as Graph Convolutional Gated Recurrent Units, Spatio- Temporal Graph Convolutional Networks and Long Short-Term Memory, in terms of Mean Squared Error, Mean Absolute Error, and Root Mean Squared Error. Moreover, the proposed framework maintains high prediction accuracy even in scenarios with severe traffic fluctuations, offering an efficient and robust solution for traffic congestion forecasting.
Junting Gao, Yangfei Lin, Zhaoyang Du, Wugedele Bao, Soufiene Djahel
VTC2025-Spring3
2025 MemTunnel: A CXL-Based Rack-Scale Host Memory Pooling Architecture for Cloud Service
abstract
Memory underutilization poses a significant challenge in cloud services, leading to performance inefficiencies and resource wastage. The tightly coupled computing and memory resources in cloud servers are identified as the root cause of this problem. To address this issue, memory pooling has been the subject of extensive research for decades, providing centralized or distributed shared memory pools as flexible memory resources for various applications running on different servers. However, existing memory disaggregation solutions sacrifice memory resources, add extra hardware (such as memory boxes/blades/drives), and degrade memory performance to achieve flexibility. To overcome these limitations, this paper proposes MemTunnel, a rack-scale host memory pooling architecture that provides a low-cost memory pooling solution based on Compute Express Link (CXL). MemTunnel is the first hardware and software architecture to offer symmetric, memory-semantic memory pooling over CXL, with an FPGA-based platform to demonstrate its feasibility in a real implementation. MemTunnel is orthogonal to the existing CXL-based memory pool and provides an additional layer of abstraction for memory disaggregation. Evaluation results show that MemTunnel achieves comparable performance to the existing CXL-based memory pool for a single machine and provides better rack-scale performance with minor hardware overheads.
Tianchan Guan, Yijin Guan, Zhaoyang Du, Jiacheng Ma 0001, Boyu Tian, Teng Ma 0006, Zheng Liu 0022, Yuan Xie 0001, Mingyu Gao 0001, Guangyu Sun 0003, Hongzhong Zheng, Dimin Niu
IEEE Trans. Parallel Distributed Syst.3
2024 Fuzzy Logic-based Enhanced Edge Server Selection for Hierarchical Federated Learning
abstract
In the rapidly evolving landscape of federated learning (FL), hierarchical architectures are pivotal for improving computational efficiency and safeguarding data privacy. A key challenge in this research area is the optimal selection of edge servers, crucial for executing distributed learning tasks across multiple clients and servers efficiently. Traditional selection methods falter due to their inability to dynamically handle the uncertainties in network conditions and server capabilities. To addressing this weakness, we propose a fuzzy logic-based approach that optimizes edge server selection in a novel smart way, thus enhancing resource allocation by efficiently handling the unpredictable nature of network environments and servers performance. This method is integrated with a previously developed scheme for selecting an optimal subset of clients, thereby establishing a comprehensive framework that significantly boosts the performance and reliability of FL networks. The performance of our approach is validated through real-world experiments and the results demonstrate its superiority over existing methods in terms of accuracy and processing time.
Zhaoyang Du, Celimuge Wu, Yangfei Lin, Soufiene Djahel, Peter Han Joo Chong
GLOBECOM1
2024 Hier-FedMeta: A Hierarchical Federated Meta-Learning Framework for Personalized and Efficient IoV Systems
abstract
The Internet of Vehicles (IoV) enhances smart city functionalities by interconnecting diverse components, yet it introduces significant challenges in terms of user privacy, communication efficiency, and energy consumption. Traditional federated learning frameworks, while adept at addressing these concerns, fall short in personalization due to heterogeneous data distributions among clients. To overcome this, we introduce Hier-FedMeta, a novel framework that combines hierarchical federated learning with meta-learning to provide tailored and efficient solutions. Our comparative analyses with four estab-lished methods show Hier-FedMeta's superior generalization capabilities and adaptability, achieving enhanced performance with minimal computational overhead after just one update step. Furthermore, our in-depth analysis of aggregation parameters offers valuable insights for the optimization of hierarchical federated meta-learning architectures, representing a significant step forward in personalized learning for IoV in smart cities.
Celimuge Wu, Zhaoyang Du, Yangfei Lin, Soufiene Djahel
VTC Spring3
2023 Communication-Efficient Federated Learning for UAV Networks with Knowledge Distillation and Transfer Learning
abstract
Federated learning (FL) in unmanned aerial ve-hicles (UAVs) networks demands considerable communication resources to transfer model data between the central server and UAVs (FL clients). However, different UAVs may have different communication capabilities due to the UAV's maneuverability and heterogeneity, where limited communication resource could be bottle neck for FL performance. In this paper, we first introduce a knowledge distillation based approach that places two different models with different sizes, namely the teacher model and student model, for FL client to make a trade-off between the FL performance and communication cost. Then, we propose a novel model switching method to switch between the teacher model and student model to adapt to the dynamic feature of UAV networks. Specifically, considering available communication bitrate and learning accuracy, we design a threshold-based model switching algorithm (TBMSA) and determine the threshold based on the k-means method (DTBKM) to accurately and quickly determine the switching point. In addition, for the knowledge transfer between models, we design a knowledge inheritance based on a transfer learning (KIBTL) algorithm, which transfers knowledge from one model to another. Experiments show that the proposed model switching algorithm achieves significant performance improvements as compared to existing baselines.
Yalong Li 0001, Celimuge Wu, Zhaoyang Du, Tsutomu Yoshinaga
GLOBECOM3
2023 Semantic Communication for Efficient Image Transmission Tasks based on Masked Autoencoders
abstract
Semantic communication, a promising candidate for 6G technology, has become a research hot spot. However, existing studies tend to focus more on image reconstruction rather than accurately transmitting semantic information at the pixel level. This paper introduces a novel approach using codec-based Masked AutoEncoders (MAE) for efficient image transmission. The proposed system compresses local information into low-dimensional latent vectors, improving system efficiency. We also design a selective module for enhanced image reconstruction and implement Noise Adversarial Training (NAT) to increase the system’s resilience to channel noise. Experimental results show that our method effectively improves downstream tasks while preserving image quality.
Celimuge Wu, Yangfei Lin, Jingjing Bao, Zhaoyang Du, Xianfu Chen, Yusheng Ji
VTC Fall5
2022 Predicting the Output Structure of Sparse Matrix Multiplication with Sampled Compression Ratio
abstract
Sparse general matrix multiplication (SpGEMM) is a fundamental building block in numerous scientific applications. One critical task of SpGEMM is to compute or predict the structure of the output matrix (i.e., the number of nonzero elements per output row) for efficient memory allocation and load balance, which impact the overall performance of SpGEMM. Existing work either precisely calculates the output structure or adopts upper-bound or sampling-based methods to predict the output structure. However, these methods either take much execution time or are not accurate enough. In this paper, we propose a novel sampling-based method with better accuracy and low costs compared to the existing sampling-based method. The proposed method first predicts the compression ratio of SpGEMM by leveraging the number of intermediate products (denoted as FLOP) and the number of nonzero elements (denoted as NNZ) of the same sampled result matrix. And then, the predicted output structure is obtained by dividing the FLOP per output row by the predicted compression ratio. We also propose a reference design of the existing sampling-based method with optimized computing overheads to demonstrate the better accuracy of the proposed method. We construct 623 test cases with various matrix dimensions and sparse structures to evaluate the prediction accuracy. Experimental results show that the absolute relative errors of the proposed method and the reference design are 1.30% and 7.93%, respectively, on average, and 25% and 158%, respectively, in the worst case.
Zhaoyang Du, Yijin Guan, Tianchan Guan, Dimin Niu, Nianxiong Tan, Xiaopeng Yu 0002, Hongzhong Zheng, Jian-Yi Meng, Xiaolang Yan, Yuan Xie 0001
ICPADS1
2022 Hyperscale FPGA-as-a-service architecture for large-scale distributed graph neural network
abstract
Graph neural network (GNN) is a promising emerging application for link prediction, recommendation, etc. Existing hardware innovation is limited to single-machine GNN (SM-GNN), however, the enterprises usually adopt huge graph with large-scale distributed GNN (LSD-GNN) that has to be carried out with distributed in-memory storage. The LSD-GNN is very different from SM-GNN in terms of system architecture demand, workflow and operators, and hence characterizations.
Shuangchen Li, Dimin Niu, Yuhao Wang 0002, Zhe Zhang 0006, Tianchan Guan, Yijin Guan, Linyong Huang, Zhaoyang Du, Yuanwei Fang, Hongzhong Zheng, Yuan Xie 0001
ISCA10
2022 EPQuant: A Graph Neural Network compression approach based on product quantization
Linyong Huang, Zhe Zhang 0006, Zhaoyang Du, Shuangchen Li, Hongzhong Zheng, Yuan Xie 0001, Nianxiong Tan
Neurocomputing3
2020 UAV-empowered Protocol for Information Sharing in VDTN
abstract
The Delay Tolerant Network (DTN) is a network architecture that plays an important role in intermittently connected networks. DTNs can support transmissions between clients even there are no end-to-end connections by using “store-carry-forward” mechanism, where some nodes called ferry nodes could be employed into DTN to enhance the network performance. In this paper, we use Unmanned Aerial Vehicles (UAVs) to act as ferry nodes, and a probabilistic routing protocol based on the encounter connection time between nodes is proposed. The proposed protocol not only considers the efficiency of message transmission but also the reliability between connected nodes. The proposed protocol and some existing protocols are compared and analyzed using ONE simulator. The simulation results show that the proposed protocol increases the delivery probability and reduces the average latency.
Zhaoyang Du, Celimuge Wu, Tsutomu Yoshinaga
MSN1
2020 A VDTN scheme with enhanced buffer management
Zhaoyang Du, Celimuge Wu, Xianfu Chen, Xiaoyan Wang 0003, Tsutomu Yoshinaga, Yusheng Ji
Wirel. Networks1