Zhao Yang 0005

dblp:21/2326-5 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0002-9525-2096ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 6 first-author · 12 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Dynamic Rank-Aware Aggregation with Graph Contrastive Learning for Federated Foundation Model Fine-Tuning
Zhao Yang 0005, Xunyun Qiu
DATE1
2026 Automated federated aggregation for dynamic systems and data in mobile edge computing
Zhao Yang 0005, Xuanyun Qiu, Weiyi Hu, Qingshuang Sun
Future Gener. Comput. Syst.1
2024 Resource-Efficient Heterogenous Federated Continual Learning on Edge
abstract
Federated learning (FL) has been widely deployed on edge devices. In practical, the data collected by edge devices exhibits temporal variations. This leads to catastrophic forgetting issue. Continual learning methods can be used to address this problem. However, when deploying these methods in FL on edge devices, it is challenging to adapt to the limited resources and heterogeneous data of the deployed devices, which reduces the efficiency and effectiveness of federated continual learning (FCL). Therefore, this article proposes a resource-efficient heterogeneous FCL framework. This framework divides the global model into an adaptation part for new knowledge and a preservation part for old knowledge. The preservation part is used to address the catastrophic forgetting problem. Only the adaptation part is trained when learning new knowledge on a new task, reducing resource consumption. Additionally, the framework mitigates the impact of heterogeneous data through an aggregation method based on feature representation. Experimental results show that our method performs well in mitigating catastrophic forgetting in a resource-efficient manner.
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Haoyang Wang 0014, Meng Zhang 0047
DATE1
2024 OFT: An accelerator with eager gradient prediction for attention training
abstract
With the tremendous success of Transformer, resource-constrained edge devices are increasingly becoming the deployment target for attention-based models. On-device attention training can address the accuracy decline caused by static models' inability to adapt to dynamic environments efficiently while protecting data privacy. However, edge devices cannot meet the resource demands caused by batched gradient backpropagation and weight updates. Due to the inherent redundancy in human language, sparsification is the primary choice to alleviate the contradiction. Current sparse training methods consume high runtime costs in both forward propagation (FP) and backward propagation (BP) computations to handle irregular sparse patterns, which makes it inefficient to convert potential speedups into actual performance improvements and energy savings.
Shengbing Zhang, Zhao Yang 0005, Meng Zhang 0047
ICCAD4
2024 Efficient knowledge management for heterogeneous federated continual learning on resource-constrained edge devices
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Haoyang Wang 0014, Meng Zhang 0047
Future Gener. Comput. Syst.1
2024 NDPGNN: A Near-Data Processing Architecture for GNN Training and Inference Acceleration
abstract
Graph neural networks (GNNs) require a large number of fine-grained memory accesses, which results in inefficient use of bandwidth resources. In this article, we introduce a near-data processing architecture tailored for GNN acceleration, named NDPGNN. NDPGNN provides different operating modes to meet the acceleration needs of various GNN frameworks while ensuring the configurability and scalability of the system. NDPGNN takes advantage of data locality characteristics to repeatedly distribute and utilize data, thereby reducing memory access requirements, and further improving memory access efficiency by combining a subgraph sparse node scheduling strategy with intermediate result reuse. We use data packaging to provide a higher effective data ratio for long-distance data transmission, thereby improving the utilization of the system’s limited bandwidth resources. Compared with the previous method, NDPGNN brings 5.68 times improvement in system performance while reducing energy consumption overhead by 8.49 times.
Haoyang Wang 0014, Shengbing Zhang, Xiaoya Fan, Zhao Yang 0005, Meng Zhang 0047
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 Equalized Aggregation for Heterogeneous Federated Mobile Edge Learning
abstract
Federated Learning (FL) is widely used in mobile edge applications. However, the heterogeneity issues of mobile edge devices pose significant challenges to the generalization of the global model in FL. In this paper, we propose LegoFL to simultaneously solve multiple heterogeneity issues in response to mobile edge computing characteristics. LegoFL identifies two types of heterogeneous behaviors in FL, namely heterogeneous parameter training and communication behaviors, to address multiple heterogeneity issues. These two types of heterogeneous behaviors result in feature and feature representation range mismatches between local communication parameters. To reduce these mismatches and improve the generalization of the global model, LegoFL dynamically distinguishes the parameter feature representation of different nodes using the global model's common feature as guidance. Then, under the connection states and system communication constraints, LegoFL dynamically selects contribution parameters on each device that can guarantee the generalization and performance of the global model for communication. Finally, to avoid the overfitting problem of the global model, heterogeneous local models are aggregated at the central server with matched feature representations. Extensive experiments on various datasets show that LegoFL achieves competitive performance. The accuracy and communication efficiency are improved by up to 12.86$\%$and 4.09× compared to state-of-the-art approaches.
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Jiaying Yang, Meng Zhang 0047
IEEE Trans. Mob. Comput.1
2023 SaGNN: a Sample-based GNN Training and Inference Hardware Accelerator
abstract
Graph neural networks (GNNs) operations contain a large number of irregular data operations and sparse matrix multiplications, resulting in the under-utilization of computing resources. The problem becomes even more complex and challenging when it comes to large graph training. Scaling GNN training is an effective solution. However, the current GNN operation accelerators do not support the mini-batch structure. We analyze the GNN operational characteristics from multiple aspects and take both the acceleration requirements in the GNN training and inference process into account, and then propose the SaGNN system structure. SaGNN offers multiple working modes to provide acceleration solutions for different GNN frameworks while ensuring system configurability and scalability. Compared to related works, SaGNN brings 5.0x improvement in system performance.
Haoyang Wang 0014, Shengbing Zhang, Kaijie Feng, Zhao Yang 0005
ISCAS5
2023 Joint heterogeneity-aware personalized federated search for energy efficient battery-powered edge computing
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Jiaying Yang, Meng Zhang 0047
Future Gener. Comput. Syst.1
2023 Energy-efficient Personalized Federated Search with Graph for Edge Computing
abstract
Federated Learning (FL) is a popular method for privacy-preserving machine learning on edge devices. However, the heterogeneity of edge devices, including differences in system architecture, data, and co-running applications, can significantly impact the energy efficiency of FL. To address these issues, we propose an energy-efficient personalized federated search framework. This framework has three key components. Firstly, we search for partial models with high inference efficiency to reduce training energy consumption and the occurrence of stragglers in each round. Secondly, we build lightweight search controllers that control the model sampling and respond to runtime variances, mitigating new straggler issues caused by co-running applications. Finally, we design an adaptive search update strategy based on graph aggregation to improve personalized training convergence. Our framework reduces the energy consumption of the training process by lowering the training overhead of each round and speeding up the training convergence rate. Experimental results show that our approach achieves up to 5.02% accuracy and 3.45× energy efficiency improvements.
Zhao Yang 0005, Qingshuang Sun
ACM Trans. Embed. Comput. Syst.1
2022 A dynamic global backbone updating for communication-efficient personalised federated learning
abstract
Federated learning (FL) is an emerging distributed machine learning technique. However, when dealing with heterogeneous data, a shared global model cannot generalise all devices' local data. Furthermore, the FL training process necessitates frequent parameter communication, which interferes with the limited bandwidth and unstable connections of participating devices. These two issues have a significant impact on FL's effectiveness and efficiency. In this paper, an enhanced communication-efficient personalised FL technique, FedGB, is proposed. Different from existing approaches, FedGB believes that only interacting common information from training results on different devices can improve local personalised training results more effectively. FedGB dynamically selects the backbone structures in the local models to represent the dynamically determined backbone information (common features) in the global model for aggregation. Only interacting common features between different nodes reduce the impact of heterogeneous data to a certain extent. The dynamic adaptive sub-model selection avoids the impact of manually setting the scale of sub-model. FedGB can thus reduce communication overheads while maintaining inference accuracy. The results obtained in a variety of experimental settings show that FedGB can effectively improve communication efficiency and inference accuracy.
Zhao Yang 0005, Qingshuang Sun
Connect. Sci.1
2022 DCNN search and accelerator co-design: Improve the adaptability between NAS frameworks and embedded platforms
Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047
Integr.4
2022 Memory-Computing Decoupling: A DNN Multitasking Accelerator With Adaptive Data Arrangement
abstract
Multiple deep neural networks (DNNs) are increasingly used in real-world intelligent applications, such as intelligent robotics and autonomous vehicles to collectively complete complicated tasks running on edge devices. Because each layer of the subtasks prefers a distinct dataflow due to the heterogeneity in shape and scale of the network layers, a variable dataflow approach on the DNN accelerators is urgently required. On DNN accelerators that enable multiple dataflows, however, we detect a dimension mismatch between parallel processing under the dataflow approach and linear data memory arrangement. When multiple DNN tasks share partial features or weights, the issue is further exacerbated. During processing, this mismatch causes a sluggish data supply from both off-chip and on-chip memory. Consequently, the overall throughput, performance, and energy efficiency suffer since DNN models are sensitive to data density. In this work, we reveal the mechanism behind this data dimension mismatch and present a series of metrics that quantify the influence on system performance. On this foundation, we offer a framework that tracks the data tensor dimension conversion and employs a flexible data arrangement over multi-DNN computation to adapt to dataflow variability. An accelerator architecture named data arrangement multi-DNN accelerator (DARMA) that features a data arrangement and distribution circuit and hierarchical memory for data dimension conversion is also presented. Since the mismatch is mitigated, the suggested accelerator outperforms current accelerators in terms of bandwidth and processing unit utilization. Through tests on VR/AR, MLperf, and other multitask applications, the evaluation results show that the proposed architecture provides both energy-efficiency and throughput improvements.
Chuxi Li, Xiaoya Fan, Xiaoti Wu, Zhao Yang 0005, Meng Zhang 0047, Shengbing Zhang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 Hardware-Aware NAS Framework with Layer Adaptive Scheduling on Embedded System
abstract
Neural Architecture Search (NAS) has been proven to be an effective solution for building Deep Convolutional Neural Network (DCNN) models automatically. Subsequently, several hardware-aware NAS frameworks incorporate hardware latency into the search objectives to avoid the potential risk that the searched network cannot be deployed on target platforms. However, the mismatch between NAS and hardware persists due to the absent of rethinking the applicability of the searched network layer characteristics and hardware mapping. A convolution neural network layer can be executed on various dataflows of hardware with different performance, with which the characteristics of on-chip data using varies to fit the parallel structure. This mismatch also results in significant performance degradation for some maladaptive layers obtained from NAS, which might achieved a much better latency when the adopted dataflow changes. To address the issue that the network latency is insufficient to evaluate the deployment efficiency, this paper proposes a novel hardware-aware NAS framework in consideration of the adaptability between layers and dataflow patterns. Beside, we develop an optimized layer adaptive data scheduling strategy as well as a coarse-grained reconfigurable computing architecture so as to deploy the searched networks with high power-efficiency by selecting the most appropriate dataflow pattern layer-by-layer under limited resources. Evaluation results show that the proposed NAS framework can search DCNNs with the similar accuracy to the state-of-the-art ones as well as the low inference latency, and the proposed architecture provides both power-efficiency improvement and energy consumption saving.
Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047
ASP-DAC4