EDBT 2026 Demo / reviewers in the wild / expert
Meng Zhang 0047
dblp:04/6901-47
· DBLP profile ↗
12ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0003-0637-249XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 10 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TPQA: Efficient attention architecture with task-aware pattern-guided quantization
Shengbing Zhang, Yichao Yuan, Yawen Zhao 0010, Meng Zhang 0047 |
Future Gener. Comput. Syst. | 7 |
| 2026 | PoDe-SORT: Robust multi-object tracking by modeling bounding box deformation
Shuaipeng Duan, Shengbing Zhang, Meng Zhang 0047 |
Knowl. Based Syst. | 3 |
| 2025 | STAPC: A Sparse Training Accelerator for Efficient On-Device DNN Learning via Position ConstraintsabstractLeveraging sparsity to eliminate redundant computation and storage overhead is essential for enhancing the efficiency of on-device deep neural network (DNN) learning. However, due to a lack of assumptions about non-zero positions, existing sparse methods incur high costs for zero-position identification and allocation, making it difficult to achieve ideal acceleration. This paper demonstrates that knowing non-zero position constraints in advance during training can bypass these sparse processing overheads. We explore non-zero position constraints among operands for three typical activation functions in edge scenarios and propose: (1) a hardware-friendly sparse training algorithm to skip redundant gradient computations, enhancing training efficiency; and (2) a high-efficiency sparse training accelerator, STAPC, that estimates non-zero gradient positions, allowing costly sparse processing to be masked in parallel to reduce energy consumption. Compared to the baseline and other sparse training methods, the proposed method achieves energy efficiency gains of 2.2x, 1.38x, and 1.46x, respectively. Shengbing Zhang, Meng Zhang 0047 |
ISCAS | 3 |
| 2024 | Resource-Efficient Heterogenous Federated Continual Learning on EdgeabstractFederated learning (FL) has been widely deployed on edge devices. In practical, the data collected by edge devices exhibits temporal variations. This leads to catastrophic forgetting issue. Continual learning methods can be used to address this problem. However, when deploying these methods in FL on edge devices, it is challenging to adapt to the limited resources and heterogeneous data of the deployed devices, which reduces the efficiency and effectiveness of federated continual learning (FCL). Therefore, this article proposes a resource-efficient heterogeneous FCL framework. This framework divides the global model into an adaptation part for new knowledge and a preservation part for old knowledge. The preservation part is used to address the catastrophic forgetting problem. Only the adaptation part is trained when learning new knowledge on a new task, reducing resource consumption. Additionally, the framework mitigates the impact of heterogeneous data through an aggregation method based on feature representation. Experimental results show that our method performs well in mitigating catastrophic forgetting in a resource-efficient manner. Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Haoyang Wang 0014, Meng Zhang 0047 |
DATE | 5 |
| 2024 | OFT: An accelerator with eager gradient prediction for attention trainingabstractWith the tremendous success of Transformer, resource-constrained edge devices are increasingly becoming the deployment target for attention-based models. On-device attention training can address the accuracy decline caused by static models' inability to adapt to dynamic environments efficiently while protecting data privacy. However, edge devices cannot meet the resource demands caused by batched gradient backpropagation and weight updates. Due to the inherent redundancy in human language, sparsification is the primary choice to alleviate the contradiction. Current sparse training methods consume high runtime costs in both forward propagation (FP) and backward propagation (BP) computations to handle irregular sparse patterns, which makes it inefficient to convert potential speedups into actual performance improvements and energy savings. Shengbing Zhang, Zhao Yang 0005, Meng Zhang 0047 |
ICCAD | 5 |
| 2024 | Efficient knowledge management for heterogeneous federated continual learning on resource-constrained edge devices
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Haoyang Wang 0014, Meng Zhang 0047 |
Future Gener. Comput. Syst. | 6 |
| 2024 | NDPGNN: A Near-Data Processing Architecture for GNN Training and Inference AccelerationabstractGraph neural networks (GNNs) require a large number of fine-grained memory accesses, which results in inefficient use of bandwidth resources. In this article, we introduce a near-data processing architecture tailored for GNN acceleration, named NDPGNN. NDPGNN provides different operating modes to meet the acceleration needs of various GNN frameworks while ensuring the configurability and scalability of the system. NDPGNN takes advantage of data locality characteristics to repeatedly distribute and utilize data, thereby reducing memory access requirements, and further improving memory access efficiency by combining a subgraph sparse node scheduling strategy with intermediate result reuse. We use data packaging to provide a higher effective data ratio for long-distance data transmission, thereby improving the utilization of the system’s limited bandwidth resources. Compared with the previous method, NDPGNN brings 5.68 times improvement in system performance while reducing energy consumption overhead by 8.49 times. Haoyang Wang 0014, Shengbing Zhang, Xiaoya Fan, Zhao Yang 0005, Meng Zhang 0047 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Equalized Aggregation for Heterogeneous Federated Mobile Edge LearningabstractFederated Learning (FL) is widely used in mobile edge applications. However, the heterogeneity issues of mobile edge devices pose significant challenges to the generalization of the global model in FL. In this paper, we propose LegoFL to simultaneously solve multiple heterogeneity issues in response to mobile edge computing characteristics. LegoFL identifies two types of heterogeneous behaviors in FL, namely heterogeneous parameter training and communication behaviors, to address multiple heterogeneity issues. These two types of heterogeneous behaviors result in feature and feature representation range mismatches between local communication parameters. To reduce these mismatches and improve the generalization of the global model, LegoFL dynamically distinguishes the parameter feature representation of different nodes using the global model's common feature as guidance. Then, under the connection states and system communication constraints, LegoFL dynamically selects contribution parameters on each device that can guarantee the generalization and performance of the global model for communication. Finally, to avoid the overfitting problem of the global model, heterogeneous local models are aggregated at the central server with matched feature representations. Extensive experiments on various datasets show that LegoFL achieves competitive performance. The accuracy and communication efficiency are improved by up to 12.86$\%$and 4.09× compared to state-of-the-art approaches. Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Jiaying Yang, Meng Zhang 0047 |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | Joint heterogeneity-aware personalized federated search for energy efficient battery-powered edge computing
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Jiaying Yang, Meng Zhang 0047 |
Future Gener. Comput. Syst. | 6 |
| 2022 | DCNN search and accelerator co-design: Improve the adaptability between NAS frameworks and embedded platforms
Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047 |
Integr. | 7 |
| 2022 | Memory-Computing Decoupling: A DNN Multitasking Accelerator With Adaptive Data ArrangementabstractMultiple deep neural networks (DNNs) are increasingly used in real-world intelligent applications, such as intelligent robotics and autonomous vehicles to collectively complete complicated tasks running on edge devices. Because each layer of the subtasks prefers a distinct dataflow due to the heterogeneity in shape and scale of the network layers, a variable dataflow approach on the DNN accelerators is urgently required. On DNN accelerators that enable multiple dataflows, however, we detect a dimension mismatch between parallel processing under the dataflow approach and linear data memory arrangement. When multiple DNN tasks share partial features or weights, the issue is further exacerbated. During processing, this mismatch causes a sluggish data supply from both off-chip and on-chip memory. Consequently, the overall throughput, performance, and energy efficiency suffer since DNN models are sensitive to data density. In this work, we reveal the mechanism behind this data dimension mismatch and present a series of metrics that quantify the influence on system performance. On this foundation, we offer a framework that tracks the data tensor dimension conversion and employs a flexible data arrangement over multi-DNN computation to adapt to dataflow variability. An accelerator architecture named data arrangement multi-DNN accelerator (DARMA) that features a data arrangement and distribution circuit and hierarchical memory for data dimension conversion is also presented. Since the mismatch is mitigated, the suggested accelerator outperforms current accelerators in terms of bandwidth and processing unit utilization. Through tests on VR/AR, MLperf, and other multitask applications, the evaluation results show that the proposed architecture provides both energy-efficiency and throughput improvements. Chuxi Li, Xiaoya Fan, Xiaoti Wu, Zhao Yang 0005, Meng Zhang 0047, Shengbing Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | Hardware-Aware NAS Framework with Layer Adaptive Scheduling on Embedded SystemabstractNeural Architecture Search (NAS) has been proven to be an effective solution for building Deep Convolutional Neural Network (DCNN) models automatically. Subsequently, several hardware-aware NAS frameworks incorporate hardware latency into the search objectives to avoid the potential risk that the searched network cannot be deployed on target platforms. However, the mismatch between NAS and hardware persists due to the absent of rethinking the applicability of the searched network layer characteristics and hardware mapping. A convolution neural network layer can be executed on various dataflows of hardware with different performance, with which the characteristics of on-chip data using varies to fit the parallel structure. This mismatch also results in significant performance degradation for some maladaptive layers obtained from NAS, which might achieved a much better latency when the adopted dataflow changes. To address the issue that the network latency is insufficient to evaluate the deployment efficiency, this paper proposes a novel hardware-aware NAS framework in consideration of the adaptability between layers and dataflow patterns. Beside, we develop an optimized layer adaptive data scheduling strategy as well as a coarse-grained reconfigurable computing architecture so as to deploy the searched networks with high power-efficiency by selecting the most appropriate dataflow pattern layer-by-layer under limited resources. Evaluation results show that the proposed NAS framework can search DCNNs with the similar accuracy to the state-of-the-art ones as well as the low inference latency, and the proposed architecture provides both power-efficiency improvement and energy consumption saving. Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047 |
ASP-DAC | 7 |