Shengbing Zhang

dblp:37/2842 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
20since 2021 · last 2026
0000-0002-2854-729XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 16 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A near CXL memory processing architecture for distributed graph neural network inference and training
abstract
Distributed Graph Neural Networks (GNNs) require efficient handling of both fine-grained memory accesses and cross memory-device communication, particularly when scaling to large graphs. However, existing acceleration solutions fail to adequately address low bandwidth utilization and scalability across different models and graph sizes. In this paper, we present OptGNN, a scalable heterogeneous distributed architecture tailored for GNNs. OptGNN addresses the challenges of fine-grained memory access by incorporating a near-memory processing mechanism, which improves internal bandwidth utilization. To optimize external communication, we introduce data packing and scheduling strategies that enhance cross memory-device data transfer efficiency. OptGNN achieves 5.7x performance improvement over baseline distributed GNN acceleration methods and 1.29x performance improvement over SOTA distributed GNN acceleration architecture CLAY. Additionally, the system is designed to support various GNN models and large-scale graphs while ensuring load balancing and high hardware utilization.
Shengbing Zhang, Xiaoya Fan
Connect. Sci.2
2026 TPQA: Efficient attention architecture with task-aware pattern-guided quantization
Shengbing Zhang, Yichao Yuan, Yawen Zhao 0010, Meng Zhang 0047
Future Gener. Comput. Syst.2
2026 PoDe-SORT: Robust multi-object tracking by modeling bounding box deformation
Shuaipeng Duan, Shengbing Zhang, Meng Zhang 0047
Knowl. Based Syst.2
2025 STAPC: A Sparse Training Accelerator for Efficient On-Device DNN Learning via Position Constraints
abstract
Leveraging sparsity to eliminate redundant computation and storage overhead is essential for enhancing the efficiency of on-device deep neural network (DNN) learning. However, due to a lack of assumptions about non-zero positions, existing sparse methods incur high costs for zero-position identification and allocation, making it difficult to achieve ideal acceleration. This paper demonstrates that knowing non-zero position constraints in advance during training can bypass these sparse processing overheads. We explore non-zero position constraints among operands for three typical activation functions in edge scenarios and propose: (1) a hardware-friendly sparse training algorithm to skip redundant gradient computations, enhancing training efficiency; and (2) a high-efficiency sparse training accelerator, STAPC, that estimates non-zero gradient positions, allowing costly sparse processing to be masked in parallel to reduce energy consumption. Compared to the baseline and other sparse training methods, the proposed method achieves energy efficiency gains of 2.2x, 1.38x, and 1.46x, respectively.
Shengbing Zhang, Meng Zhang 0047
ISCAS2
2024 Resource-Efficient Heterogenous Federated Continual Learning on Edge
abstract
Federated learning (FL) has been widely deployed on edge devices. In practical, the data collected by edge devices exhibits temporal variations. This leads to catastrophic forgetting issue. Continual learning methods can be used to address this problem. However, when deploying these methods in FL on edge devices, it is challenging to adapt to the limited resources and heterogeneous data of the deployed devices, which reduces the efficiency and effectiveness of federated continual learning (FCL). Therefore, this article proposes a resource-efficient heterogeneous FCL framework. This framework divides the global model into an adaptation part for new knowledge and a preservation part for old knowledge. The preservation part is used to address the catastrophic forgetting problem. Only the adaptation part is trained when learning new knowledge on a new task, reducing resource consumption. Additionally, the framework mitigates the impact of heterogeneous data through an aggregation method based on feature representation. Experimental results show that our method performs well in mitigating catastrophic forgetting in a resource-efficient manner.
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Haoyang Wang 0014, Meng Zhang 0047
DATE2
2024 OFT: An accelerator with eager gradient prediction for attention training
abstract
With the tremendous success of Transformer, resource-constrained edge devices are increasingly becoming the deployment target for attention-based models. On-device attention training can address the accuracy decline caused by static models' inability to adapt to dynamic environments efficiently while protecting data privacy. However, edge devices cannot meet the resource demands caused by batched gradient backpropagation and weight updates. Due to the inherent redundancy in human language, sparsification is the primary choice to alleviate the contradiction. Current sparse training methods consume high runtime costs in both forward propagation (FP) and backward propagation (BP) computations to handle irregular sparse patterns, which makes it inefficient to convert potential speedups into actual performance improvements and energy savings.
Shengbing Zhang, Zhao Yang 0005, Meng Zhang 0047
ICCAD2
2024 Efficient knowledge management for heterogeneous federated continual learning on resource-constrained edge devices
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Haoyang Wang 0014, Meng Zhang 0047
Future Gener. Comput. Syst.2
2024 NDPGNN: A Near-Data Processing Architecture for GNN Training and Inference Acceleration
abstract
Graph neural networks (GNNs) require a large number of fine-grained memory accesses, which results in inefficient use of bandwidth resources. In this article, we introduce a near-data processing architecture tailored for GNN acceleration, named NDPGNN. NDPGNN provides different operating modes to meet the acceleration needs of various GNN frameworks while ensuring the configurability and scalability of the system. NDPGNN takes advantage of data locality characteristics to repeatedly distribute and utilize data, thereby reducing memory access requirements, and further improving memory access efficiency by combining a subgraph sparse node scheduling strategy with intermediate result reuse. We use data packaging to provide a higher effective data ratio for long-distance data transmission, thereby improving the utilization of the system’s limited bandwidth resources. Compared with the previous method, NDPGNN brings 5.68 times improvement in system performance while reducing energy consumption overhead by 8.49 times.
Haoyang Wang 0014, Shengbing Zhang, Xiaoya Fan, Zhao Yang 0005, Meng Zhang 0047
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 RE-Specter: Examining the Architectural Features of Configurable CNN With Power Side-Channel
abstract
As domain-specific training data is recognized as valuable intellectual property, acquiring well-trained weights in Convolutional Neural Networks (CNN) has emerged as a new threat to the neural network design community. To design a CNN accelerator that is resilient to side-channel threats, it is crucial to have an accurate and efficient security-driven framework at the early design stage. However, there is no standard way to perform root-cause analysis on the power side channel that exists in FPGA-based CNN accelerators. Therefore, we build RE-Specter, a framework that facilitates security-driven design space exploration (DSE) across various building components, combination patterns, and parallelism configurations in CNNs. The goal is to fully understand the power side-channel effects resulting from architectural modifications or optimization decisions. We further compare the benchmarks considering precision, resource utilization, and power side-channel leakage. Finally, we experimentally explore the design space of various architectural features. The experimental results show that low-bit precision delivers more secure architectures (68.9× among DSPs, 2439× among LUTs) in Measurement-To-Disclosure (MTD), but mixed-precision strategies are necessary to maintain the model accuracy. For loop optimization, in 16-parallel scenario, accumulator-based architecture outperforms the architecture featuring an adder tree with the improvements of 8.28× in MTD and 1.38× in PST.
Lu Zhang 0074, Jingyu Wang 0004, Ruoyang Liu, Yifan He 0003, Yaolei Li, Yu Tai, Shengbing Zhang, Xiaoya Fan, Huazhong Yang, Yongpan Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2024 Equalized Aggregation for Heterogeneous Federated Mobile Edge Learning
abstract
Federated Learning (FL) is widely used in mobile edge applications. However, the heterogeneity issues of mobile edge devices pose significant challenges to the generalization of the global model in FL. In this paper, we propose LegoFL to simultaneously solve multiple heterogeneity issues in response to mobile edge computing characteristics. LegoFL identifies two types of heterogeneous behaviors in FL, namely heterogeneous parameter training and communication behaviors, to address multiple heterogeneity issues. These two types of heterogeneous behaviors result in feature and feature representation range mismatches between local communication parameters. To reduce these mismatches and improve the generalization of the global model, LegoFL dynamically distinguishes the parameter feature representation of different nodes using the global model's common feature as guidance. Then, under the connection states and system communication constraints, LegoFL dynamically selects contribution parameters on each device that can guarantee the generalization and performance of the global model for communication. Finally, to avoid the overfitting problem of the global model, heterogeneous local models are aggregated at the central server with matched feature representations. Extensive experiments on various datasets show that LegoFL achieves competitive performance. The accuracy and communication efficiency are improved by up to 12.86$\%$and 4.09× compared to state-of-the-art approaches.
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Jiaying Yang, Meng Zhang 0047
IEEE Trans. Mob. Comput.2
2023 SaGNN: a Sample-based GNN Training and Inference Hardware Accelerator
abstract
Graph neural networks (GNNs) operations contain a large number of irregular data operations and sparse matrix multiplications, resulting in the under-utilization of computing resources. The problem becomes even more complex and challenging when it comes to large graph training. Scaling GNN training is an effective solution. However, the current GNN operation accelerators do not support the mini-batch structure. We analyze the GNN operational characteristics from multiple aspects and take both the acceleration requirements in the GNN training and inference process into account, and then propose the SaGNN system structure. SaGNN offers multiple working modes to provide acceleration solutions for different GNN frameworks while ensuring system configurability and scalability. Compared to related works, SaGNN brings 5.0x improvement in system performance.
Haoyang Wang 0014, Shengbing Zhang, Kaijie Feng, Zhao Yang 0005
ISCAS2
2023 Joint heterogeneity-aware personalized federated search for energy efficient battery-powered edge computing
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Jiaying Yang, Meng Zhang 0047
Future Gener. Comput. Syst.2
2023 A high-efficiency spaceborne processor for hybrid neural networks
Shengbing Zhang, Libo Chang
Neurocomputing2
2023 A Noise-Driven Heterogeneous Stochastic Computing Multiplier for Heuristic Precision Improvement in Energy-Efficient DNNs
abstract
Stochastic computing (SC) has become a promising approximate computing solution by its negligible resource occupancy and ultralow energy consumption. As a potential replacement of accurate multiplication, SC can dramatically mitigate the problematic power consumption by DNNs. However, current SC-multipliers illustrate an extremely imbalanced accuracy across product space, i.e., neglectable noise with large products but significant noise for small ones, which is discordant to the distribution of products by the sparse matrix in neural computing. In this article, we present a heterogeneous SC-multiplier that heuristically performs three divergent approximating multiplication, including “set-to-0,” “look-up-table,” and “low-discrepancy-SC,” for appropriate precision-provision in the whole space of products. Due to those popular DNN models cannot achieve consensus on the boundaries of above operations, a training-involved method is proposed to determine the settings with limited overhead. In this way, those models successively learn the SC-operation characters and exhibit a definitely improvement on network precision. The experiment shows that, for single multiplication, the product noise can be restrained by 36.86% on average, and for multiplication in multiple network models, the accuracy improvement reaches to 5.5% on average. Furthermore, a group of proposed logic-reduction techniques can improve the energy efficiency by 65% in the system-level evaluation.
Danghui Wang, Shengbing Zhang, Xiaoya Fan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 DCNN search and accelerator co-design: Improve the adaptability between NAS frameworks and embedded platforms
Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047
Integr.3
2022 MemUnison: A Racetrack-ReRAM-Combined Pipeline Architecture for Energy-Efficient in-Memory CNNs
abstract
Though ReRAM has been greatly successful in reducing energy consumption of various neural networks, it still suffers write amplification in energy, which impedes ReRAM to provide efficient storage for the ubiquitous streaming data in CNNs, such as feature-maps. Racetrack memory, an emerging magnetic memory technique, is a proper candidate to hold streaming data since it enjoys fast sequential-access with ultra-low operating energy in read and write. In this work, we propose a hybrid processing-in-memory architecture, called MemUnison, that coordinates ReRAM and racetrack to overcome the expenditure storage of streaming data in ReRAM. By placing feature-maps in racetrack and leaving weights in ReRAM, a datapath is constructed between the two sides to form a fetch-process-writeback pipeline. As the invalid-shifts of the racetrack memory incurs a large amount of pipeline bubble, we propose a row-based access that can read and write a feature-map without any invalid-shifts. For the row-based operation, a cohesive controlling method is proposed to coordinate racetrack and ReRAM. In runtime, convolution kernels are scheduled in ReRAM banks for cross-channel calculations of one row, by which computing complexity of a convolutional layer can be reduced by 4 orders of magnitude, excessing the 2 order of reduction by traditional ReRAM.
Danghui Wang, Shengbing Zhang, Xiaoya Fan
IEEE Trans. Computers4
2022 Memory-Computing Decoupling: A DNN Multitasking Accelerator With Adaptive Data Arrangement
abstract
Multiple deep neural networks (DNNs) are increasingly used in real-world intelligent applications, such as intelligent robotics and autonomous vehicles to collectively complete complicated tasks running on edge devices. Because each layer of the subtasks prefers a distinct dataflow due to the heterogeneity in shape and scale of the network layers, a variable dataflow approach on the DNN accelerators is urgently required. On DNN accelerators that enable multiple dataflows, however, we detect a dimension mismatch between parallel processing under the dataflow approach and linear data memory arrangement. When multiple DNN tasks share partial features or weights, the issue is further exacerbated. During processing, this mismatch causes a sluggish data supply from both off-chip and on-chip memory. Consequently, the overall throughput, performance, and energy efficiency suffer since DNN models are sensitive to data density. In this work, we reveal the mechanism behind this data dimension mismatch and present a series of metrics that quantify the influence on system performance. On this foundation, we offer a framework that tracks the data tensor dimension conversion and employs a flexible data arrangement over multi-DNN computation to adapt to dataflow variability. An accelerator architecture named data arrangement multi-DNN accelerator (DARMA) that features a data arrangement and distribution circuit and hierarchical memory for data dimension conversion is also presented. Since the mismatch is mitigated, the suggested accelerator outperforms current accelerators in terms of bandwidth and processing unit utilization. Through tests on VR/AR, MLperf, and other multitask applications, the evaluation results show that the proposed architecture provides both energy-efficiency and throughput improvements.
Chuxi Li, Xiaoya Fan, Xiaoti Wu, Zhao Yang 0005, Meng Zhang 0047, Shengbing Zhang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2021 Hardware-Aware NAS Framework with Layer Adaptive Scheduling on Embedded System
abstract
Neural Architecture Search (NAS) has been proven to be an effective solution for building Deep Convolutional Neural Network (DCNN) models automatically. Subsequently, several hardware-aware NAS frameworks incorporate hardware latency into the search objectives to avoid the potential risk that the searched network cannot be deployed on target platforms. However, the mismatch between NAS and hardware persists due to the absent of rethinking the applicability of the searched network layer characteristics and hardware mapping. A convolution neural network layer can be executed on various dataflows of hardware with different performance, with which the characteristics of on-chip data using varies to fit the parallel structure. This mismatch also results in significant performance degradation for some maladaptive layers obtained from NAS, which might achieved a much better latency when the adopted dataflow changes. To address the issue that the network latency is insufficient to evaluate the deployment efficiency, this paper proposes a novel hardware-aware NAS framework in consideration of the adaptability between layers and dataflow patterns. Beside, we develop an optimized layer adaptive data scheduling strategy as well as a coarse-grained reconfigurable computing architecture so as to deploy the searched networks with high power-efficiency by selecting the most appropriate dataflow pattern layer-by-layer under limited resources. Evaluation results show that the proposed NAS framework can search DCNNs with the similar accuracy to the state-of-the-art ones as well as the low inference latency, and the proposed architecture provides both power-efficiency improvement and energy consumption saving.
Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047
ASP-DAC3
2021 SPACE: Sparsity Propagation Based DCNN Training Accelerator on Edge
Chuxi Li, Shengbing Zhang
ICA3PP (2)7
2021 A Reconfigurable Neural Network Processor With Tile-Grained Multicore Pipeline for Object Detection on FPGA
abstract
In order to improve the computational efficiency of convolutional neural networks (CNNs) for object detection on reconfigurable platforms such as field-programmable gate arrays (FPGAs), we propose a CNN processor with hierarchical pipelining and multicore reconfigurable computing based on parallel parameter constraints. First, we propose a pipelined multicore processing architecture that can adapt to computations and on-chip memory requirements of different convolutional layers. We present the design of a CNN processor that can configure the computing units while adjusting the interconnection of multicore to improve the utilization of reconfigurable computing resources. Second, we propose an elastic on-chip buffer and a data access approach by dynamically configuring addresses to better utilize on-chip memory. Meanwhile, we present a cross-layer feature map fusion strategy based on computing near memory (CNM) to reduce off-chip memory accesses. Finally, we propose a scheduling algorithm for pipelined tasks to improve the computational efficiency and throughput of the proposed processor. For evaluation, the well-known object detection methods (RetinaNet-ResNet-50, MobileNetV2-SSDLite, and YOLOv3) performed using the proposed CNN processor on the ZCU102 platform and reached the throughput of 1503, 1066, and 809 GOPS and the computational efficiency of 0.79, 0.62, and 0.36 GOPS/DSP, respectively. The designed processor realized a better tradeoff between computing efficiency and detection accuracy compared with the recently proposed object detection CNNs on FPGA.
Libo Chang, Shengbing Zhang, Huimin Du
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Towards Energy Efficient Architecture for Spaceborne Neural Networks Computation
Shengbing Zhang
ICA3PP (2)2
2012 Analog layout retargeting with geometric programming and constrains symbolization method
abstract
To satisfy the requirements of complex and special analog layout constraints, a constrains symbolization method based on geometric programming for analog layout retargeting is presented in this paper. The approach is to build symbolic template for layouts, then uses geometric programming (GP) to achieve new technology design rules, implement device symmetry and matching constraints, and manage parasitics optimization. The GP, a class of non-linear optimization problem, can be transferred or fitted into a convex optimization problem. Therefore, a global optimum solution can be achieved. The symbolization method ensures the layout retargeting automatically. The efficiency and effectiveness of the proposed algorithm, as compared with the other existing methods, are demonstrated by a basic case-study example and a two-stage Miller-compensated operational amplifier.
Shaoxi Wang, Xiaoya Fan, Shengbing Zhang, Ming-e Jing
ISCAS3