Jingyang Zhu

dblp:160/6026 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 4 first-author · 13 since 2021Systems, architecture and hardware · 10 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploiting the Irregular Input Sparsity in Systolic Array-based DNN Accelerators via Local Soft Pooling
abstract
One promising approach to mitigating the computational complexity of deep neural networks is to leverage the sparsity of input activations that results from the application of the ReLU function. However, the irregular distribution of zero-valued inputs poses a challenge for efficient implementation in existing regular architectures, such as systolic arrays. Previous works usually depend on specialized architectures to bypass the redundant computations during runtime. In contrast to these prior strategies, we propose a local soft pooling method to efficiently exploit the irregular input sparsity in systolic array-based architectures. Through local soft pooling, adjacent input rows can be safely merged at runtime, compressing the sparse input matrix into a compact format that is only 1/3 to 1/2 of its original size. The compact matrix can then be directly fed into the systolic array for computation. A computation saving of 67.78% is achieved across various networks on both CIFAR-10 and ImageNet with negligible accuracy loss. As a result, the throughput and energy efficiency are improved by 2.72 and 2.07 times, respectively.
Desheng Fu, Jingbo Jiang, Jingyang Zhu, Xizi Chen, Chi-Ying Tsui
ASP-DAC4
2026 Asynchronous Satellite Federated Learning with Intermittent Ground-to-Satellite Links
Ruanjun Li, Jingyang Zhu, Yong Zhou 0006, Yuanming Shi, Linling Kuang, Chunxiao Jiang
ICC2
2026 Unseen Cost of Space Computing: Quantifying LEO Battery Aging via Physics-Driven Modeling
abstract
Low Earth Orbit (LEO) satellite constellations in the 6G era are evolving into intelligent in-orbit computational platforms, forming Space Computing Power Networks (SCPNs) to deliver global-scale computing services. However, the intensive computation within SCPN incurs a significant "unseen cost": the frequent charge-discharge cycles accelerate the physical degradation of satellites’ life-limiting and high-cost batteries, thereby threatening the long-term operational viability of such a system. Existing approaches, often relying on indirect metrics like Depth of Discharge (DoD) and neglecting the complex, nonlinear degradation process of battery aging, fail to accurately quantify this cost. To address this, we introduce a high-fidelity, physics-driven model that quantitatively links computational workload parameters to the nonlinear battery degradation. Building on this model, we formulate a degradation-aware scheduling problem and analyze heuristic policies across different energy regimes. Simulations reveal that the optimal strategy should be adaptive: in solar-rich conditions, a myopic policy maximizing instantaneous solar utilization is superior, whereas under energy scarcity, a reactive policy leveraging real-time battery state significantly extends lifetime.
Jingyang Zhu, Yuanming Shi, Khaled Ben Letaief
ICC2
2026 Robust Information Bottleneck for Satellite Edge Inference Over MIMO Channel
Jielin Zhu, Jingyang Zhu, Youlong Wu, Ting Wang 0001, Yuanming Shi, Wei Chen 0002, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2026 Satellite Federated Fine-Tuning for Foundation Models in Space Computing Power Networks
abstract
Advancements in artificial intelligence and low-earth orbit satellites have promoted the application of large remote sensing foundation models (FMs) for various downstream tasks. However, direct downloading of these models for fine-tuning on the ground is impeded by privacy concerns and limited bandwidth. Satellite federated learning (FL) offers a solution by enabling model fine-tuning directly on-board satellites and aggregating model updates without data downloading. Nevertheless, for large FMs, the computational capacity of satellites is insufficient to support effective on-board fine-tuning in traditional satellite FL frameworks. To address these challenges, we propose a satellite-ground collaborative federated fine-tuning framework. The key of the framework lies in how to reasonably decompose and allocate model components to alleviate insufficient on-board computation capabilities. During fine-tuning, satellites exchange intermediate results with ground stations or other satellites for forward propagation and back propagation, which brings communication challenges due to the special communication topology of space transmission networks, such as intermittent satellite-ground communication, short duration of satellite-ground communication windows, and unstable inter-orbit inter-satellite links. To reduce transmission delays, we further introduce tailored communication strategies that integrate both communication and computing resources. Specifically, we propose a parallel intra-orbit communication strategy, a topology-aware satellite-ground communication strategy, and a latency-minimization inter-orbit communication strategy to reduce space communication costs. Simulation results demonstrate significant reductions in training time to 33% of on-board training time.
Jingyang Zhu, Ting Wang 0001, Yuanming Shi, Chunxiao Jiang, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.2
2025 Hierarchical Federated Learning with Integrated Sensing-Communication-Computation Over Space-Air-Ground Integrated Networks
abstract
Federated learning has achieved significant advancements in edge artificial intelligence (AI) by addressing issues related to data privacy and communication overload. Moreover, hierarchical federated learning over space-air-ground integrated networks (FedSAG), which consists of low-Earth orbit (LEO) satellites, unmanned aerial vehicles (UAVs), and edge devices, aims to provide AI services in sparsely populated regions lacking ground communication infrastructure. However, previous studies have overlooked the essential sensing process required for acquiring training data, potentially compromising training efficiency and model accuracy. In this paper, we propose an integrated sensing-communication-computation (ISCC) enabled FedSAG system, which allows remote edge devices to collect data via wireless sensing and collaboratively train a global model without sharing local data. We then analyse the convergence of the ISCC-enabled FedSAG and formulate two optimization problems. The first aims to minimize sensing variance under energy and time constraints, while the second seeks to reduce transmission energy through optimal route selection between UAVs and LEO satellite. We reformulate the problems to a minimum spanning tree and propose a Chu-Liu-Edmonds algorithm based a two-stage optimization. Simulation results demonstrate that our proposed algorithm significantly enhances convergence performance and reduces energy consumption.
Zhanpeng Yang, Jingyang Zhu, Dingzhu Wen, Yuanming Shi, Wei Chen 0002
ICC3
2025 Topology-Aware Routing for Federated Learning Over Multi-Layer Satellite Networks
abstract
Recent advancements in space computing power networks, particularly the integration of onboard computing capabilities in Low Earth Orbit (LEO) satellites, have paved the way for federated learning (FL) in satellite networks. Despite its potential, satellite FL faces unique challenges, such as the dynamic nature of satellite networks and the instability of inter-orbit communication links, which complicate global model aggregation. To address these challenges, we explore FL over multi-layer satellite networks, incorporating LEO, Medium Earth Orbit (MEO), and Geostationary Earth Orbit (GEO) satellites. Specifically, by modeling the dynamic network as a series of time-varying graph snapshots, we propose a novel topology-aware FL framework. To optimize the aggregation routing in the multi-layer satellite network, we leverage the directed minimum spanning tree (DMST) problem in graph theory and introduce a communication-efficient satellite aggregation routing algorithm (CESAR), which effectively reduces communication overhead and aggregation delays, ensuring efficient training and model updates across the satellite network. Extensive experimental results validate the efficacy of the proposed framework, demonstrating its potential to overcome the inherent challenges of satellite FL and significantly advance the capabilities of multi-layer satellite networks.
Ruanjun Li, Jingyang Zhu, Yijie Mao, Yuanming Shi, Ting Wang 0001, Chunxiao Jiang
WCNC2
2025 Satellite edge artificial intelligence with large models: architectures and technologies
Yuanming Shi, Jingyang Zhu, Chunxiao Jiang, Linling Kuang, Khaled Ben Letaief
Sci. China Inf. Sci.2
2025 Hierarchical Learning and Computing Over Space-Ground Integrated Networks
abstract
Space-ground integrated networks hold great promise for providing global connectivity, particularly in remote areas where large amounts of valuable data are generated by Internet of Things (IoT) devices, but lacking terrestrial communication infrastructure. The massive data is conventionally transferred to the cloud server for centralized artificial intelligence (AI) models training, raising huge communication overhead and privacy concerns. To address this, we propose a hierarchical learning and computing framework, which leverages the low-latency characteristic of low-earth-orbit (LEO) satellites and the global coverage of geostationary-earth-orbit (GEO) satellites, to provide global aggregation services for locally trained models on ground IoT devices. Due to the time-varying nature of satellite network topology and the energy constraints of LEO satellites, efficiently aggregating the received local models from ground devices on LEO satellites is highly challenging. By leveraging the predictability of inter-satellite connectivity, modeling the space network as a directed graph, we formulate a network energy minimization problem for model aggregation, which turns out to be aDirected Steiner Tree (DST)problem. We propose a topology-aware energy-efficient routing (TAEER) algorithm to solve theDSTproblem by finding a minimum spanning arborescence on a substitute directed graph. Extensive simulations under real-world space-ground integrated network settings demonstrate that the proposed TAEER algorithm significantly reduces energy consumption and outperforms benchmarks.
Jingyang Zhu, Yuanming Shi, Yong Zhou 0006, Chunxiao Jiang, Linling Kuang
IEEE Trans. Mob. Comput.1
2024 Satellite Federated Fine-Tuning for Foundation Models: Architecture Design and System Optimization
abstract
With the surge in the number of low earth orbit (LEO) satellites, continuous research has emerged on using satellite data to train artificial intelligence models. On one hand, traditional centralized training on the ground is not feasible due to privacy concerns and limited bandwidth for downloading raw satellite data. On the other hand, due to the limited energy and computational capability of satellites, training directly on satellites suffers from prolonged latency, especially for large models. To alleviate these issues, we propose a novel satellite-ground collaborative federated fine-tuning architecture, where ground stations (GSs) and satellites collaboratively train a global model without the need for data downloads. In this proposed architecture, satellites serve as edge devices and the ground server serves as a coordinator. However, the short satellite-ground communication windows caused by the high mobility of satellites and the substantial intra-orbit data transmission bring special challenges to the transmission process of federated edge learning. To tackle these challenges, we carefully design the satellite-ground collaborative fine-tuning architecture and utilize an optimized ring all-reduce algorithm and network flow algorithm to enhance the intra-orbit and ground-satellite transmissions, respectively. Experimental results demonstrate that our proposed architecture significantly reduces the training time by 40% compared to training solely on satellite.
Peng Yang 0027, Jingyang Zhu, Dingzhu Wen, Ting Wang 0001, Yong Zhou 0006, Yuanming Shi, Chunxiao Jiang
GLOBECOM3
2024 Latency Minimization for Wireless Federated Learning With Heterogeneous Local Model Updates
abstract
In this article, we study the latency minimization problem for a wireless federated learning (FL) system with heterogeneous computation capability, where different edge devices perform different numbers of local model updates in each communication round. We formulate a total latency minimization problem with probabilistic device selection, taking into account both the communication and computation latency in the whole FL procedure. However, it is highly challenging to optimally solve this problem due to the coupling issues of model convergence and latency minimization problem caused by the heterogeneity of local model updates. Through convergence analysis, we reveal that decoupling the resource allocation variables from the model convergence is essential to reduce the problem to a single-round latency minimization problem. To solve this simplified problem, we propose an alternating optimization scheme to jointly consider communication and computation resource allocation and mitigate the straggler effect. We prove that the resulting subproblems, i.e., bandwidth and computation capacity allocation, are both convex and can be optimally solved in closed form, respectively. Simulation results show that compared with the baseline scheme that allocates the communication and computation resources equally across edge devices, the proposed scheme can achieve up to 47.04% single-round latency reduction.
Jingyang Zhu, Yuanming Shi, Min Fu 0003, Yong Zhou 0006, Youlong Wu, Liqun Fu 0001
IEEE Internet Things J.1
2024 Over-the-Air Federated Learning and Optimization
abstract
Federated learning (FL), as an emerging distributed machine learning paradigm, allows a mass of edge devices to collaboratively train a global model while preserving privacy. In this tutorial, we focus on FL via over-the-air computation (AirComp), which is proposed to reduce the communication overhead for FL over wireless networks at the cost of compromising in the learning performance due to model aggregation error arising from channel fading and noise. We first provide a comprehensive study on the convergence of AirComp-based FEDAVG (AIRFEDAVG) algorithms under both strongly convex and non-convex settings with constant and diminishing learning rates in the presence of data heterogeneity. Through convergence and asymptotic analysis, we characterize the impact of aggregation error on the convergence bound and provide insights for system design with convergence guarantees. Then we derive convergence rates for AIRFEDAVG algorithms for strongly convex and non-convex objectives. For different types of local updates that can be transmitted by edge devices (i.e., local model, gradient, and model difference), we reveal that transmitting local model in AIRFEDAVG may cause divergence in the training procedure. In addition, we consider more practical signal processing schemes to improve the communication efficiency and further extend the convergence analysis to different forms of model aggregation error caused by these signal processing schemes. Extensive simulation results under different settings of objective functions, transmitted local information, and communication schemes verify the theoretical conclusions.
Jingyang Zhu, Yuanming Shi, Yong Zhou 0006, Chunxiao Jiang, Wei Chen 0002, Khaled Ben Letaief
IEEE Internet Things J.1
2024 Satellite Federated Edge Learning: Architecture Design and Convergence Analysis
abstract
The proliferation of low-earth-orbit (LEO) satellite networks leads to the generation of vast volumes of remote sensing data which is traditionally transferred to the ground server for centralized processing, raising privacy and bandwidth concerns. Federated edge learning (FEEL), as a distributed machine learning approach, has the potential to address these challenges by sharing only model parameters instead of raw data. Although promising, the dynamics of LEO networks, characterized by the high mobility of satellites and short ground-to-satellite link (GSL) duration, pose unique challenges for FEEL. Notably, frequent model transmission between the satellites and ground incurs prolonged waiting time and large transmission latency. This paper introduces a novel FEEL algorithm, named FEDMEGA, tailored to LEO mega-constellation networks. By integrating inter-satellite links (ISL) for intra-orbit model aggregation, the proposed algorithm significantly reduces the usage of low datarate and intermittent GSL. Our proposed method includes a ring all-reduce based intra-orbit aggregation mechanism, coupled with a network flow-based transmission scheme for global model aggregation, which enhances transmission efficiency. Theoretical convergence analysis is provided to characterize the algorithm performance. Extensive simulations show that our FEDMEGA algorithm outperforms existing satellite FEEL algorithms, exhibiting an approximate 30% improvement in convergence rate.
Yuanming Shi, Jingyang Zhu, Yong Zhou 0006, Chunxiao Jiang, Khaled Ben Letaief
IEEE Trans. Wirel. Commun.3
2023 Latency Minimization for Wireless Federated Learning with Heterogeneous Local Updates
abstract
In this paper, we study the latency minimization problem for a wireless federated learning (FL) system with heterogeneous computation capability, where different edge devices perform different numbers of local updates in each communication round. We formulate a total latency minimization problem, taking into account both the communication and computation latency in the whole FL procedure. We reveal that decoupling the resource allocation variables from the model convergence is essential to reduce the problem to a single-round latency minimization problem. To solve this simplified problem, we propose an alternating optimization scheme to jointly consider communication and computation resource allocation and mitigate the straggler effect. We prove that the resulting sub-problems, i.e., bandwidth and computation capacity allocation, are both convex and can be optimally solved in closed form, respectively. Simulations show that compared with the baseline scheme that allocates the communication and computation resources equally across edge devices, the proposed scheme can achieve single-round latency reduction.
Jingyang Zhu, Yuanming Shi, Min Fu 0003, Yong Zhou 0006, Youlong Wu, Liqun Fu 0001
WCNC1
2023 Tight Compression: Compressing CNN Through Fine-Grained Pruning and Weight Permutation for Efficient Implementation
abstract
The unstructured sparsity after pruning poses a challenge to the efficient implementation of deep learning models in existing regular architectures like systolic arrays. On the other hand, coarse-grained structured pruning is suitable for implementation in regular architectures but tends to have higher accuracy loss than unstructured pruning when the pruned models are of the same size. In this work, we propose a model compression method based on a novel weight permutation scheme to fully exploit the fine-grained weight sparsity in the hardware design. Through permutation, the optimal arrangement of the weight matrix is obtained, and the sparse weight matrix is further compressed to a small and dense format to make full use of the hardware resources. Two pruning granularities are explored. In addition to the unstructured weight pruning, we also propose a more fine-grained subword-level pruning to further improve the compression performance. Compared to the state-of-the-art works, the matrix compression rate is significantly improved from$5.88\times $to$14.13\times $. As a result, the throughput and energy efficiency are improved by 2.75 and 1.86 times, respectively.
Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Fast Convergence Algorithm for Analog Federated Learning
abstract
In this paper, we consider federated learning (FL) over a noisy fading multiple access channel (MAC), where an edge server aggregates the local models transmitted by multiple end devices through over-the-air computation (AirComp). To realize efficient analog federated learning over wireless channels, we propose an AirComp-based FedSplit algorithm, where a threshold-based device selection scheme is adopted to achieve reliable local model uploading. In particular, we analyze the performance of the proposed algorithm and prove that the proposed algorithm linearly converges to the optimal solutions under the assumption that the objective function is strongly convex and smooth. We also characterize the robustness of proposed algorithm to the ill-conditioned problems, thereby achieving fast convergence rates and reducing communication rounds. A finite error bound is further provided to reveal the relationship between the convergence behavior and the channel fading and noise. Our algorithm is theoretically and experimentally verified to be much more robust to the ill-conditioned problems with faster convergence compared with other benchmark FL algorithms.
Shuhao Xia, Jingyang Zhu, Yong Zhou 0006, Yuanming Shi, Wei Chen 0002
ICC2
2020 Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based Permutation
abstract
The unstructured sparsity after pruning poses a challenge to the efficient implementation of deep learning models in existing regular architectures like systolic arrays. The coarse-grained structured pruning, on the other hand, tends to have higher accuracy loss than unstructured pruning when the pruned models are of the same size. In this work, we propose a compression method based on the unstructured pruning and a novel weight permutation scheme. Through permutation, the sparse weight matrix is further compressed to a small and dense format to make full use of the hardware resources. Compared to the state-of-the-art works, the matrix compression rate is effectively improved from 5.88x to 10.28x. As a result, the throughput and energy efficiency are improved by 2.12 and 1.57 times, respectively.
Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui
DAC2
2019 CompRRAE: RRAM-based convolutional neural network accelerator with reduced computations through a runtime activation estimation
abstract
Recently Resistive-RAM (RRAM) crossbar has been used in the design of the accelerator of convolutional neural networks (CNNs) to solve the memory wall issue. However, the intensive multiply-accumulate computations (MACs) executed at the crossbars during the inference phase are still the bottleneck for the further improvement of energy efficiency and throughput. In this work, we explore several methods to reduce the computations for the RRAM-based CNN accelerators. First, the output sparsity resulting from the widely employed Rectified Linear Unit is exploited, and a significant portion of computations are bypassed through an early detection of the negative output activations. Second, an adaptive approximation is proposed to terminate the MAC early when the sum of the partial results of the remaining computations is considered to be within a certain range of the intermediate accumulated result and thus has an insignificant contribution to the inference. In order to determine these redundant computations, a novel runtime estimation on the maximum and minimum values of each output activation is developed and used during the MAC operation. Experimental results show that around 70% of the computations can be reduced during the inference with a negligible accuracy loss smaller than 0.2%. As a result, the energy efficiency and the throughput are improved by over 2.9 and 2.8 times, respectively, compared with the state-of-the-art RRAM-based accelerators.
Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui
ASP-DAC2
2019 SubMac: Exploiting the subword-based computation in RRAM-based CNN accelerator for energy saving and speedup
Xizi Chen, Jingbo Jiang, Jingyang Zhu, Chi-Ying Tsui
Integr.3
2018 A high-throughput and energy-efficient RRAM-based convolutional neural network using data encoding and dynamic quantization
abstract
To solve the scaling, memory wall and high power density issues, recently RRAM-based accelerators, which show a better energy and area efficiency compared with the CMOS-based counterparts, have been proposed for convolutional neural networks. However, the RRAM-based architectures still face several design challenges, including the high energy and timing overhead at the analog/digital (A/D) conversion and interfacing circuits. To address these issues, we propose several novel optimization schemes in this work. First an encoding scheme for the synaptic weights and the input feature maps is proposed to reduce the energy of the in-situ computation and the bit-resolution of the A/D conversion. Then the resolution of the A/D conversion is further optimized for a lower energy consumption. Moreover, a dynamic quantization scheme for the multiply-accumulate operations (MACs) is proposed to improve the throughput and the energy efficiency by reducing the number of partial products. Experimental results show that the throughput, the energy efficiency and the area efficiency are improved by 2 to 4 times when compared with the state-of-the-art RRAM-based accelerators.
Xizi Chen, Jingbo Jiang, Jingyang Zhu, Chi-Ying Tsui
ASP-DAC3
2018 SparseNN: An energy-efficient neural network accelerator exploiting input and output sparsity
abstract
Contemporary Deep Neural Network (DNN) contains millions of synaptic connections with tens to hundreds of layers. The large computational complexity poses a challenge to the hardware design. In this work, we leverage the intrinsic activation sparsity of DNN to substantially reduce the execution cycles and the energy consumption. An end-to-end training algorithm is proposed to develop a lightweight (less than 5% overhead) run-time predictor for the output activation sparsity on the fly. Furthermore, an energy-efficient hardware architecture, SparseNN, is proposed to exploit both the input and output sparsity. SparseNN is a scalable architecture with distributed memories and processing elements connected through a dedicated on-chip network. Compared with the state-of-the-art accelerators which only exploit the input sparsity, SparseNN can achieve a 10%-70% improvement in throughput and a power reduction of around 50%.
Jingyang Zhu, Jingbo Jiang, Xizi Chen, Chi-Ying Tsui
DATE1
2017 BHNN: A memory-efficient accelerator for compressing deep neural networks with blocked hashing techniques
abstract
In this paper, we propose a novel algorithm for compressing neural networks to reduce the memory requirements by using blocked hashing techniques. By adding blocked constraints on top of the conventional hashing technique, the test error rate is maintained while the spatial locality for the computations is preserved. Using this scheme, the synaptic connections are compressed by at least an order (10×) compared with the plain neural network with virtually no prediction accuracy loss. Compared with other compression techniques, the proposed algorithm achieves the best performance in the heavy compression regions. The blocked hashing techniques are also hardware friendly, of which the memory hierarchy of the hardware architecture can be efficiently implemented. To demonstrate the hardware efficiency, we implement the hardware architecture of the deep neural networks using the proposed blocked hashing techniques on a Xilinx Virtex-7 FPGA board. With a hardware parallelism of 32, the accelerator achieves a speed-up of 22× over the CPU, and 3~5× over the GPU in the inference phase.
Jingyang Zhu, Zhiliang Qian, Chi-Ying Tsui
ASP-DAC1
2016 LRADNN: High-throughput and energy-efficient Deep Neural Network accelerator using Low Rank Approximation
abstract
In this work, we propose an energy-efficient hardware accelerator for Deep Neural Network (DNN) using Low Rank Approximation (LRADNN). Using this scheme, inactive neurons in each layer of the DNN are dynamically identified and the corresponding computations are then bypassed. Accordingly, both the memory accesses and the arithmetic operations associated with these inactive neurons can be saved. Therefore, compared to the architectures using the direct feed-forward algorithm, LRADNN can achieve a higher throughput as well as a lower energy consumption with negligible prediction accuracy loss (within 0.1%). We implement and synthesize the proposed accelerator using TSMC 65nm technology. From the experimental results, a 31% to 53% energy reduction together with a 22% to 43% throughput increase can be achieved.
Jingyang Zhu, Zhiliang Qian, Chi-Ying Tsui
ASP-DAC1
2016 BiLink: A high performance NoC router architecture using bi-directional link with double data rate
Jingyang Zhu, Zhiliang Qian, Chi-Ying Tsui
Integr.1