Qiyang Zhang 0001

dblp:17/9915-1 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0001-5585-6613ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 3 first-author · 11 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Non-functional certification of edge-computing satellite systems
abstract
Satellite telecommunication networks are playing an increasingly pivotal role in modern communication infrastructures, owing to their expansive coverage, high reliability, and growing capabilities in computing, storage, and bandwidth. In response to evolving market demands, mobile network operators are progressively integrating satellite systems with edge-cloud computing platforms to deliver advanced networking functionalities within a unified architecture. This integration places strong demands on the non-functional assessment (e.g., reliability, availability, and resource efficiency) of satellite-based edge nodes, introducing unprecedented challenges due to their unique operational constraints. In this paper, we propose a lightweight certification framework tailored for satellite computing systems, designed to assess and validate the non-functional posture of satellite edge networks. Our approach explicitly addresses the distinctive characteristics of satellite environments, including intermittent connectivity and constrained resource availability. We validate the proposed scheme through a realistic testbed implementation, modeling a 5G-enabled satellite edge node based on the Tiansuan satellite constellation, an experimental platform jointly developed by Beijing University of Posts and Telecommunications, Spacety, and Peking University.
Filippo Berto, Marco Anisetti, Qiyang Zhang 0001, Shangguang Wang, Claudio A. Ardagna
Comput. Networks3
2026 Optimizing Multi-DNN Inference on Mobile Devices Through Heterogeneous Processor Co-Execution
abstract
Deep Neural Networks (DNNs) are increasingly adopted across various industries, driving the demand for deploying their capabilities on mobile devices. However, current mobile inference frameworks often rely on a single processor to execute each model inference, limiting hardware utilization and leading to suboptimal performance and energy efficiency. Expanding DNN accessibility on mobile platforms requires more adaptive and resource-efficient solutions to meet increasing computational demands without compromising device functionality. Nevertheless, performing parallel inference of multiple DNNs on heterogeneous processors remains a significant challenge. Existing studies have explored partitioning DNN operations into subgraphs to enable parallel execution across heterogeneous processors. However, these approaches typically generate excessive subgraphs based solely on hardware compatibility, increasing scheduling complexity and memory management overhead. To address these limitations, we propose the Advanced Multi-DNN Model Scheduling (ADMS) strategy that optimizes multi-DNN inference across heterogeneous processors on mobile devices. ADMS constructs an offline subgraph partitioning strategy that considers both hardware support for operations and scheduling granularity. It also employs a processor-state-aware scheduling algorithm to dynamically balance workloads based on real-time system conditions. This ensures efficient workload distribution and maximizes the utilization of available processors. Experimental results demonstrate that, compared to vanilla inference frameworks, ADMS achieves a 4.04× reduction in multi-DNN inference latency.
Yunquan Gao, Praveen Kumar Donta, Chinmaya Kumar Dehury, Xiujun Wang, Dusit Niyato, Qiyang Zhang 0001
IEEE Trans. Mob. Comput.7
2026 Task-Aware Collaborative Inference and Fine-Grained DNN Partitioning in MEC Networks
abstract
Mobile devices (MDs) are increasingly incorporating deep neural network (DNN) inference into their systems due to the rapid growth of intelligent applications. Mobile edge computing-based distributed DNN collaborative inference has gained popularity due to limited on-device computation and energy budgets. However, the resource competition among MDs, along with the coupling of collaborative inference tasks across MDs and servers, creates significant challenges for efficient resource management. This issue is further exacerbated by the complexity of directed acyclic graph (DAG)-structured DNNs. Most prior studies do not jointly address the dual challenges of partitioning complex-structured DNNs and leveraging advanced optimization for collaborative inference, and their resilience to channel condition fluctuations remains underexplored. To address these challenges, we propose a novel task-aware collaborative inference framework. First, we devise a fine-grained partitioning point search algorithm based on a bidirectional graph linked list, which enables one-dimensional and flexible partitioning of DAG-structured DNNs. We then reformulate the problem of minimizing collaborative inference energy consumption and latency as a task-aware Markov decision process (MDP), which partitions each user's inference task queue into consecutive task windows for resource allocation. Building on this, we propose an Embedded Multi-Agent Hybrid Proximal Policy Optimization (EMH-PPO) algorithm to learn effective policies. Extensive experiments conducted across diverse network scenarios reveal that, compared to local DNN inference on MDs, our proposed method reduces inference latency by up to 64% and energy consumption by up to 46%.
Guanlei Zhang, Qiyang Zhang 0001, Lei Feng 0001, Fanqin Zhou, Praveen Kumar Donta, Schahram Dustdar
IEEE Trans. Mob. Comput.2
2026 Latency-Optimized Scheduling for Data Aggregation in Distributed Edge Computing
abstract
In Wireless Sensor Networks (WSNs), relay sensor nodes can aggregate data from edge sensor node into a summary information before sending to the sink. Due to the vast number of sensor nodes in a distributed edge computing (DEC) network, these relay sensor nodes may receive a high number of aggregation requests. This increases the chance of conflicting transmissions, which further leads to unwanted latency. Designing a conflict-free and minimal latency data aggregation schedule remains an open question. Moreover, existing related works have been conducted in traditional WSNs. By leveraging multiple antennas, the Multiple Input Multiple Output (MIMO) and cooperative MIMO called virtual MIMO (V-MIMO) enable broadband wireless communication, thereby improving the performance of WSNs. However, compared with traditional WSNs, MIMO and V-MIMO introduce distinct interference models requiring careful consideration. The work proposes a solution to an NP-hard problem, addressing three challenges: (i) interference; (ii) latency; and (iii) dynamic changes in network topology. Firstly, to counter interference, we propose a model where multiple nodes can simultaneously send data to the same parent by connecting different antennas. Secondly, to minimize latency, we propose a novel distributed heuristic data aggregation scheduling method, which intertwines the construction of an optimal data aggregation tree and conflict-free scheduling. Finally, to handle dynamic network topology changes, we propose lightweight adaptive strategies that do not increase data aggregation latency. Simulation results and theoretical analysis demonstrate superior performance in reducing data aggregation latency. When compared with state-of-the-art solutions, our proposed method decreases data aggregation latency by at least 2.6× on average.
Yunquan Gao, Qiyang Zhang 0001, Ying Li 0037, Praveen Kumar Donta, Lauri Lovén, Schahram Dustdar
ACM Trans. Internet Techn.2
2025 FOOL: Addressing the Downlink Bottleneck in Satellite Computing With Neural Feature Compression
abstract
Nanosatellite constellations equipped with sensors capturing large geographic regions provide unprecedented opportunities for Earth observation. As constellation sizes increase, network contention poses a downlink bottleneck. Orbital Edge Computing (OEC) leverages limited onboard compute resources to reduce transfer costs by processing the raw captures at the source. However, current solutions have limited practicability due to reliance on crude filtering methods or over-prioritizing particular downstream tasks. This work presents an OEC-native and task-agnostic feature compression method that preserves prediction performance and partitions high-resolution satellite imagery to maximize throughput. Further, it embeds context and leverages inter-tile dependencies to lower transfer costs with negligible overhead. While the encoding prioritizes features for downstream tasks, we can reliably recover images with competitive scores on quality measures at lower bitrates. We extensively evaluate transfer cost reduction by including the peculiarity of intermittently available network connections in low earth orbit. Finally, we test the feasibility of our system for standardized nanosatellite form factors. We demonstrate that the proposed approach permits downlinking over 100× the data volume without relying on prior information on the downstream tasks.
Alireza Furutanpey, Qiyang Zhang 0001, Philipp Raith, Tobias Pfandzelter, Shangguang Wang, Schahram Dustdar
IEEE Trans. Mob. Comput.2
2025 SatCooper: Enhancing Cooperative Inference Analytics for Satellite Service via Multi-Exit DNNs
abstract
As a key technology of intelligent satellite-enabled services in B5G or 6G networks, deploying Deep Neural Networks (DNN) models on satellites has been a notable trend, catering to the daily demand for extensive computing-intensive and latency-sensitive tasks. The computing resources are strategically deployed on satellites where sensor data is generated or collected, facilitating the fine-grained computational inference of DNN-based tasks. However, no prior study has comprehensively explored the crucial inference challenges – e.g., the trade-off between the number of tasks completed and accuracy and partitioning models in multi-exit models – in the resource-constrained space environment. Effective scheduling frameworks cater to various streams of inference tasks are scarce because inference performance may deviate from the ideal situation due to changes in task system status, such as task profiles and network state. To this end, we first formulate a gain-aware in-orbit computing inference problem to strike a proper trade-off between inference latency and the number of tasks completed by dynamically selecting optimal early exit points and model partitioning points. We propose an offline dynamic programming-based algorithm that provides an effective solution when comprehensive system details are to be predicted. We have developed an online learning-based method to schedule inference tasks with uncertain and dynamic system statuses in real-world situations. Our evaluation shows that, compared to baseline methods, the online learning-based algorithm can improve task gain by an average of 87.3% across various tasks.
Qiyang Zhang 0001, Shangguang Wang, Jinglong Guan, Praveen Kumar Donta, Xiao Ma 0009, R. Venkatesha Prasad, Schahram Dustdar, Xuanzhe Liu
IEEE Trans. Mob. Comput.1
2025 SLICE: Energy-Efficient Satellite-Ground Co-Inference via Layer-Wise Scheduling Optimization
abstract
Recent advancements in Low Earth Orbit (LEO) satellites are facilitating the provision of Deep Neural Networks (DNNs)-inherent services to achieve ubiquitous coverage via satellite computing. However, the computational demands and energy consumption of DNN models present significant challenges for satellite computing with limited power and computation resources. Based on the layered characteristics of DNN models, a satellite-ground co-inference strategy has been introduced, which executes certain layers on satellites and the remaining layers on ground servers. Determining the optimal layers for in-orbit processing, however, is non-trivial due to the under-explored energy consumption of satellite computing across different models and restricted yet varying communication conditions of satellite-ground links. In this paper, we first conduct a comprehensive measurement to uncover energy consumption of satellite computing across different layers and models. By summarizing the key observations, we develop a layer-specific energy consumption model tailored to diverse DNN architectures and kernels. We then investigate the energy-efficient satellite-ground co-inference problem and formulate it as an integer-nonlinear programming problem, which presents high computational complexity. To tackle these difficulties, we propose a satellite-ground co-inference algorithm that employs a branch-and-bound strategy, combined with the Sobol sequence and Lagrange multiplier, to reduce complexity and ensure stability across diverse DNN architectures. To evaluate the proposed algorithm, we conduct experiments based on real-world satellite parameters. The results demonstrate that our proposed algorithm can achieve an average energy savings of 96% under various data volumes compared to the existing benchmarks.
Qiyang Zhang 0001, Ruolin Xing, Yuanzhe Li 0001, Xiao Ma 0009, Ao Zhou 0001, Shangguang Wang
IEEE Trans. Serv. Comput.2
2025 An Efficient and Stable Knowledge Service Framework for Satellite-Ground Collaboration
abstract
The rapid expansion of Low Earth Orbit (LEO) satellite constellations presents immense potential for in-orbit services. However, the large-scale and dynamic nature of LEO constellations creates unstable communication environments, where traditional methods struggle to ensure the efficiency and stability of onboard inference services. This highlights the need for an advanced knowledge service framework capable of both inference and root cause analysis of service disruptions. To address this, we propose a novel knowledge service framework that integrates data-driven and knowledge-driven models through satellite-ground collaboration. The framework leverages lightweight onboard models for real-time data processing and ground-based knowledge graphs for advanced inference and cause analysis. To further enhance stability within complex interconnected onboard systems, we propose a prediction-based algorithm for LEO satellite networks that uses joint spatio-temporal modeling to achieve accurate link prediction. Additionally, we formulate an optimization problem aimed at minimizing path distance variance and maximizing path stability across LEO topologies, and we propose a heuristic path selection strategy to ensure efficient inter-satellite routing. Extensive in-orbit deployments and simulation experiments demonstrate the feasibility and effectiveness of the proposed framework. Satellite-ground verification on the BUPT-1 satellite shows its ability to provide real-time services, while inter-satellite simulations using real constellation data indicate significant improvements in response latency and path stability. Compared with baseline methods, our proposed method significantly reduces path jitter by up to 62.6% and improves path availability by up to 17.3% across various LEO constellations.
Fei Teng 0001, Qiyang Zhang 0001, R. Venkatesha Prasad, Schahram Dustdar
IEEE Trans. Serv. Comput.3
2025 FedCLR+: Tackling Onboard Label Constraints for Accurate Federated Satellite Computing
abstract
The rapid growth of Low Earth Orbit (LEO) satellites, particularly with the increasing deployment of intelligent computing capabilities using commercial off-the-shelf (COTS) hardware, presents significant opportunities to enhance the quality of in-orbit services. However, the current onboard conditions remain insufficient to enhance model accuracy by increasing model size, and inadequate accuracy hampers the effectiveness of in-orbit services. The satellite-ground federated learning (FL) paradigm, leveraging collaborative fine-tuning, offers a promising solution to continuously improve onboard model performance. Prior studies have focused on optimizing fine-tuning under constraints like limited bandwidth and computational resources, they often overlook two critical challenges: the scarcity and skewness of labeled onboard data and the long revisit cycles of satellites. To address these challenges and better support in-orbit services, this paper designs a realistic simulation methodology for the onboard fine-tuning process and conducts a comprehensive measurement study. Based on insights from the measurement results, we propose an efficient satellite-ground federated fine-tuning system,FedCLR+. In this system, we design a FedCLR algorithm to enhance system accuracy through representation optimization. Additionally, we propose a hybrid bias-compensated strategy to further mitigate accuracy loss by enriching the diversity of aggregation information. Experimental results show thatFedCLR+significantly enhances accuracy by up to 21.61×, reduces transmission volume by an average of 7.29%, and maintaining acceptable additional overhead compared to baselines.
Chen Yang 0043, Qiyang Zhang 0001, Qibo Sun, Shufeng Ouyang, Ao Zhou 0001, Shangguang Wang, Mengwei Xu 0001
IEEE Trans. Serv. Comput.2
2025 Energy-aware computing of access service for wireless edge via distributed deep learning
Xiaoyi Jiang 0004, Fanqin Zhou, Qiyang Zhang 0001, Yang Yang 0114, Daohua Zhu, Lei Feng 0001, Dayang Wang
Wirel. Networks3
2024 Collaborative Inference in DNN-Based Satellite Systems with Dynamic Task Streams
abstract
As a driving force in the advancement of intel-ligent in-orbit applications, DNN models have been gradually integrated into satellites, producing daily latency-constraint and computation-intensive tasks. However, the substantial computation capability of DNN models, coupled with the instability of the satellite-ground link, pose significant challenges, hindering the timely completion of tasks. It becomes necessary to adapt to task stream changes when dealing with tasks requiring latency guarantees, such as dynamic observation tasks on the satellites. To this end, we consider a system model for a collaborative inference system with latency constraints, leveraging the multi-exit and model partition technology. To address this, we propose an algorithm, which is tailored to effectively address the trade-off between task completion and maintaining satisfactory task accuracy by dynamically choosing early-exit and partition points. Simulation evaluations show that our proposed algorithm signif-icantly outperforms baseline algorithms across the task stream with strict latency constraints.
Jinglong Guan, Qiyang Zhang 0001, Ilir Murturi, Praveen Kumar Donta, Schahram Dustdar, Shangguang Wang
ICC2
2024 Resource-efficient In-orbit Detection of Earth Objects
abstract
With the rapid proliferation of large Low Earth Orbit (LEO) satellite constellations, a huge amount of in-orbit data is generated and needs to be transmitted to the ground for processing. However, traditional LEO satellite constellations, which downlink raw data to the ground, are significantly restricted in transmission capability. Orbital edge computing (OEC), which exploits the computation capacities of LEO satellites and processes the raw data in orbit, is envisioned as a promising solution to relieve the downlink burden. Yet, with OEC, the bottleneck is shifted to the inelastic computation capacities. The computational bottleneck arises from two primary challenges that existing satellite systems have not adequately addressed: the inability to process all captured images and the limited energy supply available for satellite operations. In this work, we seek to fully exploit the scarce satellite computation and communication resources to achieve satellite-ground collaboration and present a satellite-ground collaborative system named TargetFuse for onboard object detection. TargetFuse incorporates a combination of techniques to minimize detection errors under energy and bandwidth constraints. Extensive experiments show that TargetFuse can reduce detection errors by 3.4× on average, compared to onboard computing. TargetFuse achieves a 9.6× improvement in bandwidth efficiency compared to the vanilla baseline under the limited bandwidth budget constraint.
Qiyang Zhang 0001, Ruolin Xing, Zimu Zheng, Xiao Ma 0009, Mengwei Xu 0001, Schahram Dustdar, Shangguang Wang
INFOCOM1
2024 Energy-Aware Satellite-Ground Co-Inference via Layer-Wise Processing Schedule Optimization
abstract
Recent advancements in Low Earth Orbit (LEO) satellites are facilitating the provision of Deep Neural Networks (DNNs)-inherent services to achieve ubiquitous coverage via satellite computing. However, the computational demands and energy consumption of DNN models pose significant challenges for satellite computing with limited power and computation resources. Based on the hierarchical characteristics of DNN models, we propose a satellite-ground co-inference strategy that executing certain layers on satellites and the remaining layers on ground servers. However, identifying the optimal layers for in-orbit processing with latency constraints is challenging due to the uncertain energy consumption across diverse models. To explore the correlation between energy consumption and layer types, we conduct comprehensive measurements on a hardware device commonly found in commercial LEO satellites and develop a layer-based energy consumption prediction model. Then, we formulate an optimization problem of minimizing the energy consumption on the satellite within the latency constraint as an integer nonlinear programming problem. Solving this problem is difficult due to combinatorial explosion in the discrete solution space. To address this, we propose an improved algorithm based on genetic algorithms. Using configurations from a real satellite, we conduct simulation experiments, concluding that our algorithm significantly improves energy savings by an average of 27 ×.
Qiyang Zhang 0001, Ruolin Xing, Yuanzhe Li 0001, Xiao Ma 0009, Chaoxin Yu, Ao Zhou 0001, Shangguang Wang
Internetware2
2024 A Comprehensive Deep Learning Library Benchmark and Optimal Library Selection
abstract
Deploying deep learning (DL) on mobile devices has been a notable trend in recent years. To support fast inference of on-device DL, DL libraries play a critical role as algorithms and hardware do. Unfortunately, no prior work ever dives deep into the ecosystem of modern DL libraries and provides quantitative results on their performance. In this paper, we first build a comprehensive benchmark that includes 6 representative DL libraries and 15 diversified DL models. Then we perform extensive experiments on 10 mobile devices, and the results reveal the current landscape of mobile DL libraries. For example, we find that the best-performing DL library is severely fragmented across different models and hardware, and the gap between DL libraries can be rather huge. In fact, the impacts of DL libraries can overwhelm the optimizations from algorithms or hardware, e.g., model quantization and GPU/DSP-based heterogeneous computing. Motivated by the fragmented performance of DL libraries across models and hardware, we propose an effective DL Library selection framework to obtain the optimal library on a new dataset that has been created. We evaluate the DL Library selection algorithm, and the results show that the framework at it can improve the prediction accuracy by about 10% than benchmark approaches on average.
Qiyang Zhang 0001, Xiangying Che, Xiao Ma 0009, Mengwei Xu 0001, Schahram Dustdar, Xuanzhe Liu, Shangguang Wang
IEEE Trans. Mob. Comput.1
2024 Toward Efficient Satellite Computing Through Adaptive Compression
abstract
The rapid development of Low Earth Orbit (LEO) satellite constellations offers significant potential for in-orbit services, particularly in mitigating the impact of sudden natural disasters. However, the massive data collected by these satellites are often large and severely constrained by limited transmission capabilities when sending data to the ground. Satellite computing, which utilizes onboard computational capacity to process data before transmission, presents a promising solution to alleviate the downlink burden. Nonetheless, this paradigm introduces another bottleneck: limited onboard computing capacity, resulting in slow in-orbit processing and poor results. Current satellite computing systems struggle to efficiently address both data transmission and computing bottlenecks, particularly for urgent disaster services that demand accurate and timely results. Thus, we introduce an efficient satellite computing system designed to jointly mitigate these bottlenecks, thereby providing better service. The core idea is to utilize onboard computing capacity for swift in-orbit annotation of image regions, enabling adaptive compression and download based on annotation confidence and perceived downlink availability. Once the data is downloaded, image restoration and re-inference are performed on the ground to enhance accuracy. Compared to satellite-only inference, our system demonstrates an average improvement in inference accuracy of 3.8%. Furthermore, compared to ground-only inference, with only a 2.8% accuracy loss, our system achieves a 38.4% reduction in response time and saves 71.6% of downlink volume on average.
Chen Yang 0043, Qibo Sun, Qiyang Zhang 0001, Claudio A. Ardagna, Shangguang Wang, Mengwei Xu 0001
IEEE Trans. Serv. Comput.3
2023 Energy and Time-Aware Inference Offloading for DNN-based Applications in LEO Satellites
abstract
In recent years, Low Earth Orbit (LEO) satellites have witnessed rapid development, with inference based on Deep Neural Network (DNN) models emerging as the prevailing technology for remote sensing satellite image recognition. However, the substantial computation capability and energy demands of DNN models, coupled with the instability of the satellite-ground link, pose significant challenges, burdening satellites with limited power intake and hindering the timely completion of tasks. Existing approaches, such as transmitting all images to the ground for processing or executing DNN models on the satellite, is unable to effectively address this issue. By exploiting the internal hierarchical structure of DNNs and treating each layer as an independent subtask, we propose a satellite-ground collaborative computation partial offloading approach to address this challenge. We formulate the problem of minimizing the inference task execution time and onboard energy consumption through offloading as an integer linear programming (ILP) model. The complexity in solving the problem arises from the combinatorial explosion in the discrete solution space. To address this, we have designed an improved optimization algorithm based on branch and bound. Simulation results illustrate that, compared to the existing approaches, our algorithm improve the performance by 10%-18%.
Qiyang Zhang 0001, Xiao Ma 0009, Ao Zhou 0001
ICNP2
2022 A Comprehensive Benchmark of Deep Learning Libraries on Mobile Devices
abstract
Deploying deep learning (DL) on mobile devices has been a notable trend in recent years. To support fast inference of on-device DL, DL libraries play a critical role as algorithms and hardware do. Unfortunately, no prior work ever dives deep into the ecosystem of modern DL libs and provides quantitative results on their performance. In this paper, we first build a comprehensive benchmark that includes 6 representative DL libs and 15 diversified DL models. We then perform extensive experiments on 10 mobile devices, which help reveal a complete landscape of the current mobile DL libs ecosystem. For example, we find that the best-performing DL lib is severely fragmented across different models and hardware, and the gap between those DL libs can be rather huge. In fact, the impacts of DL libs can overwhelm the optimizations from algorithms or hardware, e.g., model quantization and GPU/DSP-based heterogeneous computing. Finally, atop the observations, we summarize practical implications to different roles in the DL lib ecosystem.
Qiyang Zhang 0001, Xiang Li 0067, Xiangying Che, Xiao Ma 0009, Ao Zhou 0001, Mengwei Xu 0001, Shangguang Wang, Yun Ma 0002, Xuanzhe Liu
WWW1