Heng Zhang 0032

dblp:55/826-32 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-4874-6162ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 CSQoS: Continual Sparse QoS Measurement for Edge Clouds With GNN-Based Variational Bayesian
abstract
Edge computing, an emerging paradigm, utilizes decentralized edge nodes to offer low-latency, high-quality network services. Quality of Service (QoS) is a crucial metric to measure network service quality, and network resource scheduling relies on QoS measurement results. However, current QoS measurement methods often measure all QoS data among edge nodes, these dense measurement approaches introduce significant costs. Besides, edge nodes may adopt varied network access manners, these factors cause fluctuations in QoS between edge nodes. But existing QoS measurement works often focus on measuring exact QoS values while merely considering QoS fluctuations, resulting in unreliable measured QoS data. In addition, existing QoS measurement methods often can not support online QoS measurement, leading to stale offline QoS data affecting network resource scheduling. To tackle the two issues, we propose a novel Continual Sparse QoS range measurement method (CSQoS) with four innovative designs: (1) To reduce measurement costs, we propose to measure QoS by only sampling partial QoS data and using them to impute unmeasured QoS data. To achieve sparse QoS imputation, we propose a novel variational Bayesian model (BayGNN) with an edge-enhanced Graph Neural Network (GNN) as the encoder for feature extraction and a Multilayer Perceptron (MLP) as the decoder to predict unmeasured QoS data. (2) To assess QoS data ranges, we design the proposed BayGNN model to produce uncertainty simultaneously. (3) To fulfill reliable online QoS predictions, we incorporate continual learning and residual connections in BayGNN. Experimental results on 2 real-world datasets demonstrate that CSQoS has minimal QoS imputation error with the lowest measurement costs, reducing 17.6% RMSE and 20% sampling costs.
Heng Zhang 0032, Liping Yi, Xiaofei Wang 0001
IEEE Internet Things J.1
2026 DynEformer: A Unified Framework for Robust Workload Prediction Under Dynamic Environment
abstract
Workload prediction in multi-tenant edge cloud platforms (MT-ECP) is crucial for efficient application deployment and resource provisioning. However, the heterogeneous application patterns, variable infrastructure performance, and frequent deployments in MT-ECP pose significant challenges for accurate prediction. Existing clustering-based methods often incur excessive costs due to maintaining multiple data clusters and models, while end-to-end time-series prediction methods struggle with dynamic environments. To address these challenges, we perform a comprehensive analysis on a large-scale workload dataset in real-world MT-ECP and propose DynEformer, an end to-end framework with global pooling and static context aware ness, offering a unified workload prediction scheme for dynamic MT-ECP. Meticulously designed global pooling and information merging mechanisms can effectively identify and utilize global application patterns to drive local workload predictions. The integration of static content-aware mechanisms enhances model robustness in real-world scenarios. We also extend DynEformer's capabilities to Long-term workload forecasting (LTLF) and Long-period service (LPS) tasks. Experiments on six real-world datasets demonstrate that DynEformer achieves state-of-the-art performance, with a 32% relative improvement on nine baselines and a 52% improvement in application switching and new entity scenarios. Additional experiments on long-term prediction and online learning further confirm its effectiveness for LTLF and LPS tasks.
Shaoyuan Huang, Zheng Wang 0001, Heng Zhang 0032, Xiaofei Wang 0001, Cheng Zhang 0019
IEEE Trans. Knowl. Data Eng.3
2026 pFedLoRA: Model-Heterogeneous Personalized Federated Learning With Homogeneous Low-Rank Adapter Sharing on Mobile Edge Devices
abstract
Federated learning (FL) is an emerging machine learning paradigm in which a central server coordinates multiple participants (FL clients) collaboratively to train on decentralized data. In practice, FL often faces data, system, and model heterogeneity, which inspires the field of Model-Heterogeneous Personalized Federated Learning (MHPFL). However, existing MHPFL methods rely on extra public data or ignore the relationship between local private heterogeneous models and shared homogeneous models across clients. This leads to unsatisfactory model performance, computational overheads, and communication costs. To bridge this gap, we propose a novel and efficient model-heterogeneouspersonalizedFederated learning framework (pFedLoRA) based on sharing homogeneous Low-Rank Adapter (LoRA) which is popular for fine-tuning pre-trained models. Specifically, we devise a lightweight homogeneous adapter, rather than apply the typical LoRA, to facilitate each client's heterogeneous local model training with our proposed iterative training for global-local bidirectional knowledge exchange. The homogeneous small local adapters are aggregated on the FL server to generate a global adapter. We theoretically prove its$\mathcal {O}(1/T)$non-convex convergence rate. Experiments on 5 datasets demonstratepFedLoRAoutperforms 9 state-of-the-art baselines in model accuracy with$11.81 \times$computation and$7.41\times$communication cost saving.
Liping Yi, Heng Zhang 0032, Han Yu 0001, Gang Wang 0001, Xiaoguang Liu 0001, Qinghua Hu
IEEE Trans. Mob. Comput.2
2024 Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
abstract
As live streaming services skyrocket, Crowdsourced Cloud-edge service Platforms (CCPs) have surfaced as pivotal intermediaries catering to the mounting demand. Despite the role of stream scheduling to CCPs’ Quality of Service (QoS) and throughput, conventional optimization strategies struggle to enhancing CCPs’ revenue, primarily due to the intricate relationship between resource utilization and revenue. Additionally, the substantial scale of CCPs magnifies the difficulties of time-intensive scheduling. To tackle these challenges, we propose Seer, a proactive revenue-aware scheduling system for live streaming services in CCPs. The design of Seer is motivated by meticulous measurements of real-world CCPs environments, which allows us to achieve accurate revenue modeling and overcome three key obstacles that hinder the integration of prediction and optimal scheduling. Utilizing an innovative Preschedule-Execute-Re-schedule paradigm and flexible scheduling modes, Seer achieves efficient revenue-optimized scheduling in CCPs. Extensive evaluations demonstrate Seer’s superiority over competitors in terms of revenue, utilization, and anomaly penalty mitigation, boosting CCPs revenue by 147% and expediting scheduling 3.4× faster.
Shaoyuan Huang, Zheng Wang 0001, Zhongtian Zhang, Heng Zhang 0032, Xiaofei Wang 0001
INFOCOM4
2024 QM-RGNN: An Efficient Online QoS Measurement Framework with Sparse Matrix Imputation for Distributed Edge Clouds
abstract
Measurements for the quality of end-to-end network services (QoS) are crucial to ensure stability, reliability, and user experience for distributed edge clouds. Measuring all QoS data brings significant costs. Existing QoS measurement methods attempt to use sparse measured QoS data to estimate unmeasured QoS data. But they suffer from limited estimation accuracy when facing QoS data with high sparsity or significant volatility. Moreover, they also consume high sampling and training costs during continuously online measurements. Our preliminary analysis reveals that end-to-end QoS is strongly temporal-spatial related. It inspires us to leverage partially measured QoS data to impute temporal-spatial-related unmeasured QoS data for reducing measurement costs. To predict unmeasured QoS data precisely with low computational costs, we propose a novel QoS Measurement framework based on Residual Graph Neural Network (QM-RGNN), which inputs QoS data as a graph and outputs the prediction of unmeasured QoS data. It consists of three core components: 1) an encoder-decoder model QM-GNN with GCN as the encoder and MLP as the decoder is devised for efficient QoS prediction, and a residual module is introduced in QM-GNN to tackle highly sparse and volatile QoS data; 2) a dynamic adaptive sample ratio is proposed to reduce the sampling costs; 3) an online learning pattern is designed to reduce continuous training costs. Experiments on two real-world industrial edge cloud datasets demonstrate the superiority of QM-RGNN in QoS measurement. It obtains at least a 37.5% reduction of relative RMSE between ground-truth and predicted QoS data with up to 90% training cost reduction and 22.7% sampling cost reduction.
Heng Zhang 0032, Zixuan Cui, Shaoyuan Huang, Deke Guo, Xiaofei Wang 0001
INFOCOM1
2024 Large-Scale Measurements and Optimizations on Latency in Edge Clouds
abstract
The emergence of next-generation latency-critical applications places strict requirements on network latency and stability. Edge cloud, an instantiated paradigm for edge computing, is gaining more and more attention due to its benefits of low latency. In this work, we make an in-depth investigation into the network QoS, especially end-to-end latency, at both spatial and temporal dimensions on a nationwide edge computing platform. Through the measurements, we collect a multi-variable large-scale real-world dataset on latency. We then quantify how the spatial-temporal factors affect the end-to-end latency, and verify the predictability of end-to-end latency. The results reveal the limitation of centralized clouds and illustrate how could edge clouds provide low and stable latency. Our results also point out that existing edge clouds merely increase the density of servers and ignore spatial-temporal factors, so they still suffer from high latency and fluctuations. Based on a quantified latency impact factor, we have proposed several optimization strategies for edge cloud latency and validated their effectiveness. We also propose a robust prototype edge cloud model based on lessons we learn from the measurement and evaluate its performance in the production environment. Evaluation result shows that edge clouds achieve 84.1% latency reduction with 0.5 ms latency fluctuation and 73.3% QoS improvement compared with the centralized clouds.
Heng Zhang 0032, Shaoyuan Huang, Mengwei Xu 0001, Deke Guo, Xiaofei Wang 0001, Xin Wang 0030, Victor C. M. Leung
IEEE Trans. Cloud Comput.1
2024 Fine-Grained Spatio-Temporal Distribution Prediction of Mobile Content Delivery in 5G Ultra-Dense Networks
abstract
The 5G networks have extensively promoted the growth of mobile users and novel applications, and with the skyrocketing user requests for a large amount of popular content, the consequent content delivery services (CDSs) have been bringing a heavy load to mobile service providers. As a key mission in intelligent networks management, understanding and predicting the distribution of CDSs benefits many tasks of modern network services such as resource provisioning and proactive content caching for content delivery networks. However, the revolutions in novel ubiquitous network architectures led by ultra-dense networks (UDNs) make the task extremely challenging. Specifically, conventional methods face the challenges of insufficient spatio precision, lacking generalizability, and complex multi-feature dependencies of user requests, making their effectiveness unreliable in CDSs prediction under 5G UDNs. In this article, we propose to adopt a series of encoding and sampling methods to model CDSs of known and unknown areas at a tailored fine-grained level. Moreover, we design a spatio-temporal-social multi-feature extraction framework for CDSs hotspots prediction, in which a novel edge-enhanced graph convolution block is proposed to encode dynamic CDSs networks based on the social relationships and the spatio features. Besides, we introduce the Long-Short Term Memory (LSTM) to further capture the temporal dependency. Extensive performance evaluations with real-world measurement data collected in two mobile content applications demonstrate the effectiveness of our proposed solution, which can improve the prediction area under the curve (AUC) by 40.5% compared to the state-of-the-art proposals at a spatio granularity of 76m, with up to 80% of the unknown areas.
Shaoyuan Huang, Heng Zhang 0032, Xiaofei Wang 0001, Min Chen 0003, Jianxin Li 0001, Victor C. M. Leung
IEEE Trans. Mob. Comput.2
2023 How Far Have Edge Clouds Gone? A Spatial-Temporal Analysis of Edge Network Latency In the Wild
abstract
The emergence of next-generation latency-critical applications places strict requirements on network latency and stability. Edge cloud, an instantiated paradigm for edge computing, is gaining more and more attention due to its benefits of low latency. In this work, we make an in-depth investigation into the network QoS, especially end-to-end latency, at both spatial and temporal dimensions on a nationwide edge computing platform. Through the measurements, we collect a multi-variable large-scale real-world dataset on latency. We then quantify how the spatial-temporal factors affect the end-to-end latency, and verified the predictability of end-to-end latency. The results reveal the limitation of centralized clouds and illustrate how could edge clouds provide low and stable latency. Our results also point out that existing edge clouds merely increase the density of servers and ignore spatial-temporal factors, so they still suffer from high latency and fluctuations. Based on the observations, we propose a robust prototype edge cloud model based on lessons we learn from the measurement and evaluate its performance in the production environment. The further evaluation result shows that edge clouds achieve 84.1% latency reduction with 0.5ms latency fluctuation and 73.3% QoS improvement compared with the centralized clouds.
Heng Zhang 0032, Shaoyuan Huang, Mengwei Xu 0001, Deke Guo, Xiaofei Wang 0001, Victor C. M. Leung
IWQoS1
2023 One for All: Unified Workload Prediction for Dynamic Multi-tenant Edge Cloud Platforms
abstract
Workload prediction in multi-tenant edge cloud platforms (MT-ECP) is vital for efficient application deployment and resource provisioning. However, the heterogeneous application patterns, variable infrastructure performance, and frequent deployments in MT-ECP pose significant challenges for accurate and efficient workload prediction. Clustering-based methods for dynamic MT-ECP modeling often incur excessive costs due to the need to maintain numerous data clusters and models, which leads to excessive costs. Existing end-to-end time series prediction methods are challenging to provide consistent prediction performance in dynamic MT-ECP. In this paper, we propose an end-to-end framework with global pooling and static content awareness, DynEformer, to provide a unified workload prediction scheme for dynamic MT-ECP. Meticulously designed global pooling and information merging mechanisms can effectively identify and utilize global application patterns to drive local workload predictions. The integration of static content-aware mechanisms enhances model robustness in real-world scenarios. Through experiments on five real-world datasets, DynEformer achieved state-of-the-art in the dynamic scene of MT-ECP and provided a unified end-to-end prediction scheme for MT-ECP.
Shaoyuan Huang, Zheng Wang 0001, Heng Zhang 0032, Xiaofei Wang 0001, Cheng Zhang 0007
KDD3
2023 A Measurement-Driven Analysis and Prediction of Content Propagation in the Device-to-Device Social Networks
abstract
In the 5 G era, data traffic has been growing rapidly. A small number of popular data files may dominate the network traffic and lead to heavy network congestion. Device-to-Device (D2D) communication can be used for caching and offloading significant data traffic. D2D social networks are instantiated paradigms of D2D communication. Existing studies maximize the performances of caching and offloading in D2D social networks by predicting potential content propagation paths. However, predicting such paths still faces many challenges, such as limitation of user spatial-temporal features, fragility of D2D social networks, and uncertainty of participants. As a solution, we first measure users' multi-dimensional features and content propagation paths to explore the distributions of D2D activities. Then we propose a D2D-LSTM model to predict complete content propagation paths hierarchically and design a prototype-user model for new participants. Experimental results demonstrate the state-of-the-art performances of D2D-LSTM. D2D-LSTM achieves at most 95% and at least 84.6% average precision in predicting terminal prototype-user class. Tree generation tests show that the generated trees have at most 64% and at least 17% similarity with ground-truth trees.
Heng Zhang 0032, Shaoyuan Huang, Xin Wang 0030, Jianxin Li 0001, Xiaofei Wang 0001, Victor C. M. Leung
IEEE Trans. Knowl. Data Eng.1
2022 Multitask Offloading Strategy Optimization Based on Directed Acyclic Graphs for Edge Computing
abstract
With the advancement of the user application service demands, the IoT system tends to offload the tasks to the edge server for execution. Most of the current studies on edge computation offloading ignore the dependencies between components of the application. The few pieces of research on edge computing offloading which focus on the topology of application are primarily applied in single-user scenarios. Unlike previous work, our work mainly solves dependent task offloading with edge computing in multiuser scenarios, which is more in line with reality. In this article, the dependent task offloading problem is modeled as a Markov decision process (MDP) first. Then, we propose an actor–critic mechanism with two embedding layers for directed acyclic graphs (DAGs)-based multiple dependent tasks computation offloading, namely, ACED, by jointly considering the topology of the application and the channel interference between several users. Finally, the results of simulations also show the priorities of the proposed ACED algorithm.
Yajun Yang, Chenyang Wang 0001, Heng Zhang 0032, Chao Qiu, Xiaofei Wang 0001
IEEE Internet Things J.4
2021 Spatio-Temporal-Social Multi-Feature-based Fine-Grained Hot Spots Prediction for Content Delivery Services in 5G Era
abstract
The arrival of 5G networks has extensively promoted the growth of content delivery services (CDSs). Understanding and predicting the spatio-temporal distribution of CDSs are beneficial to mobile users, Internet Content Providers and carriers. Conventional methods for predicting the spatio-temporal distribution of CDSs are mostly base-stations (BSs) centric, leading to weak generalization and spatio coarse-grained. To improve the spatio accuracy and generalization of modeling, we propose user-centric methods for CDSs spatio-temporal analysis. With geocoding and spatio-temporal graphs modeling algorithms, CDSs records collected from mobile devices are modeled as dynamic graphs with spatio-temporal attributes. Moreover, we propose a spatio-temporal-social multi-feature extraction framework for spatio fine-grained CDSs hot spots prediction. Specifically, an edge-enhanced graph convolutional block is designed to encode CDSs information based on the social relations and the spatio dependence features. Besides, we introduce the Long Short Term Memory (LSTM) to further capture the temporal dependence. Experiments on two real-world CDSs datasets verified the effectiveness of the proposed framework, and ablation studies are taken to evaluate the importance of each feature.
Shaoyuan Huang, Heng Zhang 0032, Xiaofei Wang 0001, Min Chen 0003, Jianxin Li 0001, Victor C. M. Leung
CIKM2
2020 D2D-LSTM: LSTM-Based Path Prediction of Content Diffusion Tree in Device-to-Device Social Networks
abstract
With the proliferation of mobile device users, the Device-to-Device (D2D) communication has ascended to the spotlight in social network for users to share and exchange enormous data. Different from classic online social network (OSN) like Twitter and Facebook, each single data file to be shared in the D2D social network is often very large in data size, e.g., video, image or document. Sometimes, a small number of interesting data files may dominate the network traffic, and lead to heavy network congestion. To reduce the traffic congestion and design effective caching strategy, it is highly desirable to investigate how the data files are propagated in offline D2D social network and derive the diffusion model that fits to the new form of social network. However, existing works mainly concern about link prediction, which cannot predict the overall diffusion path when network topology is unknown. In this article, we propose D2D-LSTM based on Long Short-Term Memory (LSTM), which aims to predict complete content propagation paths in D2D social network. Taking the current user's time, geography and category preference into account, historical features of the previous path can be captured as well. It utilizes prototype users for prediction so as to achieve a better generalization ability. To the best of our knowledge, it is the first attempt to use real world large-scale dataset of mobile social network (MSN) to predict propagation path trees in a top-down order. Experimental results corroborate that the proposed algorithm can achieve superior prediction performance than state-of-the-art approaches. Furthermore, D2D-LSTM can achieve 95% average precision for terminal class and 17% accuracy for tree path hit.
Heng Zhang 0032, Xiaofei Wang 0001, Chenyang Wang 0001, Jianxin Li 0001
AAAI1