Chenyu Gong

dblp:297/0174 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2026 CLAIR: SLA-Aware Inference Routing in Converged Cloud-Network Systems
Mulei Ma, Qixuan Li, Tailiang Liu, Zeyun Du, Tian Huang, Chenyu Gong, Yang Yang 0001
INFOCOM7
2025 Multi-Tier Multi-Node Scheduling of LLM for Collaborative AI Computing
Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang 0001
INFOCOM2
2024 Multi-dimensional Quality of Experience for Customized User Requirements
abstract
Designing an everyone-centric customized service system is a crucial stage in the intelligent transformation of the digital world. As such, Quality of Experience (QoE) for user requirements design has become an essential research topic. Among the existing works, part of them assumes user requirements through ideal distributions. Another portion uses real-world data, but they analyze it from the system side. Both do not give a pervasive description of user requirements to support future research efforts. To tackle the above challenges, based on the Service Requirements Zone (SRZ), we propose an extended integrated multi-dimensional QoE, Acceptable Performance Zone (APZ), which includes both preferred and acceptable user requirements. We also detail the eight Key Performance Indicators (KPIs) in the SRZ. To provide more math support to the user requirements, we adopt real-world cluster trace datasets from Alibaba for analysis, aiming to explore the characteristics of real-world user requirements. The results show that the characteristics of user tasks generally obey some specific distributions. Specifically, the task size and memory requirement both follow bimodal log-Gaussian distributions, whereas the delay and computing requirements follow unimodal log-Gaussian distributions. At the same time, the energy consumption follows the log-beta distribution.
Yingzhi Liu, Mulei Ma, Chenyu Gong, Yang Yang 0001
APCC3
2024 MOGR: Multi-task Offloading via Graph Representation in Heterogeneous Computing Network
abstract
In the rapidly evolving field of heterogeneous computing networks, efficient task offloading plays a pivotal role in optimizing system throughput and resource utilization. However, existing task offloading methods often fall short of adequately modeling the dependency topology relationships between of-floaded tasks, which limits their effectiveness in capturing the complex interdependencies of task features. To address this limitation, we propose a framework named MOGR: Multi-task Offloading via Graph Representation. Our modeling approach takes into account factors such as task characteristics, network conditions, and available resources at the edge, and embeds these captured features into the graph structure. By utilizing Graph Convolutional Networks (GCN), our mechanism can capture and analyze the intricate relationships between task features, enabling a more comprehensive understanding of the underlying dependency topology. Through extensive evaluations in heteroge-neous networks, our proposed algorithm improves 15.1%-30.5% over greedy and approximate algorithms in optimizing system throughput and resource utilization. Our experiments showcase the advantage of considering the intricate interplay of task features using GCN-based modeling.
Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang 0001
ICC2
2024 FlocOff: Data Heterogeneity Resilient Federated Learning With Communication-Efficient Edge Offloading
abstract
Federated Learning (FL) has emerged as a fundamental learning paradigm to harness massive data scattered at geo-distributed edge devices in a privacy-preserving way. Given the heterogeneous deployment of edge devices, however, their data are usually Non-IID, introducing significant challenges to FL including degraded training accuracy, intensive communication costs, and high computing complexity. Towards that, traditional approaches typically utilize adaptive mechanisms, which may suffer from scalability issues, increased computational overhead, and limited adaptability to diverse edge environments. To address that, this paper instead leverages the observation that the computation offloading involves inherent functionalities such as node matching and service correlation to achieve data reshaping and proposesFederatedlearning basedoncomputingOffloading (FlocOff) framework, to address data heterogeneity and resource-constrained challenges. Specifically, FlocOff formulates the FL process with Non-IID data in edge scenarios and derives rigorous analysis on the impact of imbalanced data distribution. Based on this, FlocOff decouples the optimization in two steps, namely: 1) Minimizes the Kullback-Leibler (KL) divergence via Computation Offloading scheduling (MKL-CO); 2) Minimizes the Communication Cost through Resource Allocation (MCC-RA). Extensive experimental results demonstrate that the proposed FlocOff effectively improves model convergence and accuracy by 14.3%-32.7% while reducing data heterogeneity under various data distributions.
Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang 0001, Liantao Wu
IEEE J. Sel. Areas Commun.2
2023 FLIRRAS: Fast Learning With Integrated Reward and Reduced Action Space for Online Multitask Offloading
abstract
With the rapid development of edge data intelligence, task offloading (TO) and resource allocation (RA) optimization in multiaccess edge computing networks can significantly improve the Quality of Service (QoS). However, for the online scenario, traditional methods (e.g., game theory and numerical methods) cannot adapt to dynamic environments. Deep reinforcement learning (DRL) is applied to adjust the policy to get long-term rewards. Nevertheless, since the joint problem of TO and RA is nonconvex and NP-hard, existing DRL methods cannot guarantee high efficiency because of the large action space. To solve the above problem, we propose a fast learning with integrated reward and reduced action space-based DRL framework (FLIRRAS), which adopts a low-complexity approach to jointly optimize TO and RA strategies. The FLIRRAS framework combines DRL with numerical methods to iteratively pursues the discrete TO and continuous RA. Specifically, a deep neural network (DNN) is used to learn environmental information, which can get prior knowledge of the offloading decision. Furthermore, a novel reward integrating the utility of TO and RA is designed to motivate the agent to find the optimal policy. To solve the dilemma that the action space is too large, low-complexity convex optimization methods, i.e., subgradient projection and KKT condition, are used to supplement and adjust the decision, which reduces the network parameters and the decision space. In addition, given the dynamic online environment, we introduce the experience replay mechanism, where policy is updated regularly to reflect the best mapping between states. The experiment results show that the performance of FLIRRAS is better than greedy and other DRL approaches, and it outperforms the latest DRL method by over 18.0% in terms of execution time.
Mulei Ma, Chenyu Gong, Liantao Wu, Yang Yang 0001
IEEE Internet Things J.2
2022 Task Offloading and Resource Allocation in CPU-GPU Heterogeneous Networks
abstract
With the massive use of GPU, task scheduling under CPU-GPU clusters has become an indispensable research topic. Unlike existing models, we propose an innovative framework that users offload their tasks in CPU-GPU heterogeneous Edge Clusters (ECs) instead of general-purpose CPU clusters. The framework takes full advantage of the GPU's powerful parallel computing capabilities. Specifically, we decompose each user task into sequential segments and parallel segments, which can be offloaded to CPUs and GPUs of the ECs, respectively. By dis-cretizing the GPU's computing capability, we formulate a Mixed Integer Nonlinear Programming (MINLP), which involves jointly optimizing the task offloading decision, the uplink transmission power of users, and computing resource allocation. To tackle this challenging problem, we propose a Joint Simulated Annealing and Convex Optimization (JSAC) based algorithm to minimize the total overhead consisting of delay and energy consumption. Our experimental simulation results demonstrate that the JSAC algorithm can make full use of GPU's powerful parallel computing capability via allocating GPU resources effectively. In particular, the JSAC algorithm achieves optimal performance in terms of system overhead, number of beneficial UEs, and speedup.
Chenyu Gong, Mulei Ma, Liantao Wu, Yong Zhou 0006, Yang Yang 0001
GLOBECOM1
2022 UCL: Unit Competition of Layers for Streaming Tasks in Heterogeneous Networks
abstract
Partitioning and offloading the deep neural network (DNN) model over multi-tier computing units have been recently proposed to shorten the inference time. However, the state-of-the-art cannot adapt to large-scale offloading problems for streaming tasks because of its exponential complexity. Besides, as an essential kind of DNNs, the offloading of grouped con-volutional neural networks (GCNNs) has not been explored yet. Motivated by the above facts, in this paper, we concentrate on the offloading of chained DNNs (CDNNs) and GCNNs for streaming tasks. Consider a typical heterogeneous network consisting of various computing units, the user equipment (UE) publishes computation-intensive and delay-sensitive streaming DNN tasks while computing units accomplish them collaboratively. To mini-mize the delay of processing the task stream, DNN layers should be offloaded to appropriate units, which is the streaming-task multi-unit (STMU) problem. To tackle this problem, we formulate a non-cooperative potential game called unit competition of layers (UCL). The theoretical analysis proves the existence of the Nash equilibrium (NE), and the corresponding algorithm with linear complexity is developed to achieve the NE. Finally, extensive experiments demonstrate that UCL outperforms the state-of-the-art significantly in large-scale scenarios while maintaining similar performance on small-scale tasks.
Liantao Wu, Guoliang Gao, Chenyu Gong
GLOBECOM4