VLDB 2026 Research / reviewers in the wild / expert
Mulei Ma
dblp:320/8373
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0002-6032-1418ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 6 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLAIR: SLA-Aware Inference Routing in Converged Cloud-Network Systems
Mulei Ma, Qixuan Li, Tailiang Liu, Zeyun Du, Tian Huang, Chenyu Gong, Yang Yang 0001 |
INFOCOM | 1 |
| 2025 | Multi-Tier Multi-Node Scheduling of LLM for Collaborative AI Computing
Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang 0001 |
INFOCOM | 1 |
| 2024 | Multi-dimensional Quality of Experience for Customized User RequirementsabstractDesigning an everyone-centric customized service system is a crucial stage in the intelligent transformation of the digital world. As such, Quality of Experience (QoE) for user requirements design has become an essential research topic. Among the existing works, part of them assumes user requirements through ideal distributions. Another portion uses real-world data, but they analyze it from the system side. Both do not give a pervasive description of user requirements to support future research efforts. To tackle the above challenges, based on the Service Requirements Zone (SRZ), we propose an extended integrated multi-dimensional QoE, Acceptable Performance Zone (APZ), which includes both preferred and acceptable user requirements. We also detail the eight Key Performance Indicators (KPIs) in the SRZ. To provide more math support to the user requirements, we adopt real-world cluster trace datasets from Alibaba for analysis, aiming to explore the characteristics of real-world user requirements. The results show that the characteristics of user tasks generally obey some specific distributions. Specifically, the task size and memory requirement both follow bimodal log-Gaussian distributions, whereas the delay and computing requirements follow unimodal log-Gaussian distributions. At the same time, the energy consumption follows the log-beta distribution. Yingzhi Liu, Mulei Ma, Chenyu Gong, Yang Yang 0001 |
APCC | 2 |
| 2024 | MOGR: Multi-task Offloading via Graph Representation in Heterogeneous Computing NetworkabstractIn the rapidly evolving field of heterogeneous computing networks, efficient task offloading plays a pivotal role in optimizing system throughput and resource utilization. However, existing task offloading methods often fall short of adequately modeling the dependency topology relationships between of-floaded tasks, which limits their effectiveness in capturing the complex interdependencies of task features. To address this limitation, we propose a framework named MOGR: Multi-task Offloading via Graph Representation. Our modeling approach takes into account factors such as task characteristics, network conditions, and available resources at the edge, and embeds these captured features into the graph structure. By utilizing Graph Convolutional Networks (GCN), our mechanism can capture and analyze the intricate relationships between task features, enabling a more comprehensive understanding of the underlying dependency topology. Through extensive evaluations in heteroge-neous networks, our proposed algorithm improves 15.1%-30.5% over greedy and approximate algorithms in optimizing system throughput and resource utilization. Our experiments showcase the advantage of considering the intricate interplay of task features using GCN-based modeling. Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang 0001 |
ICC | 1 |
| 2024 | FlocOff: Data Heterogeneity Resilient Federated Learning With Communication-Efficient Edge OffloadingabstractFederated Learning (FL) has emerged as a fundamental learning paradigm to harness massive data scattered at geo-distributed edge devices in a privacy-preserving way. Given the heterogeneous deployment of edge devices, however, their data are usually Non-IID, introducing significant challenges to FL including degraded training accuracy, intensive communication costs, and high computing complexity. Towards that, traditional approaches typically utilize adaptive mechanisms, which may suffer from scalability issues, increased computational overhead, and limited adaptability to diverse edge environments. To address that, this paper instead leverages the observation that the computation offloading involves inherent functionalities such as node matching and service correlation to achieve data reshaping and proposesFederatedlearning basedoncomputingOffloading (FlocOff) framework, to address data heterogeneity and resource-constrained challenges. Specifically, FlocOff formulates the FL process with Non-IID data in edge scenarios and derives rigorous analysis on the impact of imbalanced data distribution. Based on this, FlocOff decouples the optimization in two steps, namely: 1) Minimizes the Kullback-Leibler (KL) divergence via Computation Offloading scheduling (MKL-CO); 2) Minimizes the Communication Cost through Resource Allocation (MCC-RA). Extensive experimental results demonstrate that the proposed FlocOff effectively improves model convergence and accuracy by 14.3%-32.7% while reducing data heterogeneity under various data distributions. Mulei Ma, Chenyu Gong, Liekang Zeng, Yang Yang 0001, Liantao Wu |
IEEE J. Sel. Areas Commun. | 1 |
| 2023 | FLIRRAS: Fast Learning With Integrated Reward and Reduced Action Space for Online Multitask OffloadingabstractWith the rapid development of edge data intelligence, task offloading (TO) and resource allocation (RA) optimization in multiaccess edge computing networks can significantly improve the Quality of Service (QoS). However, for the online scenario, traditional methods (e.g., game theory and numerical methods) cannot adapt to dynamic environments. Deep reinforcement learning (DRL) is applied to adjust the policy to get long-term rewards. Nevertheless, since the joint problem of TO and RA is nonconvex and NP-hard, existing DRL methods cannot guarantee high efficiency because of the large action space. To solve the above problem, we propose a fast learning with integrated reward and reduced action space-based DRL framework (FLIRRAS), which adopts a low-complexity approach to jointly optimize TO and RA strategies. The FLIRRAS framework combines DRL with numerical methods to iteratively pursues the discrete TO and continuous RA. Specifically, a deep neural network (DNN) is used to learn environmental information, which can get prior knowledge of the offloading decision. Furthermore, a novel reward integrating the utility of TO and RA is designed to motivate the agent to find the optimal policy. To solve the dilemma that the action space is too large, low-complexity convex optimization methods, i.e., subgradient projection and KKT condition, are used to supplement and adjust the decision, which reduces the network parameters and the decision space. In addition, given the dynamic online environment, we introduce the experience replay mechanism, where policy is updated regularly to reflect the best mapping between states. The experiment results show that the performance of FLIRRAS is better than greedy and other DRL approaches, and it outperforms the latest DRL method by over 18.0% in terms of execution time. Mulei Ma, Chenyu Gong, Liantao Wu, Yang Yang 0001 |
IEEE Internet Things J. | 1 |
| 2022 | Task Offloading and Resource Allocation in CPU-GPU Heterogeneous NetworksabstractWith the massive use of GPU, task scheduling under CPU-GPU clusters has become an indispensable research topic. Unlike existing models, we propose an innovative framework that users offload their tasks in CPU-GPU heterogeneous Edge Clusters (ECs) instead of general-purpose CPU clusters. The framework takes full advantage of the GPU's powerful parallel computing capabilities. Specifically, we decompose each user task into sequential segments and parallel segments, which can be offloaded to CPUs and GPUs of the ECs, respectively. By dis-cretizing the GPU's computing capability, we formulate a Mixed Integer Nonlinear Programming (MINLP), which involves jointly optimizing the task offloading decision, the uplink transmission power of users, and computing resource allocation. To tackle this challenging problem, we propose a Joint Simulated Annealing and Convex Optimization (JSAC) based algorithm to minimize the total overhead consisting of delay and energy consumption. Our experimental simulation results demonstrate that the JSAC algorithm can make full use of GPU's powerful parallel computing capability via allocating GPU resources effectively. In particular, the JSAC algorithm achieves optimal performance in terms of system overhead, number of beneficial UEs, and speedup. Chenyu Gong, Mulei Ma, Liantao Wu, Yong Zhou 0006, Yang Yang 0001 |
GLOBECOM | 2 |
| 2022 | Data-aware Hierarchical Federated Learning via Task OffloadingabstractTo cope with the high communication overhead caused by frequent aggregation of Federated Learning (FL) in Multi-access Edge Computing (MEC) scenarios, Hierarchical Federated Edge Learning (HFEL) is proposed as an evolving framework. HFEL offloads tasks to edge servers for partial model aggregation to reduce network traffic. However, most of the existing research focuses on resource optimization for HFEL without considering the impact of data characteristics and cannot guarantee the quality of FL training. To this end, we propose a task offloading approach based on data and resource heterogeneity under HFEL to improve training performance and reduce system cost. Specifically, we leverage information entropy to incorporate data statistical features into the cost function to reshape edge datasets. In addition, we applied Multi-Agent Deep Deterministic Policy Gradient (MADDPG) with a resource allocation module to generate distributed offloading policy more efficiently. Our algorithm not only adopts local observations to obtain the optimal action but also takes into account device heterogeneity, which can adapt to the unstable edge environment. Extensive experiments under multiple datasets and baselines are carried out, which demonstrate that our algorithm can effectively improve the accuracy of aggregated models while reducing system cost. Mulei Ma, Liantao Wu, Nanxi Chen, Ziyu Shao, Yang Yang 0001 |
GLOBECOM | 1 |