Wenkai Lv

dblp:319/6970 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-2276-053XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CHIME: Cost-Constrained Hybrid Popularity-Aware Intelligent Service Caching Framework for MEC
Tianyang Zheng, Pengfei Yang 0001, Chenlu Zhai, Wenkai Lv, Yueli Ding, Quan Wang 0006
IEEE Internet Things J.5
2025 Alternating optimization for energy consumption-oriented task offloading in SAGIN
Pengfei Yang 0001, Tianyang Zheng, Weidi Su, Bijie Yi, Wenkai Lv, Quan Wang 0006
Comput. Networks6
2025 PipeMCTS: Pipeline Inference Optimization for Edge Computing via Surrogate Model-Guided MCTS
abstract
For deep neural network (DNN) deployment on Internet of Things (IoT) devices, pipeline-based inference coordinates diverse computing resources in heterogeneous multiprocessor system-on-chips (HMPSoCs) to achieve efficient execution. However, determining the optimal pipeline configuration is challenging due to the exponentially expanding search space and intricate layer-wise dependencies. Existing methods formulate this as a single-shot optimization problem, which struggles to efficiently explore the search space and incurs substantial resource overhead from repeated evaluations of suboptimal configurations. This article proposes PipeMCTS, which reformulates pipeline deployment as a sequential optimization problem solved via Monte Carlo tree search (MCTS). By incrementally constructing the search tree in a layer-wise manner, PipeMCTS effectively prunes unpromising branches and accelerates the search process. PipeMCTS incorporates two key components to enhance search efficiency: 1) a temperature-controlled selection strategy that balances exploration and exploitation in the complex search space and 2) an uncertainty-aware simulation strategy that accelerates search convergence through Gaussian process (GP)-guided evaluation. Experimental results demonstrate that PipeMCTS achieves a 66.42% improvement in throughput compared to arm compute library (ARM-CL).
Jianjun Ding, Dan Xian, Yusheng Jin, Dejun Hua, Wenbin Hua, Wenkai Lv, Chengmin Lin
IEEE Internet Things J.7
2025 Cortex: Enhancing Resource Utilization in Edge Clusters Through Efficient Co-Location of LC and BE Workloads
abstract
In edge computing environments, the co-location of latency-critical (LC) services and best-effort (BE) jobs is a key strategy for enhancing resource utilization. However, existing analysis-based co-location strategies incur high analytical costs and struggle to rapidly adapt to the evolving fields of edge computing and microservice architectures, often failing to effectively meet the demands of edge computing environments. Feedback-based co-location strategies, while reducing analytical overhead, lack sufficient research in multi-node environments, resulting in overly coarse-grained deployment strategies. These strategies do not adequately consider the dynamic workloads and resource constraints inherent in edge computing, leading to improper resource allocation and degraded performance. This paper introduces Cortex, a Kubernetes-based co-location framework for edge device clusters that addresses these challenges by innovatively transforming the co-location deployment problem into a Minimum Cost Maximum Flow (MCMF) problem and employing the Network Simplex Algorithm (NSA) to optimize resource allocation and ensure QoS of LC services. Cortex also features a dynamic adjustment mechanism that adapts to changes in the request load of LC services, thereby minimizing the performance loss of BE jobs and reducing resource wastage. Our experiments in real edge device clusters demonstrate that Cortex significantly improves system resource utilization by 12.81%, increases the QoS satisfaction rate by 17.86%, and boosts the number of BE jobs by 51.96% compared to existing methods.
Tianyang Zheng, Pengfei Yang 0001, Quan Wang 0006, Wenkai Lv
IEEE Internet Things J.6
2024 Graph-Reinforcement-Learning-Based Dependency-Aware Microservice Deployment in Edge Computing
abstract
Microservice architecture is a design philosophy that achieves decoupling by decomposing a monolithic application into multiple lightweight microservices. Meanwhile, edge computing can significantly reduce service latency and network congestion by extending computation and storage resources to the network edge. Therefore, in the microservice-oriented edge computing platform, a fundamental problem is how to efficiently deploy microservices with complex dependencies on the resource-constrained edge servers to satisfy the Quality of Service (QoS) constraints of users. Most of the existing studies ignore multiple call graphs with differentiated dependencies for an application, which often result in the violation of QoS. To address this issue, in this article, we first model the request response time of multiple instances and multiple call graphs scenario with service conflicts. Then, different from the existing heuristic or approximation algorithms which rely heavily on expert knowledge, we propose a graph-reinforcement-learning-based deployment (GRLD) framework. GRLD uses a graph convolutional network (GCN) to extract the graph data required for multiple call graphs with messages passing and aggregation, and the generated feature is fed into the underlying network of deep-reinforcement-learning (DRL). Experimental results show that GRLD outperforms counterparts in reducing service deployment overhead while satisfying QoS constraints of multiple call graphs.
Wenkai Lv, Pengfei Yang 0001, Tianyang Zheng, Chengmin Lin, Minwen Deng, Quan Wang 0006
IEEE Internet Things J.1
2024 Performance Prediction for Deep Learning Models With Pipeline Inference Strategy
abstract
For Heterogeneous Multi-Processor System-on-Chips (HMPSoCs), a reasonable pipeline design can significantly improve the inference performance of Deep Learning (DL) models. The pipeline design optimization can be modeled as a search problem where an accurate prediction model can efficiently speed up the search process. However, the performance prediction of DL models for the pipeline inference strategy is challenging because of the inter-layer effect, inference details, and variety of model structures. In this paper, we propose TPPNet, a transformer-based model for predicting the inference performance of various DL models with the pipeline inference strategy. TPPNet represents the DL model as an execution sequence with operators and hardware details to extract the hidden factors between layers. Moreover, we apply the Multi-task Learning (MTL) method to accurately predict throughput and latency metrics by constructing a predictive model. To the best of our knowledge, this is the first study dedicated to pipeline inference performance prediction for the DL model on HMPSoCs. We evaluate TPPNet on six well-known DL models using RK3399. The experimental outcomes affirm the high accuracy of TPPNet and its capability to significantly reduce the time overhead associated with pipeline exploration.
Pengfei Yang 0001, Linwei Hu, Wenkai Lv, Chengmin Lin, Quan Wang 0006
IEEE Internet Things J.5
2024 Fine-grained complexity-driven latency predictor in hardware-aware neural architecture search using composite loss
Chengmin Lin, Pengfei Yang 0001, Wenkai Lv, Quan Wang 0006
Inf. Sci.5
2024 Flexi-BOPI: Flexible granularity pipeline inference with Bayesian optimization for deep learning models on HMPSoC
Pengfei Yang 0001, Linwei Hu, Wenkai Lv, Chengmin Lin, Quan Wang 0006
Inf. Sci.5
2024 SLAPP: Subgraph-level attention-based performance prediction for deep learning models
Pengfei Yang 0001, Linwei Hu, Chengmin Lin, Wenkai Lv, Quan Wang 0006
Neural Networks6
2023 Energy Consumption and QoS-Aware Co-Offloading for Vehicular Edge Computing
abstract
By deploying computing, storage, and bandwidth resources at the user side, vehicular edge computing (VEC) provides low-delay services for vehicle users. However, due to the limited resources of edge servers, how to efficiently meet the Quality-of-Service (QoS) requirements of multiple tasks and save the total energy consumption in a dynamic environment is an important issue in VEC. In this article, we first propose an energy consumption and QoS-aware co-offloading model. Unlike most previous studies, our goal is to minimize the total energy consumption while guaranteeing the QoS constraints of tasks, thus avoiding the overallocation of resources and high energy consumption caused by the one-sided pursuit of delay minimization. Then, without the requirements for domain experts, we propose Bayesian optimization-based computation offloading (BOCO) method to find the optimal offloading decision. To the best of our knowledge, this work is the first to apply Bayesian optimization to computation offloading in VEC. Furthermore, we conduct a series of experiments and comparisons with other offloading methods to analyze the effectiveness and performance of the proposed algorithm. Experimental results verify that our proposed BOCO outperforms counterparts.
Wenkai Lv, Pengfei Yang 0001, Tianyang Zheng, Bijie Yi, Yunqing Ding, Quan Wang 0006, Minwen Deng
IEEE Internet Things J.1
2023 Efficient and accurate compound scaling for convolutional neural networks
Chengmin Lin, Pengfei Yang 0001, Quan Wang 0006, Zeyu Qiu, Wenkai Lv
Neural Networks5
2022 Microservice Deployment in Edge Computing Based on Deep Q Learning
abstract
The microservice deployment strategy is promising in reducing the overall service response time in the microservice-oriented edge computing platform. However, existing works ignore the effect of different interaction frequencies among microservices and the decrease in service execution performance caused by the increased node loads. In this article, we first model the invocation relationships among microservices as an undirected and weighted interaction graph to characterize the communication overhead. Then, we propose a multi-objective microservice deployment problem (MMDP) in edge computing. MMDP aims to minimize the communication overhead while achieving load balance between edge nodes. Without the requirement for domain experts, we propose Reward Sharing Deep Q Learning (RSDQL), a learning-based algorithm, to solve MMDP and obtain the optimal deployment strategy. In addition, to improve the scalability of the services, we propose an Elastic Scaling algorithm (ES) based on heuristics to deal with the dynamic pressure of requests. Finally, we conduct a series of experiments in Kubernetes to evaluate the performance of our approach. Experimental results indicate that, compared with interaction-aware strategy and Kubernetes default strategy, RSDQL has shorter response times, more balanced resource loads, and makes services scale elastically according to the request pressure.
Wenkai Lv, Quan Wang 0006, Pengfei Yang 0001, Yunqing Ding, Bijie Yi, Chengmin Lin
IEEE Trans. Parallel Distributed Syst.1