EDBT 2026 Demo / reviewers in the wild / expert
Mengfei Zhu
dblp:284/3197
· DBLP profile ↗
14ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0002-8338-4219ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 8 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reflex: A Bi-Modal Failure Recovery Mechanism for Clusters under Control-Plane DegradationabstractModern clusters rely on centralized control planes, but decisions can be slow and fragile under control-plane degradations. We present Reflex, a bi-modal recovery design for clusters under control-plane degradation. Reflex adds a reflex-arc-like path that makes rapid takeover decisions from preprocessed local priorities while suppressing contention via lightweight coordination. After services are runnable, it performs steady-state reconstruction. Our evaluation shows bounded latency and robust conflict suppression under control-plane degradation and bursty failures. Mengfei Zhu, Rui Kang 0002, Jiuxiang Zhu, Tong Li 0014 |
APNet | 1 |
| 2026 | Robust Prewarming Orchestration for Multi-Model Elastic Inference under Traffic Uncertainty
Mengfei Zhu, Rui Kang 0002, Tong Li 0014 |
INFOCOM | 1 |
| 2026 | POSTER: ActShare: Coordinating Reusable Action Requests in Multi-Agent WorkflowsabstractMulti-agent workflows often generate repeated action requests when specialized agents interact with shared runtime objects. When requests target the same object under explicit compatibility conditions, a single execution result can serve multiple agents. This paper presents ActShare, a pre-execution coordination layer for reusable action requests. Before invoking a tool or a model, each agent emits a schema-constrained structured action request. ActShare compares the request against an active action table and a recent result cache. A request is attached only to a compatible owner action; otherwise, it is executed independently. When the owner action completes, ActShare writes the result once and dispatches it to all attached requests. We prototype ActShare in a stateful multi-agent debugging workflow and evaluate effect on repeated action reduction, token and tool-call savings. Rui Kang 0002, Mengfei Zhu, Tong Li 0014 |
SIGCOMM | 2 |
| 2026 | POSTER: CLEX: Contract-Bounded Local Execution for Device-Level Network ControlabstractModern networks increasingly suffer from gray failures that are too fine-grained and short-lived for global re-optimization, while existing local mechanisms lack the context and control boundaries needed for effective bounded response. We propose CLEX, a contract-bounded local execution framework in which the controller defines per-device action boundaries and the device-side execution agent combines local semantics, runtime signals, and short-lived memory to select bounded local policy states that are realized through the local execution substrate. Implementation shows that CLEX enables agile local responses while reducing unnecessary action oscillation. Mengfei Zhu, Rui Kang 0002, Tong Li 0014 |
SIGCOMM | 1 |
| 2023 | Robust function deployment against uncertain recovery time in different protection types with workload-dependent failure probability
Mengfei Zhu, Eiji Oki |
Comput. Networks | 1 |
| 2022 | Robust Function Deployment against Uncertain Recovery Time with Workload-Dependent Failure ProbabilityabstractThis paper proposes a robust function deployment model against uncertain recovery time with satisfying an expected recovery time guarantee in a cost-efficient manner. We consider that each node fails with a workload-dependent failure probability, which is a non-decreasing function that reveals the empirical relationship between the workload and the failure probability. The preventively deployed backup resources can recover an unavailable function hosted by a failed node in a period of time, which is related to the backup strategies and failure and recovery scenarios. We introduce an uncertainty set that considers the upper and lower bounds of the recovery time of a function by each node that protects it and the upper bound of the average recovery time among nodes. The robust optimization technique is applied to handle the worst case of expected recovery time satisfying a time guarantee under an uncertain recovery time. With this technique, the model is formulated as a mixed integer linear programming problem. The numerical results reveal that the proposed model saves the deployment cost on average 24% compared to a baseline that uses the deterministic recovery time in our tested cases. Mengfei Zhu, Fujun He, Eiji Oki |
CCNC | 1 |
| 2022 | Implementation of Real-time Function Deployment with Resource Migration in KubernetesabstractPrompt function deployment and management is a key role in network function virtualization to improve the continuity and reliability of network services. Kubernetes is a system to deploy and manage functions automatically. Existing tools in Kubernetes do not provide automatic function deployment and management in a real-time and optimal manner. It does not provide a resource type to manage the migratable resource, either. This paper designs and implements a two-layer controller structure in Kubernetes to achieve the function deployment in a limited computation time with considering resource migration for allocation optimality. A controller in the lower layer manages the Pods for an intermediate allocation with a model or a heuristic algorithm to respond to requests promptly. A controller in the upper layer manages instances by optimizing resource allocations with considering resource migration; it maintains the Pods by keeping the current state (intermediate allocation) consistent with the desired state (optimal allocation). Our demonstration validates that the controller automatically manages the resources promptly and correctly. Mengfei Zhu, Rui Kang 0002, Eiji Oki |
NOMS | 1 |
| 2022 | Optimization Model for Primary and Backup Resource Allocation With Workload-Dependent Failure ProbabilityabstractThis paper proposes an optimization model to derive a primary and backup resource allocation considering a workload-dependent failure probability to minimize the maximum expected unavailable time (MEUT). The workload-dependent failure probability is a non-decreasing function which reveals the relationship between the workload and the failure probability. The proposed model adopts hot backup and cold backup strategies to provide protection. The cold backup strategy is a protection strategy, in which the requested loads of backup resources are not activated before failures occur to reduce resource utilization with the cost of longer recovery time. The hot backup strategy is a protection strategy, in which the backup resources are activated and synchronized with the primary resources to recover promptly with the cost of higher workload. We formulate the optimization problem as a mixed integer linear programming (MILP) problem. We prove that MEUT of the proposed model is equal to the smaller value between the two MEUTs obtained by applying only hot backup and cold backup strategies with the same total requested load. A heuristic algorithm inspired by the water-filling algorithm is developed with the proved theorem. The numerical results show that the proposed model suppresses MEUT compared with the conventional model which does not consider the workload-dependent failure probability. The developed heuristic algorithm is approximately 105times faster than the MILP approach with 10−2performance penalty on MEUT. Mengfei Zhu, Fujun He, Eiji Oki |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2022 | Resource Allocation Model Against Multiple Failures With Workload-Dependent Failure ProbabilityabstractFault tolerance and load balancing are two key roles in resource allocation against failures. This paper proposes a primary and backup resource allocation model with preventive recovery priority setting to minimize a weighted value of unavailable probability (W-UP) against multiple failures. W-UP considers the probability of unsuccessful recovery and the maximum unavailable probability after recovery among physical nodes. We consider that each node fails with a workload-dependent failure probability; each failure pattern occurs with a probability. The workload-dependent failure probability is a non-decreasing function revealing an empirical relationship between the workload and the failure probability for each physical node. We introduce a recovery strategy to handle the workload variation which is determined at the operation start time and can be applied for each failure pattern. Once a failure pattern occurs, the recoveries are operated according to the priority setting to promptly recover the functions hosted by failed nodes. We also discuss an approach to obtain unsuccessful recovery probability with considering the maximum number of arbitrary recoverable functions by a set of available nodes without the priority setting. We formulate the optimization problem as a mixed integer linear programming (MILP) problem. We develop a heuristic algorithm to solve larger size problems in a practical time. The developed heuristic algorithm is approximately 729 times faster than the MILP approach with 1.6% performance penalty on W-UP. The numerical results observe that the proposed model reduces W-UP compared with baselines. Mengfei Zhu, Fujun He, Eiji Oki |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Implementation of Backup Resource Management Controller for Reliable Function Allocation in KubernetesabstractResource allocation and management is a key role in network function virtualization to improve the reliability of network services. Kubernetes is a system to deploy and manage the virtual network functions automatically. Existing tools in Kubernetes does not provide a resource type to define the backup Pods. It does not provide automatic resource management based on the user requests for the backup Pods, either. This paper designs and implements a custom resource and the corresponding controller in Kubernetes to manage the primary and backup resources of network functions. The custom resource is a set of Pods with different types, which includes primary, hot backup, and cold backup Pods. The controller manages the set of Pods and maintains the current state of the different types of Pods to keep the current state consistent with the desired state of each type of Pod. Demonstration validates that the controller automatically manage the primary and backups resources correctly. Mengfei Zhu, Rui Kang 0002, Fujun He, Eiji Oki |
NetSoft | 1 |
| 2021 | Implementation of Virtual Network Function Allocation with Diversity and Redundancy in KubernetesabstractDiversity in network function virtualization is to use a group of thin replicas to provide the network services under the required processing ability. Redundancy is to provide a certain number of replicas against function failures and improve network reliability. Kubernetes is a system to deploy and manage virtual network functions automatically. Existing tools in Kubernetes do not provide a resource type to provide required functions jointly considering VNF diversity and redundancy. This paper designs and implements a custom resource and the corresponding controller in Kubernetes to manage the VNF diversity and redundancy jointly. The controller selects suitable replicas from a pool of replica templates to satisfy the required processing ability with the minimum required number of replicas and converts the backup functions to the primary functions when the primary functions cannot provide the required ability. Demonstration validates that the controller automatically manages the resources correctly, improves the resource utilization, and increases the number of acceptable requests. Rui Kang 0002, Mengfei Zhu, Fujun He, Eiji Oki |
Networking | 2 |
| 2021 | Optimization Model for Multiple Backup Resource Allocation With Workload-Dependent Failure ProbabilityabstractThis paper proposes a multiple backup resource allocation model with a workload-dependent failure probability to minimize the maximum expected unavailable time (MEUT) under a protection priority policy. The workload-dependent failure probability is a non-decreasing function which reveals the relationship between the workload and the failure probability. The proposed model adopts hot backup and cold backup strategies to provide protection. For protection of each function with multiple backup resources, it is required to adopt a suitable priority policy to determine the expected unavailable time. We analyze the superiority of the protection priority policy for multiple backup resources in the proposed model; we provide the theorems that clarify the influence of policies on MEUT. We formulate the optimization problem as a mixed integer linear programming (MILP) problem. We provide a lower bound of the optimal objective value in the proposed model. We prove that the decision version of the multiple resource allocation problem in the proposed model is NP-complete. A heuristic algorithm inspired by the water-filling algorithm is developed with providing an upper bound of the expected unavailable time obtained by the algorithm. The numerical results show that the proposed model reduces MEUT compared to baselines. The priority policy adopted in the proposed model suppresses MEUT compared with other priority policies. The developed heuristic algorithm is approximately 106times faster than the MILP approach with 10-4performance penalty on MEUT. Mengfei Zhu, Fujun He, Eiji Oki |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Fault-Tolerant Control for Dynamic Positioning Vessel With Thruster Faults Based on the Neural Modified Extended State ObserverabstractA new fault-tolerant control (FTC) method based on the neural modified extended state observer (NMESO) is proposed for dynamic positioning (DP) vessel with thruster faults in this article. Through incorporating a compound orthogonal neural network (CONN) into the design process of the modified extended state observer (MESO), the NMESO is developed to estimate the uncertainties in the DP control system, such as the environmental disturbances and the unknown dynamics, as well as the thruster faults simultaneously without knowing any prior information of them. With the help of the accurate estimation of the total uncertainties by NMESO, a PD-like feedback controller is established to realize the FTC of the DP vessel toward the thruster faults. By utilizing the Lyapunov stability analysis, it is proved that all the error signals in the closed-loop cascade system formed by the NMESO and PD-like feedback controller are uniformly ultimately bounded (UUB) and the bounds could be arbitrarily small by choosing appropriate parameters. Simulation experiments on two typical thruster fault scenarios are carried out to validate the effectiveness and the performance of the proposed NMESO-FTC compared with the conventional ESO-FTC. The simulation results show the proposed approach has better fault-tolerant performance. Wenzhao Yu, Haixiang Xu, Yahao Chen, Mengfei Zhu |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2020 | Multiple Backup Resource Allocation with Workload-Dependent Failure ProbabilityabstractThis paper proposes a multiple backup resource allocation model with a workload-dependent failure probability to minimize the maximum expected unavailable time (MEUT) under a protection priority policy. The workload-dependent failure probability is a monotonically increasing function which reveals the relationship between computing workload and failure probability. The proposed model adopts hot backup and cold backup strategies to provide protection. The cold backup strategy is a protection strategy, in which the requested loads of backup resources are not processed as active workloads before failures occur to reduce resource utilization with the cost of long recover time. The hot backup strategy is a protection strategy, in which the backup resources execute at the same time with functions to recover promptly with the cost of high workload. For protection of each function with multiple backup resources, it is required to adopt a suitable priority policy to determine the expected unavailable time. We analyze the superiority of the protection priority policy for multiple backup resources in the proposed model and provide the theorems that clarify the influence of policies on MEUT. The numerical results show that the proposed model reduces MEUT compared with the single backup model in which each function is protected by only one server without protection priority of servers. The priority policy adopted in the proposed model specifying that the server which adopts the hot backup strategy has higher priority than that with the cold backup strategy for multiple backup resources suppresses MEUT compared with other priority policies. Mengfei Zhu, Fujun He, Eiji Oki |
GLOBECOM | 1 |