VLDB 2026 Research / reviewers in the wild / expert
Lin Cui 0001
dblp:27/1570-1
· DBLP profile ↗
63ranked-venue papers
7as first author
42since 2021 · last 2026
0000-0001-7961-3261ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 41 · 4 first-author · 26 since 2021Systems, architecture and hardware · 15 · 3 first-author · 11 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPRINT: Line-Rate In-band Network Telemetry Recovery for Application Optimization
Bingzhen Chen, WaiMing Lau, Xiaoquan Zhang, Fung Po Tso 0001, Lin Cui 0001 |
INFOCOM | 5 |
| 2026 | Monic: In-Network Mixture-of-Experts Inference on Programmable Data Planes
Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001, Zhen Zhang 0017, Kaimin Wei, Weijia Jia 0001, Lin Cui 0001 |
INFOCOM | 8 |
| 2026 | FaasOrc: A bi-level function scheduling and caching framework for serverless edge computingabstractServerless computing, underpinned by an event-driven approach with transient stateless containers, significantly enhances resource efficiency and simplifies function development. To maintain an acceptable Quality of Service (QoS) agreed in the Service Level Agreement (SLA), service providers need to improve the response latency while considering resource efficiency. However, cold-start delays in container initialization often lead to considerable latency in these applications. Existing mitigation strategies, such as pre-warming and function caching, are inadequate due to workload skewness and oscillation across edge nodes. These limitations are particularly critical in resource-constrained edge environments. We must jointly consider multiple factors, such as node status, function resource requirement and function popularity. To overcome these limitations, this paper presents FaasOrc , a bi-level function orchestration framework to mitigate workload skewness and oscillation across edge nodes. FaasOrc uses a cluster-level scheduler to schedule requests and a node-level manager to detect popular functions. Our comprehensive evaluation, consisting of two parts: simulations and a real-system prototype over Knative , benchmarks the proposed solution against existing scheduling and caching strategies. The findings highlight our method’s capability to reduce the response latency by 31.4%. Chen Chen 0073, Lars Nagel 0001, Lin Cui 0001, Weijia Jia 0001, Fung Po Tso 0001 |
J. Netw. Comput. Appl. | 3 |
| 2026 | dVRM: Cross-Switch Memory Sharing and Self-Adaptive Allocation in Distributed Data PlaneabstractProgrammable switches have revolutionized networking by enabling a new spectrum of applications, such as network telemetry, in-network computation, and machine learning. These applications heavily utilize register memory but their performance is significantly constrained by the scarcity of on-chip resources, such as the 15 MB of SRAM available on a Tofino switch. To effectively accommodate increasingly memorydemanding applications, we aim to pool register resources across multiple switches, creating a larger unified register memory space. This resource pooling approach addresses the limitations of existing single-switch Virtual Register Memory (VRM) solutions, which cannot meet the demands of these applications in distributed environments. To achieve this, we propose dVRM, a distributed VRM deployment framework that enables crossswitch memory sharing and self-dynamic memory allocation on the data plane. dVRM introduces three innovations: (1) Grouped Multi-Switch Registers (GMRs), virtualizing distributed pipeline stages into a unified memory pool; (2) a self-adaptive, bit-width allocation mechanism driven by real-time data-plane feedback; and (3) lightweight heuristics for concurrent application deployment with distributed VRM, formulated as a mixed-integer linear programming (MILP) problem. We have implemented dVRM on both P4 hardware switches (with Intel Tofino ASIC) and BMv2. Experimental results show that dVRM significantly reduces hash unit consumption by up to 26% and achieves an improvement in accuracy (ARE) of up to 57.3% across various workloads. Mimi Qian, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Zhen Zhang 0017, Weijia Jia 0001 |
IEEE Trans. Computers | 2 |
| 2026 | Communication-Efficient Federated Learning by Exploiting Spatio-Temporal Correlations of GradientsabstractCommunication overhead is a critical challenge in federated learning, particularly in bandwidth-constrained networks. Although many methods have been proposed to reduce communication overhead, most focus solely on compressing individual gradients, overlooking the temporal correlations among them. Prior studies have shown that gradients exhibit spatial correlations, typically reflected in low-rank structures. Through empirical analysis, we further observe a strong temporal correlation between client gradients across adjacent rounds. Based on these observations, we propose GradESTC, a compression technique that exploits both spatial and temporal gradient correlations. GradESTC exploits spatial correlations to decompose each full gradient into a compact set of basis vectors and corresponding combination coefficients. By exploiting temporal correlations, only a small portion of the basis vectorsneed to be dynamically updated in each round. GradESTC significantly reduces communication overhead by transmitting lightweight combination coefficients and a limited number of updated basis vectors instead of the full gradients. Extensive experiments show that, upon reaching a target accuracy level near convergence, GradESTC reduces uplink communication by an average of 39.79% compared to the strongest baseline, while maintaining comparable convergence speed and final accuracy to uncompressed FedAvg. By effectively leveraging spatio-temporal gradient structures, GradESTC offers a practical and scalable solution for communication-efficient federated learning. Shenlong Zheng, Zhen Zhang 0017, Yuhui Deng 0001, Geyong Min, Lin Cui 0001 |
IEEE Trans. Computers | 5 |
| 2026 | AGCB: Adaptive Garbage Collection for Enhancing Lifetime and Performance of Bit-Alterable Flash MemoryabstractBit-alterable flash-based SSDs, offering page-level erase operation, allows individual flash pages in a block to be erased independently. The page-level erase operation alleviates the overhead of page migration during garbage collection and improves the SSD lifetime. However, when the number of invalid pages within a block exceeds a certain threshold, the latency of page-level garbage collections using page-level erase may exceed that of block-level garbage collections. In bit-alterable flash memory, existing garbage collection strategies dynamically choose between page-level and block-level garbage collections based on their latency. This often fails to fully exploit the advantage of page-level garbage collection in reducing write amplification under low-load conditions.To address this limitation, we propose an adaptive garbage collection strategy called AGCB to dynamically adjust garbage collection operations by the runtime workload of flash channels, thereby enhancing SSD performance and lifetime. Specifically, AGCB classifies flash channels as busy or idle by monitoring the depth of the transaction queue in cache. According to this classification, AGCB selectively applies page-level or block-level garbage collection operations, aiming to minimize the impact of garbage collections with host I/O requests. Meanwhile, we introduce a staged victim block selection scheme to further improve garbage collection efficiency and wear leveling. The experimental results unveil that compared with the existing schemes, AGCB reduces the number of garbage collection operations, average response time, and blocked user requests by an average of 14.6%, 14.7%, and 17.3%, respectively. Laifu Zhang, Yuhui Deng 0001, Peng Zhou 0032, Shujie Pang, Zhaorui Wu, Lin Cui 0001, Zhen Zhang 0017 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | PAAP: A Graph-Based VNF Deployment Framework for Embedding Bidirectional SFC in Mobile Edge NetworksabstractIn the realm of Mobile Edge Computing (MEC) networks, mobile interactive applications, such as multiplayer online games, federated learning, and interactive multi-view video, are becoming increasingly popular. The embedding of Bidirectional Service Function Chains (BSFCs) for these applications has been studied. However, existing BSFC embedding research only considers static users, and the prevailing Virtual Network Functions (VNFs) placement methods do not account for the impact of individual node resources on the path, thus failing to maximize node resource utilization at the network level. Considering the aforementioned issues, we introduce a framework for BSFC embedding tailored for mobile users, named Path as a Point (PAAP). This framework integrates the resources of the paths and then globally considers the impact of the computing resources of edge nodes on the paths connecting them, maximizing edge node resource utilization across the entire network. We propose a three-phase algorithm within the PAAP framework. In the first phase, we introduce a singleuser algorithm for VNFs deployment during BSFC embedding. In the second phase, this optimization is extended to multiple users by establishing inter-user association rules. In the third phase, candidate application placement positions are evaluated across the entire network, culminating in the determination of optimal placement strategies. Empirical evaluations confirm the effectiveness of the proposed algorithm, demonstrating significant performance improvements over baseline methods. Dehui Ou, Zhen Zhang 0017, Yuhui Deng 0001, Geyong Min, Lin Cui 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | P-UCB: A Preference-Based Upper Confidence Bound Strategy for Efficient Edge Server PlacementabstractMobile Edge Computing (MEC) is an emerging network architecture designed to enhance Quality of Service (QoS) by bringing computational resources and storage closer to end users. In MEC environments, strategic placement of edge servers is crucial to minimize costs and optimize QoS. The Edge Server Placement (ESP) problem has been effectively addressed by using Multi-Armed Bandit (MAB) algorithms, known for their efficiency and adaptability, with the Upper Confidence Bound (UCB) algorithm being particularly prominent due to its stable performance and low dependency on parameters. However, UCB suffers excessive exploration and high uncertainty in reward estimation during its initial phase, leading to slow convergence. To overcome these challenges, we propose a novel Preference-based UCB (P-UCB) algorithm, which integrates the preference function into the UCB framework, drawing inspiration from the Gradient Bandit (GB) method. This modification not only accelerates convergence, but also improves overall efficiency. Furthermore, to address the Base Station Allocation (BSA) issue within the ESP context, we introduce a Weighted Base Station Allocation (WBSA) algorithm, which helps better manage access delay and workload balance. The P-UCB algorithm is evaluated through a comprehensive metric that based on access delay and workload balance, showing significant improvements over existing methods such as Multiple Choice (MC)-UCB, Q- Particle Swarm Optimization (QPSO), and Genetic Algorithm (GA). Experimental results on a real-world dataset show that P-UCB achieves a notable performance increase of at least 12.9%, effectively optimizing access delay and workload balance across various experimental settings, including the number of base stations, edge servers, and other relevant system parameters. Dongjiong Zhu, Zhen Zhang 0017, Yuhui Deng 0001, Shun Long, Lin Cui 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | PMPHD: A High Performance Virtual Machine Consolidation Strategy Based on Dynamic Threshold AdjustmentabstractVirtual machine (VM) consolidation strategies are widely deployed in Cloud Data Centers (CDCs) to optimize resource utilization and improve the Quality of Service (QoS). However, the host overload detection algorithms in current VM consolidation strategies are static. That means, once the overload threshold is calculated, it will not change until the next recalculation. The current algorithms are not suitable for the environment of highly dynamic workloads which results in additional energy consumption and potential Service Level Agreement Violations (SLAVs) which will affect the QoS of CDC. In PMPHD, a novel host dynamic threshold adjustment algorithm is proposed. In the proposed algorithm, the PMs are classified into mildly overloaded, normal, and severely overloaded based on the resource utilization. If the PM is predicted to be severely overloaded in the next moment, the threshold of this PM will be proactively reduced. The PM is determined to be overloaded, and some VMs in this PM will be migrated in advance. Thus, this PM will be in normal in the next moment, and the VM performance degradation resulting from SLAV and VM migration overlap in the next moment will be avoided. If the PM is predicted to be mildly overloaded, the threshold will be appropriately increased to transit it to be in normal state in the next moment, and the VM in the PM will not be migrated. Since the PMs’ workloads are dynamic, the PMPHD overload algorithm predicts the resource utilization rate of PM continuously, and adjusts the overload threshold of PM. Compared with other algorithms, PMPHD maintains high efficiency while having lower ESV (a combination metric for balancing energy consumption and SLAV). Zhen Zhang 0017, Zhenyu He 0003, Yuhui Deng 0001, Shenlong Zheng, Dongjiong Zhu, Lin Cui 0001 |
IEEE Trans. Netw. Serv. Manag. | 8 |
| 2026 | sketchPro: Identifying Top-k Items Based on Probabilistic Update on Programmable Data PlaneabstractDetecting the top-k heaviest items in network traffic is fundamental to traffic engineering, congestion control, and security analytics. Controller-side solutions suffer from high communication latency and heavy resource overhead, motivating the migration of this task to programmable data planes (PDP). However, PDP hardware (e.g., Tofino ASIC) offers only a few megabytes of on-chip SRAM per pipeline stage and supports neither loops nor complex arithmetic, making accurate top-k detection highly challenging. This paper proposes sketchPro, a novel sketch-based solution that employs a probabilistic update scheme to retain large items, enabling accurate top-k identification on PDP with minimal memory. sketchPro dynamically adjusts the probability of updates based on the current statistical size of the items and the frequency of hash collisions, thus allowing sketchPro to effectively detect top-k items. We have implemented sketchPro on PDP, including P4 software switch (i.e., BMv2) and hardware switch (Intel Tofino ASIC). Extensive evaluation results demonstrate that sketchPro can achieve more than 95% precision with only 10KB of memory. Keke Zheng, Mai Zhang, Mimi Qian, Waiming Lau, Lin Cui 0001 |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2026 | Latency-Sensitive and Resource-Efficient Parallel VNF Placement in Mobile Edge Networks: A Dynamic Graph Weighting Approach
Zhen Zhang 0017, Yuhui Deng 0001, Geyong Min, Lin Cui 0001 |
IEEE Trans. Netw. | 6 |
| 2025 | Planner: A Generative Graph Learning Framework for Noisy and Dynamic In-band Network TelemetryabstractIn-band network telemetry (INT) enables real-time network monitoring by embedding telemetry data into packets. The advent of programmable switches further enhances the flexibility of INT by enabling dynamic customization of telemetry collection at the hardware level. However, the practical application of INT is hampered by significant challenges arising from data noise (due to packet loss, delay, and measurement inaccuracies) and network dynamics (such as changing INT paths and feature requirements). These issues severely degrade the performance of machine learning models used for analyzing INT data, hindering the accurate capture of spatio-temporal network characteristics. This paper presents Planner, a novel generative graph learning framework designed to address these limitations. Planner enables the collection of network features at various levels of granularity on programmable switches. Crucially, it constructs dynamic graphs representing evolving INT paths and employs a hybrid Graph Neural Network (GNN) and Recurrent Neural Network (RNN) architecture to effectively learn spatial and temporal dependencies. Furthermore, Planner incorporates variational inference to generate robust latent representations, mitigating the detrimental effects of noise and instability in INT data. We have implemented a testbed prototype of Planner using Intel Tofino ASIC switches. Extensive experiments demonstrate the performance superiority and robustness of Planner over the baseline methods, achieving a 23.2% improvement in F1 score. Xiaoquan Zhang, Waiming Lau, Lin Cui 0001, Fung Po Tso 0001, Zhuoqian Liang, Zhen Zhang 0017, Yuhui Deng 0001 |
ICNP | 3 |
| 2025 | Quark: Implementing Convolutional Neural Networks Entirely on Programmable Data Plane
Mai Zhang, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Zhen Zhang 0017, Yuhui Deng 0001, Zhetao Li |
INFOCOM | 2 |
| 2025 | Against Eavesdropping in Vehicular Networks: Joint Optimization of Power and Spectrum Based on Multi-Agent Reinforcement LearningabstractWith the continuous development of vehicular networks, the security of in-vehicle data is receiving more and more attention. Different from traditional cryptographic techniques, the solution of physical layer security offers the advantages of lower algorithmic complexity and reduced computational demands on devices. This paper addresses the issue of resource allocation for vehicular networks in the presence of eavesdroppers. We formulate the joint optimization problem of power control and spectrum allocation under the quality of service requirements for vehicle-to-infrastructure (V2I) links with high capacity and vehicle-to-vehicle (V2V) links with low latency while ensuring the security of V2V transmission. To cope with the fast-changing channel state information in the high-mobility vehicular scenario, the optimization problem is transformed into a Markov decision process and solved by a multi-agent deep deterministic policy gradients (MADDPG) based deep reinforcement learning (DRL) algorithm. Simulation results show that the proposed MADDPG-based algorithm can effectively ensure high capacity of V2I links while ensuring the low latency of V2V secure transmission. Zhuoyan Feng, Xiujie Huang, Lin Cui 0001, Renzhang Chen, Quanlong Guan |
WCNC | 3 |
| 2025 | Reducing tail latency for multi-bottleneck in datacenter networks: A compound approachabstractThe effectiveness of network congestion control fundamentally depends on the accuracy and granularity of congestion feedback . In datacenter networks, precise feedback is essential for achieving high performance. Most existing approaches use either Explicit Congestion Notification (ECN) or network delay (e.g., RTT) independently as congestion indicators . However, in multi-bottleneck networks, the limitations of these signals become more pronounced: ECN struggles with large cumulative end-to-end latency, while RTT lacks the precision needed to control queuing delays at individual hops. To address these challenges, we propose Cocktail , a simple yet effective transport protocol for datacenter networks that combines both ECN and RTT congestion signals to more effectively handle multi-bottleneck scenarios. By leveraging the ECN signal, Cocktail bounds per-hop queue lengths, enhancing its ability to control single-hop latency and prevent packet loss . Additionally, by estimating RTT, Cocktail effectively manages end-to-end delay, resulting in lower Flow Completion Time (FCT). Extensive experimental evaluations in Mininet demonstrate that Cocktail significantly reduces the average and 99th-percentile completion times for small flows by up to 20% and 29%, respectively, compared to current practices under production workloads. Yuxiang Zhang 0007, Lin Cui 0001, Fung Po Tso 0001, Xiaolin Lei |
Comput. Networks | 2 |
| 2025 | Enhancing In-Network Computing Deployment via Collaboration Across PlanesabstractThe new paradigm of In-network computing (INC) permits service computation to be executed within network paths, rather than solely on dedicated servers. Although the programmable data plane has showcased notable performance advantages for INC application deployments, its effectiveness is constrained by resource limitations, potentially impeding the expressiveness and scalability of these deployments. Conversely, delegating computational tasks to the control plane, supported by general-purpose servers with abundant resources, offers increased flexibility. Nonetheless, this strategy compromises efficiency to a considerable extent, particularly when the system operates under heavy load. To simultaneously exploit the efficiency of data plane and the flexibility of control plane, we proposeCarlo, a cross-plane collaborative optimization framework to support the network-wide deployment of multiple INC applications across both the control and data plane.Carlofirst analyzes resource requirements of various INC applications across different planes. It then establishes mathematical models for resource allocation in cross-plane and automatically generates solutions using proposed algorithms. We have implemented the prototype ofCarloon Intel Tofino ASIC switches and DPDK. Experimental results demonstrate thatCarlocan effectively trade off between computation time and deployment performance while avoiding performance degradation. Xiaoquan Zhang, Lin Cui 0001, Waiming Lau, Fung Po Tso 0001, Yuhui Deng 0001, Weijia Jia 0001 |
IEEE Trans. Computers | 2 |
| 2025 | Stable Task Allocation in Mobile Crowdsensing: An Interruption-Driven ApproachabstractIn mobile crowdsensing, task interruptions can cause failures and reduce system stability. Despite the significance of this issue, few studies have addressed task allocation under interruptions. To bridge this gap, we propose IT-STA, an interruption-based stable task allocation algorithm that reallocates interrupted tasks to improve completion rates and maintain system stability. First, an efficient detection mechanism is designed to promptly identify interrupted tasks, ensuring timely intervention. Second, a distributed reallocation strategy is developed to assign interrupted tasks to suitable participants, leveraging a novel individual migration strategy that enables parallel coordination among nodes, ensuring efficient global matching and avoiding suboptimal solutions. Experimental results demonstrate IT-STA’s superiority over baselines in task allocation stability and performance. Kaimin Wei, Guozi Qi, Lin Cui 0001, Jinpeng Chen 0001, Ke Xu 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2025 | DisPLOY: Target-Constrained Distributed Deployment for Network Measurement Tasks on Data PlaneabstractIn programmable networks, measurement tasks are placed on programmable switches to monitor network traffic at line rate. These tasks typically require substantial resources (e.g., significant SRAM), while programmable switches are constrained by limited resources due to their hardware design (e.g., Tofino ASIC), making distributed deployment essentially. Measurement tasks must monitor specific network locations or traffic flows, introducing significant complexity in deployment optimization. This target-constrained nature makes task optimization on switches (e.g., task merging) become device-dependent and order-dependent, which can lead to deployment failures or performance degradation if ignored. In this paper, we introduceDisPLOY, a novel target-constrained distributed deployment framework specifically designed for network measurement tasks on the data plane.DisPLOYenables operators to specify monitoring targets—network traffic or device/link—across multiple switches. Given the monitoring targets,DisPLOYeffectively minimizes redundant operations and optimizes deployment to achieve both resource efficiency (e.g., minimizing stage consumption) and high-performance monitoring (e.g., high accuracy). We implement and evaluateDisPLOYthrough deployment on both P4 hardware switches (Intel Tofino ASIC) and BMv2. Experimental results show thatDisPLOYsignificantly reduces stage consumption by up to 66% and improves ARE by up to 78.4% in flow size estimation while maintaining end-to-end performance. Mimi Qian, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001, Zhetao Li, Weijia Jia 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | Monte: SFCs Migration Scheme in the Distributed Programmable Data PlaneabstractService function chains (SFCs) are sequences of network functions that provide specific services to meet operators’ needs in today's ISPs and datacenter networks. To improve the performance of SFCs, programmable data planes are used to leverage their low latency and high performance packet processing. However, SFCs need to be adaptable to dynamics such as changes in requirements and attributes. Therefore, the ability to migrate SFCs is essential. Unfortunately, migrating SFCs in distributed programmable data planes is challenging due to the risk of degraded performance and failure to meet SFCs requirements and resource constraints in switches. In this paper, we proposeMonte, which provides an effective SFCs migration scheme in distributed programmable data planes. We build a novel integer programming model to represent the migration process with constraints on resource limitations of switches and SFCs attributes in the distributed data plane. Additionally, an SFCs migration algorithm is designed to optimize the migration cost by deeply analyzing resource allocation in the switch pipeline.Montehas been implemented on both P4 software switches (Bmv2) and hardware switches (Intel Tofino ASIC). Extensive evaluation results show that the migration cost inMonteis 94.03% lower on average than the state-of-the-art deployment scheme, andMontecan effectively save pipeline resources. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Zhetao Li, Weijia Jia 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | FlxVRM: Enabling Online Configuring Memory via Virtualization on Programmable Data PlaneabstractProgrammable data plane (PDP) has emerged as a powerful platform for line-rate packet processing, utilizing on-chip register memory to execute stateful applications. Yet most existing efforts concentrate on static approaches for allocating register memory, necessitating switch restarting and service interruption. Despite the availability of research on sharing memory for concurrent applications, the rigid requirement of limiting memory sharing to the same pipeline stages hampers application flexibility and poses scalability challenges. To address this limitation, we presentFlxVRM,a flexible register memory virtualization layerfor data plane P4 programs which supports high-flexibility sharing of register memory for concurrent applications on PDP.FlxVRMenables memory allocation at any stage and location of the pipeline on PDP for each application at run time. To reduce resource usage during virtualization in the data plane pipeline,FlxVRMfurther merges different tables and actions with similar structures within P4 programs. Additionally,FlxVRMprovides a compiler to generate data plane programs for virtualization as well as the control plane API configuration. A prototype ofFlxVRMis implemented based on P4 hardware switches with Intel Tofino ASIC. Our experiment results show thatFlxVRMsignificantly improves the allocatable memory space for applications by up to 50%, while reducing the resource of the table up to 68%. Mimi Qian, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Zhen Zhang 0017, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | An Energy-Aware Virtual Machine Scheduling Approach for Cloud Data CentersabstractThe reduction of energy consumption will be even more urgent in cloud data centers due to the explosive increase of application data. Virtual machine (VM) integration is a relatively standard technology currently applied for computing facilities of data centers. However, excessive VM consolidation can easily lead to local hot spots that lower the energy efficiency and reliability of data centers. In addition, on account of the impact of heat recirculation in data centers, the traditional VM scheduling strategy cannot comprehensively ponder optimizing the holistic data center energy, which encompasses both server energy and cooling energy. To handle these issues, we proposedEAVMS- an Energy-Aware VM Scheduling approach for minimizing the holistic energy consumption of data centers. EAVMS adopts a two-phase approach to gain energy efficiency while guaranteeing QoS. First, EAVMS leverages a Blended Genetic algorithm and Simulated Annealing algorithm (BGSA) to optimize the initial placement of VMs. Second, EAVMS utilizes a dynamic migration algorithm to achieve effective migration by setting a maximum server temperature threshold without violating the service level agreement (SLA) that cuts down energy consumption by moderating the hot spots of servers. We conducted extensive experiments using two real-world traces (i.e., PlanetLab and Google Cluster datasets) to evaluate the effectiveness of EAVMS. The experimental results unveil that our approach is capable of saving 3.23$ \%$–43.07$ \%$in the holistic energy consumption of cloud data centers with only a tiny service performance degradation compared to other state-of-the-art alternatives (e.g., MJPM, GRANITE, TAS, XINT-GA, and Random). Jie Li 0067, Yuhui Deng 0001, Zijie Zhong, Zhaorui Wu, Shujie Pang, Lin Cui 0001, Geyong Min |
IEEE Trans. Sustain. Comput. | 6 |
| 2024 | Carlo: Cross-Plane Collaboration for Multiple In-network Computing ApplicationsabstractIn-network computing (INC) is a new paradigm that allows applications to be executed within the network, rather than on dedicated servers. Conventionally, INC applications have been exclusively deployed on the data plane (e.g., programmable ASICs), offering impressive performance capabilities. However, the data plane’s efficiency is hindered by limited resources, which can prevent a comprehensive deployment of applications. On the other hand, offloading compute tasks to the control plane, which is underpinned by general-purpose servers with ample resources, provides greater flexibility. However, this approach comes with the tradeoff of significantly reduced efficiency, especially when the system operates under heavy load. To simultaneously exploit the efficiency of data plane and the flexibility of control plane, we propose Carlo, a cross-plane collaborative optimization framework to support the network-wide deployment of multiple INC applications across both the control and data plane. Carlo first analyzes resource requirements of various INC applications across different planes. It then establishes mathematical models for resource allocation in cross-plane and automatically generates solutions using proposed algorithms. We have implemented the prototype of Carlo on Intel Tofino ASIC switches and DPDK. Experimental results demonstrate that Carlo can compute solutions in a short time while avoiding performance degradation caused by the deployment scheme. Xiaoquan Zhang, Lin Cui 0001, Waiming Lau, Fung Po Tso 0001, Yuhui Deng 0001, Weijia Jia 0001 |
INFOCOM | 2 |
| 2024 | Exact Computation of Network Reliability with Sentential Decision DiagramsabstractModern society largely depends on various network systems such as computer networks, communications, and power networks, it is crucial to exactly compute the reliability of these network systems. The exact computation of network reliability is known to be #P-Complete problem. The state-of-the-art exact computation method is based on binary decision diagrams (BDDs). Sentential decision diagrams (SDDs) are a new canonical representation of Boolean functions which is a strict superset of BDD. In both theoretical and practical perspectives, SDDs are a more compact representation than BDDs. In addition, it maintains canonicity and polynomial-time operations on Boolean functions. In this paper, we propose a top-down compilation algorithm based on SDD for representing the set of subgraphs that makes the network active. Then the exact computation of network reliability is based on the constructed SDD. The experiments show that in most real networks and all synthetic networks, the resulting SDD is smaller than the BDD representing the same set of subgraphs. Thus, SDD-based method to exact computation of network reliability is more efficient than BDD-based. Delong Li, Jiayu Zeng, Liangda Fang, Chaonan Wang, Lin Cui 0001, Quanlong Guan |
ISSRE | 5 |
| 2024 | Enabling locality-sensitive machine learning towards low predictive overhead in flow classification
Wenzhi Li, Lin Cui 0001, Xiaoquan Zhang |
Comput. Networks | 2 |
| 2024 | DNN acceleration in vehicle edge computing with mobility-awareness: A synergistic vehicle-edge and edge-edge framework
Lin Cui 0001, Fung Po Tso 0001, Zhetao Li, Weijia Jia 0001 |
Comput. Networks | 2 |
| 2024 | Optimizing the performance of OpenFlow Protocol over QUICabstractIn the last decade, Software-defined networking (SDN) has been an eye-catching network architecture . As the essential protocol of SDN, OpenFlow provides the ability to control switches forwarding plane by the remote controller. Currently, OpenFlow protocol is mostly implemented based on TCP. However, due to the limitations of the TCP protocol (e.g., head-of-line blocking), OpenFlow suffers various problems in practical networks, such as performance degradation and increasing network overhead. This paper investigates these issues through experiments and introduces QUIC to handle these issues. Moreover, a scheduling algorithm , called Extended Performance Modular (EPM), is also proposed to further optimize the performance of OpenFlow protocol with QUIC. By considering both message properties and network conditions, EPM effectively utilizes multiple streams of QUIC to avoid head-of-line blocking and improves the efficiency of OpenFlow. The proposed OpenFlow-QUIC with EPM scheme has been implemented and evaluated on both RYU and Open vSwitch (OVS), and the code has been open-sourced. Extensive experiment results show that, compared to traditional OpenFlow, the proposed scheme can reduce applications latency over OpenFlow by 39.3%, while saving 51.6% network overhead on average. Waiming Lau, Kakei Wong, Lin Cui 0001 |
J. Netw. Comput. Appl. | 3 |
| 2024 | A Combined Trend Virtual Machine Consolidation Strategy for Cloud Data CentersabstractVirtual machine (VM) consolidation strategies are widely used in cloud data centers (CDC) to optimize resource utilization and reduce total energy consumption. Although existing strategies consider current and future resource utilization, the impact of sudden bursts in historical resource utilization on the hosts has been underestimated in uncertain future periods. Insufficient analysis of historical resource utilization may increase the risk of host overloading and Service Level Agreement Violation (SLAV). By defining historical and future trends based on resource utilization, we propose a novel combined trend VM consolidation (CTVMC) strategy which can effectively reduce energy consumption and SLAV. The VMs with the largest combined trend are selected for migration to prevent host overloading. Based on the temporal locality and prediction technique, CTVMC then employs the past, present, and future resource utilization to filter candidate hosts, and identifies the most complementary host to place VM using combined trends. We conduct extensive simulation experiments with PlanetLab Trace and Google Cluster Trace in the CloudSim simulator. Compared with the well-known strategies, CTVMC strategy using the PlanetLab Trace can reduce the number of migrations by over 72.39%, SLAV by over 75.85%, and ESV (a combined metric that judges the trade-off between energy consumption and SLAV) by over 81.54%. According to the Google Cluster Trace, our strategy can reduce the number of migrations by over 61.51%, SLAV by over 37.37%, and ESV by over 35.30%. Zhen Zhang 0017, Yuhui Deng 0001, Geyong Min, Lin Cui 0001 |
IEEE Trans. Computers | 5 |
| 2024 | HVMM: A Holistic Virtual Machine Management Strategy for Cloud Data CentersabstractCloud computing has emerged as an infrastructure in the era of digital economy and has been widely applied in various fields. Virtual Machine(VM) management is the key mechanism in a Cloud Data Center(CDC). A typical VM management system is responsible for VM allocation and VM reallocation, and it is usually designed to optimize specific objectives, especially for the metrics of energy consumption, resource wastage, communication cost, and Service Level Agreement Violations (SLAV). However, it is greatly challenging to optimize these metrics at the same time, and most existing VM management strategies focus on optimizing part of the above four metrics. In this paper, we propose a Holistic Virtual Machine Management (HVMM) strategy to optimize the energy consumption, resource wastage, communication cost, and SLAV simultaneously. First, we define two parameters, the Compatibility and the Performance-to-Power Ratio (PPR), for VM allocation to optimize energy consumption and resource wastage. Then, we propose a reallocation approach based on spectral clustering that can handle dynamic traffic between VMs without a priori knowledge of the traffic between VMs, and it takes a slight expense of resource wastage and energy consumption to reduce communication cost between VMs and ensures low SLAV risk. To evaluate the performance of HVMM, we compared the proposed strategy with the state-of-the-art strategies in various experiments on real-world traces. Compared with the other strategies, the resource wastage of HVMM is reduced by 62%. Simultaneously, the communication cost is reduced by about 26%, the energy consumption is reduced by about 5%, and SLAV is much lower than that of the others. Piao Lv, Zhen Zhang 0017, Yuhui Deng 0001, Lin Cui 0001, Longxin Lin |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2024 | OffsetINT: Achieving High Accuracy and Low Bandwidth for In-Band Network TelemetryabstractNetwork measurement is essential for efficient network management and operations. In-band network telemetry (INT) offers fine-grained per-device per-packet information which could provide full-visibility for networks. However, the existing solutions fall short in achieving high accuracy, generality, and low overhead simultaneously. To address this limitation, we introduceOffsetINTto meet these three criteria. The key idea ofOffsetINTis to use minimal bits to carry collected states during monitoring, which is based on our observation that the value of telemetry states are usually very close (e.g., the time of adjacent arrival packets) or small (e.g., only a few tens of microseconds for processing latency) for most of the time in real networks. Instead of embedding complete values of state in packets,OffsetINToptimizes bit usage by encoding an offset (using fewer bits), which is carried in-band by passing packets to the end-hosts for recovery and analysis. We theoretically derive the bounds of bandwidth mitigation forOffsetINT. We have implementedOffsetINTin both P4 hardware switches (with Intel Tofino ASIC) and BMv2. Expensive evaluation results show thatOffsetINTcan achieve an accuracy of up to 100% compared to the original INT while reducing INT bandwidth by up to 48%. Mimi Qian, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | Fine-grained HTTP/3 prioritization via reinforcement learningabstractAs the latest version of HTTP, HTTP/3 reduces the loading time of web pages and improves the user experience by replacing TCP and TLS with QUIC. Previous studies have already demonstrated the importance of prioritization for optimizing the performance of HTTP. By default, HTTP/3 schedules all the requests in a round-robin (RR) way on the server. However, in practice, in addition to the different characteristics of web resources, the differences and dynamic changes of network conditions and user equipment (e.g., mobile and PC) will also have a great impact on prioritization. Without considering these factors, existing prioritization schemes deployed on both clients and servers cannot always ensure optimal performance for HTTP/3. In light of the above issues, in this paper, we proposed D ynamic R esources P rioritization via R einforcement L earning (DRP-RL) to provide fine-grained resource prioritization for HTTP/3 by considering the effects of both network and user equipment. Reinforcement learning (RL) is adopted, in which the RL agent can leverage network, user and web page information to learn the best prioritization of resources across different user groups. The HTTP/3 server is instructed to send the resources in a particular order for different clients dynamically to improve performance. DRP-RL has been implemented based on quic-go , and extensive evaluations indicate that DRP-RL minimizes 3.1% ∼ 21.0% of First Content Paint (FCP) and saves 5.3% ∼ 23.7% of Page Load Time (PLT) across various web pages when compared with RR. Kakei Wong, Lin Cui 0001 |
Comput. Networks | 2 |
| 2023 | A survey on sliding window sketch for network measurement
Zijie Zeng, Lin Cui 0001, Mimi Qian, Zhen Zhang 0017, Kaimin Wei |
Comput. Networks | 2 |
| 2023 | HashCache: Accelerating Serverless Computing by Skipping Duplicated Function ExecutionabstractServerless computing is a leading force behind deploying and managing software in cloud computing. One inherent challenge in serverless computing is the increased overall latency due to duplicate computations. Our initial investigation into the function invocations of serverless applications reveals an abundance of duplicate invocations. Inspired by this critical observation, we introduceHashCache, a system designed to cache duplicate function invocations, thereby mitigating duplicate computations. In HashCache, serverless functions are classified into three categories, namely, computational functions, stateful functions, and environment-related functions. On the grounds of such a function classification, HashCache associates the stateful functions and their states to build an adaptive synchronization mechanism. With this support, HashCache exploits the cached results of computational and stateful functions to serve upcoming invocation requests to the same functions, thereby reducing duplicate computations. Moreover, HashCache stores remote files probed by stateful functions into a local cache layer, which further curtails invocation latency. We implement HashCache within theApache OpenWhiskto forge a cache-enabled serverless computing platform. We conduct extensive experiments to quantitatively evaluate the performance of HashCache in terms of invocation latency and resource utilization. We compare HashCache against two state-of-the-art approaches -FaaSCacheandOpenWhisk. The experimental results unveil that our HashCache remarkably reduces invocation latency and resource overhead. More specifically, HashCache curbs the 99-tail latency of FaaSCache and OpenWhisk by up to 91.37% and 95.96% in real-world serverless applications. HashCache also slashes the resource utilization of FaaSCache and OpenWhisk by up to 31.62% and 35.51%, respectively. Zhaorui Wu, Yuhui Deng 0001, Yi Zhou 0009, Lin Cui 0001, Xiao Qin 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Compiling Service Function Chains via Fine-Grained Composition in the Programmable Data PlaneabstractService function chains (SFCs) are fundamental services in today's datacenters and ISP networks. Explosive volume of network traffic creates high demands for low latency and high performance. The emergence of programmable data planes has offered a new way to overcome the problem. However, limited by pipeline constraints in hardware architecture, implementing multiple network functions on programmable data planes is challenging. Besides, considering various types of network functions, e.g., stateful network functions, a general model is essential for abstracting distinct network functions. In this article, we proposepSFCwhich provides a fine-grained SFCs deployment scheme in programmable data planes. Control flow graph (CFG) is proposed to abstract and analyze various network functions. Then we model pipeline constraints in the hardware architecture using an ILP (Integer Linear Programming), and model the SFCs deployment in the substrate network as a one big switch (OBS) problem. To reduce deployment cost,pSFCfirst composes multiple SFCs to a compound CFG for eliminating redundant logics within SFCs, further decomposes the compound CFG based on the resource limitation per stage, and finally maps the OBS into the substrate network. We have implementedpSFCin both bmv2 software switch and P4 hardware switch (i.e., Intel Tofino ASIC). Evaluation results show thatpSFCreduces switch costs by 45.7% and decreases average latency by 22% without compromising throughput. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | Dapper: Deploying Service Function Chains in the Programmable Data Plane Via Deep Reinforcement LearningabstractNetwork functions perform specific packet processing on network traffic. To meet operators' needs, forming service function chains (SFCs) is a fundamental technique used in today's ISPs and datacenter networks. Implementing SFCs in the programmable data plane with high throughput and low latency is a new approach to satisfy demands of ever-growing network traffic. Previous works have proposed different solutions to solve the problem, but they all inevitably have to make trade-offs between running time and performance. For example, an ILP (Integer Linear Programming) can optimize cost but suffers from long running time in large-scale network topologies. Heuristic algorithms depend strongly on manual designs and usually have a performance gap with the optimal solution. In this paper, we proposeDapper, a framework for deploying SFCs in the programmable data plane using DRL (Deep Reinforcement Learning) with graph convolutional network. In order to expand the searching space to prevent the optimal value from being missed,Dapperallows the RL (Reinforcement Learning) agent to simultaneously extract features from both the substrate network and the hardware pipeline, and exploit a graph convolutional network to enhance performance. Moreover, a mask mechanism is also designed to accelerateDapperand improve its scalability.Dapperhas been implemented and extensively evaluated on both P4 hardware switches (equipped with Intel Tofino ASIC) and software switches (i.e., bmv2). Experimental results show thatDappercan automatically generate deployment solutions in a few seconds of running time after training. They also demonstrate thatDapperreduces hardware stage usage and the latency of SFCs by up to 17.8% and 50$\sim$73% respectively on average when compared with heuristics. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Zhetao Li, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | pSFC: Fine-grained Composition of Service Function Chains in the Programmable Data PlaneabstractDynamic service function chains (SFC) are enabled by network function virtualization on general purpose servers. The emergence of programmable data planes (PDP) has offered a new way for the deployment of SFC. However, the implementation of network functions is constrained by resource limitations in PDPs (e.g., compute and memory resource). Moreover, most of existing works do not consider the optimization of state information (e.g., registers), which is essential for stateful network functions. In this paper, we propose pSFC which provides a fine-grained SFC deployment scheme in the PDP to tackle the problem. We first model network functions as control flow graphs (CFG) and the process of deployment as a one big switch (OBS) problem, and then propose an ILP (Integer Linear Programming) model for resource optimization for the OBS problem, which is NP-hard. To solve this problem efficiently, pSFC first composes multiple SFCs for eliminating redundant resources, decomposes the compound CFG based on the resource limitation per stage, and finally maps OBS into the substrate network. We have implemented pSFC in both bmv2 software switch and P4 hardware switch (i.e., Intel Tofino). Evaluation shows that pSFC reduces switch costs 45.7% and average latency 15% while providing the correctness of the process of SFC. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001 |
CCGRID | 2 |
| 2022 | B-Scale: Bottleneck-aware VNF Scaling and Flow Routing in Edge CloudsabstractWith the ever-growing demand for low-latency network applications, edge computing emerges as a new paradigm that provides computation and storage resources in close proximity to end-users. Many research efforts have resorted to network function virtualization, wherein network applications are provisioned as service function chains at edge clouds. However, due to the traffic dynamics and limited resource capacity at the network edge, how to efficiently embed service chains with latency optimization and resource efficiency remains as a challenging problem. As most existing research efforts largely overlook the bottle-necked resources of VNFs in the VNF scaling, we seek a more realistic approach to provisioning VNF instances across multiple edge clouds. Also, given the limited resources at the edge, it is of significant importance to improve the VNF utilization rate. Specifically, we formulate the VNF scaling problem as an integer linear programming (ILP) problem, aiming to minimize the end-to-end latency for service function chains. To solve this problem, we devise a novel bottleneck-aware algorithm that manages the number and deployment of newly created instances. After that, we propose an online algorithm for traffic steering to improve the utilization rates of VNF instances and avoid congestion on hotspot links. The proposed algorithm is shown to provide good performance by trace-driven simulation in real-world topologies. Chen Chen 0073, Lars Nagel 0001, Lin Cui 0001, Fung Po Tso 0001 |
ISCC | 3 |
| 2022 | Distributed federated service chaining: A scalable and cost-aware approach for multi-domain networksabstractFuture networks are expected to support cross-domain, cost-aware and fine-grained services in an efficient and flexible manner. Service Function Chaining (SFC) has been introduced as a promising approach to deliver these services. In the literature, centralized resource orchestration is usually employed to process SFC requests and manage computing and network resources. However, centralized approaches inhibit the scalability and domain autonomy in multi-domain networks. They also neglect location and hardware dependencies of service chains. In this paper, we propose Distributed Federated Service Chaining (DFSC), a framework for orchestrating and maintaining SFC placement in a distributed fashion while sharing only a minimal amount of domain information and control. First, a deployment cost minimization problem is formulated as an Integer Linear Programming (ILP) problem with fine-grained constraints for location and hardware dependencies. We show that this problem is NP-hard. Then, a placement algorithm is devised to use information only on inter-domain paths and border nodes. Our extensive experimental results demonstrate that DFSC efficiently optimizes the deployment cost, supports domain autonomy and enables faster decision-making. The results also show that DFSC finds solutions within a factor 1.15 of the optimal solution on average. Compared to a centralized approach in the literature, DFSC reduces the deployment cost by up to 20% and uses 70% less decision-making time. Chen Chen 0073, Lars Nagel 0001, Lin Cui 0001, Fung Po Tso 0001 |
Comput. Networks | 3 |
| 2022 | dDrops: Detecting silent packet drops on programmable data plane
Mimi Qian, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001 |
Comput. Networks | 2 |
| 2022 | Optimizing multipath QUIC transmission over heterogeneous paths
Hongxin Zeng, Lin Cui 0001, Fung Po Tso 0001, Zhen Zhang 0017 |
Comput. Networks | 2 |
| 2022 | Low-latency service function chain migration in edge-core networks based on open Jackson networks
Lin Cui 0001, Fung Po Tso 0001 |
J. Syst. Archit. | 2 |
| 2021 | A survey on stateful data plane in software defined networks
Xiaoquan Zhang, Lin Cui 0001, Kaimin Wei, Fung Po Tso 0001, Yangyang Ji, Weijia Jia 0001 |
Comput. Networks | 2 |
| 2021 | pHeavy: Predicting Heavy Flows in the Programmable Data PlaneabstractSince heavy flows account for a significant fraction of network traffic, being able to predict heavy flows has benefited many network management applications for mitigating link congestion, scheduling of network capacity, exposing network attacks and so on. Existing machine learning based predictors are largely implemented on the control plane of Software Defined Networking (SDN) paradigm. As a result, frequent communication between the control and data planes can cause unnecessary overhead and additional delay in decision making. In this paper, we presentpHeavy, a machine learning based scheme for predicting heavy flows directly on the programmable data plane, thus eliminating network overhead and latency to SDN controller. Considering the scarce memory and limited computation capability in the programmable data plane,pHeavyincludes a packet processing pipeline which deploys pre-trained decision tree models for in-network prediction. We have implementedpHeavyin both bmv2 software switch and P4 hardware switch (i.e., Barefoot Tofino). Evaluation results demonstrate thatpHeavyhas achieved 85% and 98% accuracy after receiving the first 5 and 20 packets of a flow respectively, while being able to reduce the size of decision tree by 5.4x on average. More importantly,pHeavycan predict heavy flows at line rate on the P4 hardware switch. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2019 | Autonomous Flying WiFi Access PointabstractUnmanned aerial vehicles (UAVs), aka drones, are widely used civil and commercial applications. A promising one is to use the drones as relying nodes to extend the wireless coverage. However, existing solutions only focus on deploying them to predefined locations. After that, they either remain stationary or only move in predefined trajectories throughout the whole deployment. In the open outdoor scenarios such as search and rescue or large music events, etc., users can move and cluster dynamically. As a result, network demand will change constantly over time and hence will require the drones to adapt dynamically. In this paper, we present a proof of concept implementation of an UAV access point (AP) which can dynamically reposition itself depends on the users movement on the ground. Our solution is to continuously keeping track of the received signal strength from the user devices for estimating the distance between users devices and the drone, followed by trilateration to localise them. This process is challenging because our on-site measurements show that the heterogeneity of user devices means that change of their signal strengths reacts very differently to the change of distance to the drone AP. Our initial results demonstrate that our drone is able to effectively localise users and autonomously moving to a position closer to them. Gareth J. Nunns, Yu-Jia Chen, Deng-Kai Chang, Kai-Min Liao, Fung Po Tso 0001, Lin Cui 0001 |
ISCC | 6 |
| 2019 | Extensive evaluation on the performance and behaviour of TCP congestion control protocols under varied network scenarios
Jinting Lin, Lin Cui 0001, Yuxiang Zhang 0007, Fung Po Tso 0001, Quanlong Guan |
Comput. Networks | 2 |
| 2019 | Mystique: A Fine-Grained and Transparent Congestion Control Enforcement SchemeabstractTCP congestion control is a vital component for the latency of Web services. In practice, a single congestion control mechanism is often used to handle all TCP connections on a Web server, e.g., Cubic for Linux by default. Considering complex and ever-changing networking environment, the default congestion control may not always be the most suitable one. Adjusting congestion control to meet different networking scenarios usually requires modification of TCP stacks on a server. This is difficult, if not impossible, due to various operating system and application configurations on production servers. In this paper, we propose Mystique, a light-weight, flexible, and dynamic congestion control switching scheme that allows network or server administrators to deploy any congestion control schemes transparently without modifying existing TCP stacks on servers. We have implemented Mystique in Open vSwitch (OVS) and conducted extensive test-bed experiments in both public and private cloud environments. Experiment results have demonstrated that Mystique is able to effectively adapt to varying network conditions, and can always employ the most suitable congestion control for each TCP connection. More specifically, Mystique can significantly reduce latency by 18.13% on average when compared with individual congestion controls. Yuxiang Zhang 0007, Lin Cui 0001, Fung Po Tso 0001, Quanlong Guan, Weijia Jia 0001, Jipeng Zhou |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2019 | Enabling Heterogeneous Network Function ChainingabstractToday's data center operators deploy network policies in both physical (e.g., middleboxes, switches) and virtualized (e.g., virtual machines on general purpose servers) network function boxes (NFBs), which reside in different points of the network, to exploit their efficiency and agility respectively. Nevertheless, such heterogeneity has resulted in a great number of independent network nodes that can dynamically generate and implement inconsistent and conflicting network policies, making correct policy implementation a difficult problem to solve. Since these nodes have varying capabilities, services running atop are also faced with profound performance unpredictability. In this paper, we propose a Heterogeneous netwOrk Policy Enforcement (HOPE) scheme to overcome these challenges. HOPE guarantees that network functions (NFs) that implement a policy chain are optimally placed onto heterogeneous NFBs such that the network cost of the policy is minimized. We first experimentally demonstrate that the processing capacity of NFBs is the dominant performance factor. This observation is then used to formulate the Heterogeneous Network Policy Placement problem, which is shown to be NP-Hard. To solve the problem efficiently, an online algorithm is proposed. Our experimental results demonstrate that HOPE achieves the same optimality as Branch-and-bound optimization but is 3 orders of magnitude more efficient. Lin Cui 0001, Fung Po Tso 0001, Song Guo 0001, Weijia Jia 0001, Kaimin Wei, Wei Zhao 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | Latency-aware joint virtual machine and policy consolidation for mobile edge computingabstractTo guarantee an efficient and high-performance environment for mobile devices to perform offloading with low end-to-end delay, it is important to ensure no network policies are violated. In this paper, we explore the simultaneous, dynamic virtual machine (VM) and policy consolidation, and formulate the Policy-VM Latency-aware Consolidation problem for Mobile Edge Computing, which is shown to be NP-Hard. We propose the PL-Edge, an efficient scheme to jointly consolidate network policies and virtual machines for mobile edge computing to reduce communication end-to-end delays among devices and virtual machines. Our simulation results demonstrate that the proposed PL-Edge can significantly reduces policy-flows end-to-end delay by nearly 45% while adhering strictly to the requirements of network policies. Thiago A. L. Genez, Fung Po Tso 0001, Lin Cui 0001 |
CCNC | 3 |
| 2018 | Dynamic Network Function Chain Composition for Mitigating Network LatencyabstractNetwork Function Virtualisation (NFV) enables rapid deployment of new services in networks on an on-demand basis using general purpose servers. Multiple virtual network functions (VNFs) can be dynamically chained in an ordered sequence for the delivery of end-to-end services. Nevertheless, network latency caused by the sequential order of packet processing on every VNF can hurt the performance of latency-sensitive applications. To reduce such network latency, existing solutions only consider the maximum capacity of individual virtual network functions (VNFs) and do not take into account the fact that performance of VNFs, as with any software applications, is bottlenecked by either CPU or I/O peripheral capacity of the server they run on and their underneath implementation such as singleor multi-threaded.By exploiting this knowledge, we can better determine the number of required VNF instances and distribute the network traffic among them for any given VNF chain. In this paper, we formulate the VNF Scaling and Traffic Distribution problem and prove that it is NP-hard. We then present the design and implementation of Natif, an efficient VNF-Aware VNF insTantIation and traFfic distribution scheme. Through our OpenStack-based testbed evaluations, we demonstrate that Natif can significantly improve the network latency by 188% on average as compared to other approaches. As a chain composition scheme, Natif can effectively work with any VNF chaining algorithms. Wajdi Hajji, Thiago A. L. Genez, Fung Po Tso 0001, Lin Cui 0001, Iain Phillips 0002 |
ISCC | 4 |
| 2018 | Modest BBR: Enabling Better Fairness for BBR Congestion ControlabstractAs a vital component of TCP, congestion control defines TCP's performance characteristics. Hence, it is important for congestion control to provide high link utilization and low queuing delay. Recent BBR tries to estimate available bottleneck capacity to achieve this goal. However, its aggressiveness characteristics generate a massive amount of packet retransmission which harms loss-based congestion control protocol such as Cubic. In this paper, we first dive into this issue and reveal that the aggressiveness of BBR can degrade the performance of Cubic, as well as the overall Internet transmission. Then we present Modest BBR, a simple yet effective solution based on BBR, by responding to retransmission less aggressively. Through extensive testbed experiments and Mininet simulation, we show Modest BBR can preserve high throughput and short convergence time while improve the overall performance when coexisting with Cubic. For example, Modest BBR gets similar throughput compared to BBR, while it improves 7.1% of the overall throughput and achieves better fairness to loss-based schemes. Yuxiang Zhang 0007, Lin Cui 0001, Fung Po Tso 0001 |
ISCC | 2 |
| 2018 | Enforcing network policy in heterogeneous network function box environment
Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
Comput. Networks | 1 |
| 2018 | A survey on software defined networking with multiple controllers
Lin Cui 0001, Yuxiang Zhang 0007 |
J. Netw. Comput. Appl. | 2 |
| 2017 | Heterogeneous NetwOrk Policy Enforcement in data centersabstractWith the emergence of network function virtualization, data center start to deploy a variety of network function boxes (NFBs) in both physical and virtual form factors in order to combines inherent efficiency offered by physical NFBs with the agility and flexibility of virtual ones. However, existing schemes are limited to exclusively consider physical or virtual NFBs, which may reduce the performance efficiency of services running atop. In this paper, we propose a Heterogeneous NetwOrk Policy Enforcement scheme (HOPE) to overcome these challenges. An efficient algorithm that can closely approximate optimal latency-wise NF service chaining is proposed. The experimental results have also shown that HOPE can outperform greedy algorithm by 25% in terms of network latency and is 56× more efficient than naive depth-first search algorithm. Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
IM | 1 |
| 2017 | Experimental evaluation of SDN-controlled, joint consolidation of policies and virtual machinesabstractMiddleboxes (MBs) are ubiquitous in modern data centre (DC) due to their crucial role in implementing network security, management and optimisation. In order to meet network policy's requirement on correct traversal of an ordered sequence of MBs, network administrators rely on static policy based routing or VLAN stitching to steer traffic flows. However, dynamic virtual server migration in virtual environment has greatly challenged such static traffic steering. In this paper, we design and implement Sync, an efficient and synergistic scheme to jointly consolidate network policies and virtual machines (VMs), in a readily deployable Mininet environment. We present the architecture of Sync framework and open source its code. We also extensively evaluate Sync over diverse workload and policies. Our results show that in an emulated DC of 686 servers, 10k VMs, 8k policies, and 100k flows, Sync processes a group of 900 VMs and 10 VMs in 634 seconds and 4 seconds respectively. Wajdi Hajji, Fung Po Tso 0001, Lin Cui 0001, Dimitrios P. Pezaros |
ISCC | 3 |
| 2017 | TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
Yuxiang Zhang 0007, Lin Cui 0001, Fung Po Tso 0001, Quanlong Guan, Weijia Jia 0001 |
NPC | 2 |
| 2017 | A stable matching based elephant flow scheduling algorithm in data center networks
Yuxiang Zhang 0007, Lin Cui 0001 |
Comput. Networks | 2 |
| 2017 | PLAN: Joint Policy- and Network-Aware VM Management for Cloud Data CentersabstractPolicies play an important role in network configuration and therefore in offering secure and high performance services especially over multi-tenant Cloud Data Center (DC) environments. At the same time, elastic resource provisioning through virtualization often disregards policy requirements, assuming that the policy implementation is handled by the underlying network infrastructure. This can result in policy violations, performance degradation and security vulnerabilities. In this paper, we define PLAN, a PoLicy-Aware and Network-aware VM management scheme to jointly consider DC communication cost reduction through Virtual Machine (VM) migration while meeting network policy requirements. We show that the problem is NP-hard and derive an efficient approximate algorithm to reduce communication cost while adhering to policy constraints. Through extensive evaluation, we show that PLAN can reduce topology-wide communication cost by 38 percent over diverse aggregate traffic and configuration policies. Lin Cui 0001, Fung Po Tso 0001, Dimitrios P. Pezaros, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Synergistic policy and virtual machine consolidation in cloud data centersabstractIn modern Cloud Data Centers (DC)s, correct implementation of network policies is crucial to provide secure, efficient and high performance services for tenants. It is reported that the inefficient management of network policies accounts for 78% of DC downtime, challenged by the dynamically changing network characteristics and by the effects of dynamic Virtual Machine (VM) consolidation. While there has been significant research in policy and VM management, they have so far been treated as disjoint research problems. In this paper, we explore the simultaneous, dynamic VM and policy consolidation, and formulate the Policy-VM Consolidation (PVC) problem, which is shown to be NP-Hard. We then propose Sync, an efficient and synergistic scheme to jointly consolidate network policies and virtual machines. Extensive evaluation results and a testbed implementation of our controller show that policy and VM migration under Sync significantly reduces flow end-to-end delay by nearly 40%, and network-wide communication cost by 50% within few seconds, while adhering strictly to the requirements of network policies. Lin Cui 0001, Richard Cziva, Fung Po Tso 0001, Dimitrios P. Pezaros |
INFOCOM | 1 |
| 2016 | Multiple Region of Interest Coverage in Camera Sensor Networks for Tele-Intensive Care UnitsabstractCamera sensor networks (CSNs) are gradually being used in a tele-intensive care unit (tele-ICU), providing useful patient information to remote intensivists. Intensivists wish to focus on different regions of interest (RoIs) containing their patients. We consider a situation where preinstalled camera sensors' locations remain static and they can change the fields of view only by rotating orientations, while the RoIs are dynamically changed in location and size because of changes in the number of patients and care unit configuration. Therefore, an important issue is how to enhance the coverage of these RoIs by controlling the camera sensors' orientations. Previous studies on coverage optimization either focus on single area coverage or point(s) coverage. However, ignoring those multiple RoIs or simply treating them as points can cause unwanted coverage, resulting in performance degradation. In this paper, we investigate a novel multiple RoI coverage (MRC) problem in a CSN-based tele-ICU, aiming to maximize the lowest coverage ratio of all RoIs. The MRC problem is nondeterministic polynomial-time hard, so we propose an efficient heuristic algorithm MRC-Priority to solve it. We have implemented a CSN testbed to evaluate the performance of our proposed algorithm. Experimental results show that our proposed algorithm can improve the lowest coverage ratio up to 200% as compared with existing solutions. Bo Cheng 0011, Lin Cui 0001, Weijia Jia 0001, Wei Zhao 0001, Gerhard P. Hancke 0002 |
IEEE Trans. Ind. Informatics | 2 |
| 2015 | Policy-Aware Virtual Machine Management in Data Center NetworksabstractPolicies play an important role in network configuration and, therefore, in offering secure and high performance services, especially over multi-tenant Cloud Data Center (DC) environments. At the same time, elastic resource provisioning through virtualization often disregards policy requirements, assuming that the policy implementation is handled by the underlying network infrastructure. In this paper, we define PLAN, a Policy-Aware virtual machine management scheme to jointly consider DC communication cost reduction through Virtual Machine (VM) migration while meeting network policy requirements. Lin Cui 0001, Fung Po Tso 0001, Dimitrios P. Pezaros, Weijia Jia 0001, Wei Zhao 0001 |
ICDCS | 1 |
| 2015 | Fincher: Elephant flow scheduling based on stable matching in data center networksabstractWith the development of cloud computing in recent years, data center networks have become a hot topic in both industrial and academic communities. Previous studies have shown that elephant flows, which usually carry large amount of data, are critical to the efficiency of data centers. In this paper, we study the flow scheduling problem in data centers with a focus on elephant flows. By applying stable matching theory, the scheduling problem is modeled and some useful method is complemented. Then, we propose Fincher, an efficient scheme leveraging Software-Defined Networking (SDN) to reduce latency and avoid congestions in data centers. Yuxiang Zhang 0007, Lin Cui 0001, Qiao Chu |
IPCCC | 2 |
| 2013 | Cyclic stable matching for three-sided networking services
Lin Cui 0001, Weijia Jia 0001 |
Comput. Networks | 1 |
| 2013 | DragonNet: A Robust Mobile Internet Service System for Long-Distance TrainsabstractAbstract—Wide range wireless networks often suffer from annoying service deterioration due to fickle wireless environment. This is especially the case with passengers on long distance train (LDT) to connect onto the Internet. To improve the service quality of wide range wireless networks, we present the DragonNet protocol with its implementation. The DragonNet system is a chained gateway which consists of a group of interlinked DragonNet routers working specifically for mobile chain transport systems. The protocol makes use of the spatial diversity of wireless signals that not all spots on a surface see the same level of radio frequency radiation. In the case of a LDT of around 500 meters, it is highly possible that some of the spanning routers still see sound signal quality, when the LDT is partially blocked from wireless Internet. DragonNet protocol fully utilizes this feature to amortize single point router failure over the whole router chain by intelligently rerouting traffics on failed ones to sound ones. We have implemented the DragonNet system and tested it in real railways over a period of three months. Our results have pinpointed two fundamental contributions of DragonNet protocol. First, DragonNet significantly reduces average temporary communication blackout (i.e. no Internet connection) to 1.5 seconds compared with 6 seconds that without DragonNet protocol. Second, DragonNet efficiently doubles the aggregate throughput on average. Fung Po Tso 0001, Lin Cui 0001, Lizhuo Zhang, Weijia Jia 0001, Di Yao 0006, Jin Teng, Dong Xuan |
IEEE Trans. Mob. Comput. | 2 |
| 2011 | DragonNet: A robust mobile Internet service system for long distance trainsabstractWide range wireless networks often suffer from annoying service deterioration due to fickle wireless environment. This is especially the case with passengers on long distance train (LDT) to connect onto the Internet. To improve the service quality of wide range wireless networks, we present the DragonNet protocol with its implementation. The DragonNet system is a chained gateway which consists of a group of interlinked DragonNet routers working specifically for mobile chain transport systems. The protocol makes use of the spatial diversity of wireless signals that not all spots on a surface see the same level of radio frequency radiation. In the case of a LDT of around 500 meters, it is highly possible that some of the spanning routers still see sound signal quality, when the LDT is partially blocked from wireless Internet. DragonNet protocol fully utilizes this feature to amortize single point router failure over the whole router chain by intelligently rerouting traffics on failed ones to sound ones. We have implemented the DragonNet system and tested it in real railways over a period of three months. Our results have pinpointed two fundamental contributions of DragonNet protocol. First, DragonNet significantly reduces average temporary communication blackout (i.e. no Internet connection) to 1.5 seconds compared with 6 seconds that without DragonNet protocol. Second, DragonNet efficiently doubles the aggregate throughput on average. Fung Po Tso 0001, Lin Cui 0001, Lizhuo Zhang, Weijia Jia 0001, Di Yao 0006, Jin Teng, Dong Xuan |
INFOCOM | 2 |