EDBT 2026 Demo / reviewers in the wild / expert
Fung Po Tso 0001
dblp:75/803 · also Posco Tso
· DBLP profile ↗
64ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0001-9366-8285ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 35 · 6 first-author · 15 since 2021Systems, architecture and hardware · 16 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorArtificial intelligence and machine learning · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPRINT: Line-Rate In-band Network Telemetry Recovery for Application Optimization
Bingzhen Chen, WaiMing Lau, Xiaoquan Zhang, Fung Po Tso 0001, Lin Cui 0001 |
INFOCOM | 4 |
| 2026 | Monic: In-Network Mixture-of-Experts Inference on Programmable Data Planes
Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001, Zhen Zhang 0017, Kaimin Wei, Weijia Jia 0001, Lin Cui 0001 |
INFOCOM | 3 |
| 2026 | FaasOrc: A bi-level function scheduling and caching framework for serverless edge computingabstractServerless computing, underpinned by an event-driven approach with transient stateless containers, significantly enhances resource efficiency and simplifies function development. To maintain an acceptable Quality of Service (QoS) agreed in the Service Level Agreement (SLA), service providers need to improve the response latency while considering resource efficiency. However, cold-start delays in container initialization often lead to considerable latency in these applications. Existing mitigation strategies, such as pre-warming and function caching, are inadequate due to workload skewness and oscillation across edge nodes. These limitations are particularly critical in resource-constrained edge environments. We must jointly consider multiple factors, such as node status, function resource requirement and function popularity. To overcome these limitations, this paper presents FaasOrc , a bi-level function orchestration framework to mitigate workload skewness and oscillation across edge nodes. FaasOrc uses a cluster-level scheduler to schedule requests and a node-level manager to detect popular functions. Our comprehensive evaluation, consisting of two parts: simulations and a real-system prototype over Knative , benchmarks the proposed solution against existing scheduling and caching strategies. The findings highlight our method’s capability to reduce the response latency by 31.4%. Chen Chen 0073, Lars Nagel 0001, Lin Cui 0001, Weijia Jia 0001, Fung Po Tso 0001 |
J. Netw. Comput. Appl. | 5 |
| 2026 | dVRM: Cross-Switch Memory Sharing and Self-Adaptive Allocation in Distributed Data PlaneabstractProgrammable switches have revolutionized networking by enabling a new spectrum of applications, such as network telemetry, in-network computation, and machine learning. These applications heavily utilize register memory but their performance is significantly constrained by the scarcity of on-chip resources, such as the 15 MB of SRAM available on a Tofino switch. To effectively accommodate increasingly memorydemanding applications, we aim to pool register resources across multiple switches, creating a larger unified register memory space. This resource pooling approach addresses the limitations of existing single-switch Virtual Register Memory (VRM) solutions, which cannot meet the demands of these applications in distributed environments. To achieve this, we propose dVRM, a distributed VRM deployment framework that enables crossswitch memory sharing and self-dynamic memory allocation on the data plane. dVRM introduces three innovations: (1) Grouped Multi-Switch Registers (GMRs), virtualizing distributed pipeline stages into a unified memory pool; (2) a self-adaptive, bit-width allocation mechanism driven by real-time data-plane feedback; and (3) lightweight heuristics for concurrent application deployment with distributed VRM, formulated as a mixed-integer linear programming (MILP) problem. We have implemented dVRM on both P4 hardware switches (with Intel Tofino ASIC) and BMv2. Experimental results show that dVRM significantly reduces hash unit consumption by up to 26% and achieves an improvement in accuracy (ARE) of up to 57.3% across various workloads. Mimi Qian, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Zhen Zhang 0017, Weijia Jia 0001 |
IEEE Trans. Computers | 3 |
| 2025 | Planner: A Generative Graph Learning Framework for Noisy and Dynamic In-band Network TelemetryabstractIn-band network telemetry (INT) enables real-time network monitoring by embedding telemetry data into packets. The advent of programmable switches further enhances the flexibility of INT by enabling dynamic customization of telemetry collection at the hardware level. However, the practical application of INT is hampered by significant challenges arising from data noise (due to packet loss, delay, and measurement inaccuracies) and network dynamics (such as changing INT paths and feature requirements). These issues severely degrade the performance of machine learning models used for analyzing INT data, hindering the accurate capture of spatio-temporal network characteristics. This paper presents Planner, a novel generative graph learning framework designed to address these limitations. Planner enables the collection of network features at various levels of granularity on programmable switches. Crucially, it constructs dynamic graphs representing evolving INT paths and employs a hybrid Graph Neural Network (GNN) and Recurrent Neural Network (RNN) architecture to effectively learn spatial and temporal dependencies. Furthermore, Planner incorporates variational inference to generate robust latent representations, mitigating the detrimental effects of noise and instability in INT data. We have implemented a testbed prototype of Planner using Intel Tofino ASIC switches. Extensive experiments demonstrate the performance superiority and robustness of Planner over the baseline methods, achieving a 23.2% improvement in F1 score. Xiaoquan Zhang, Waiming Lau, Lin Cui 0001, Fung Po Tso 0001, Zhuoqian Liang, Zhen Zhang 0017, Yuhui Deng 0001 |
ICNP | 4 |
| 2025 | Quark: Implementing Convolutional Neural Networks Entirely on Programmable Data Plane
Mai Zhang, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Zhen Zhang 0017, Yuhui Deng 0001, Zhetao Li |
INFOCOM | 4 |
| 2025 | Reducing tail latency for multi-bottleneck in datacenter networks: A compound approachabstractThe effectiveness of network congestion control fundamentally depends on the accuracy and granularity of congestion feedback . In datacenter networks, precise feedback is essential for achieving high performance. Most existing approaches use either Explicit Congestion Notification (ECN) or network delay (e.g., RTT) independently as congestion indicators . However, in multi-bottleneck networks, the limitations of these signals become more pronounced: ECN struggles with large cumulative end-to-end latency, while RTT lacks the precision needed to control queuing delays at individual hops. To address these challenges, we propose Cocktail , a simple yet effective transport protocol for datacenter networks that combines both ECN and RTT congestion signals to more effectively handle multi-bottleneck scenarios. By leveraging the ECN signal, Cocktail bounds per-hop queue lengths, enhancing its ability to control single-hop latency and prevent packet loss . Additionally, by estimating RTT, Cocktail effectively manages end-to-end delay, resulting in lower Flow Completion Time (FCT). Extensive experimental evaluations in Mininet demonstrate that Cocktail significantly reduces the average and 99th-percentile completion times for small flows by up to 20% and 29%, respectively, compared to current practices under production workloads. Yuxiang Zhang 0007, Lin Cui 0001, Fung Po Tso 0001, Xiaolin Lei |
Comput. Networks | 3 |
| 2025 | Enhancing In-Network Computing Deployment via Collaboration Across PlanesabstractThe new paradigm of In-network computing (INC) permits service computation to be executed within network paths, rather than solely on dedicated servers. Although the programmable data plane has showcased notable performance advantages for INC application deployments, its effectiveness is constrained by resource limitations, potentially impeding the expressiveness and scalability of these deployments. Conversely, delegating computational tasks to the control plane, supported by general-purpose servers with abundant resources, offers increased flexibility. Nonetheless, this strategy compromises efficiency to a considerable extent, particularly when the system operates under heavy load. To simultaneously exploit the efficiency of data plane and the flexibility of control plane, we proposeCarlo, a cross-plane collaborative optimization framework to support the network-wide deployment of multiple INC applications across both the control and data plane.Carlofirst analyzes resource requirements of various INC applications across different planes. It then establishes mathematical models for resource allocation in cross-plane and automatically generates solutions using proposed algorithms. We have implemented the prototype ofCarloon Intel Tofino ASIC switches and DPDK. Experimental results demonstrate thatCarlocan effectively trade off between computation time and deployment performance while avoiding performance degradation. Xiaoquan Zhang, Lin Cui 0001, Waiming Lau, Fung Po Tso 0001, Yuhui Deng 0001, Weijia Jia 0001 |
IEEE Trans. Computers | 4 |
| 2025 | DisPLOY: Target-Constrained Distributed Deployment for Network Measurement Tasks on Data PlaneabstractIn programmable networks, measurement tasks are placed on programmable switches to monitor network traffic at line rate. These tasks typically require substantial resources (e.g., significant SRAM), while programmable switches are constrained by limited resources due to their hardware design (e.g., Tofino ASIC), making distributed deployment essentially. Measurement tasks must monitor specific network locations or traffic flows, introducing significant complexity in deployment optimization. This target-constrained nature makes task optimization on switches (e.g., task merging) become device-dependent and order-dependent, which can lead to deployment failures or performance degradation if ignored. In this paper, we introduceDisPLOY, a novel target-constrained distributed deployment framework specifically designed for network measurement tasks on the data plane.DisPLOYenables operators to specify monitoring targets—network traffic or device/link—across multiple switches. Given the monitoring targets,DisPLOYeffectively minimizes redundant operations and optimizes deployment to achieve both resource efficiency (e.g., minimizing stage consumption) and high-performance monitoring (e.g., high accuracy). We implement and evaluateDisPLOYthrough deployment on both P4 hardware switches (Intel Tofino ASIC) and BMv2. Experimental results show thatDisPLOYsignificantly reduces stage consumption by up to 66% and improves ARE by up to 78.4% in flow size estimation while maintaining end-to-end performance. Mimi Qian, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001, Zhetao Li, Weijia Jia 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Monte: SFCs Migration Scheme in the Distributed Programmable Data PlaneabstractService function chains (SFCs) are sequences of network functions that provide specific services to meet operators’ needs in today's ISPs and datacenter networks. To improve the performance of SFCs, programmable data planes are used to leverage their low latency and high performance packet processing. However, SFCs need to be adaptable to dynamics such as changes in requirements and attributes. Therefore, the ability to migrate SFCs is essential. Unfortunately, migrating SFCs in distributed programmable data planes is challenging due to the risk of degraded performance and failure to meet SFCs requirements and resource constraints in switches. In this paper, we proposeMonte, which provides an effective SFCs migration scheme in distributed programmable data planes. We build a novel integer programming model to represent the migration process with constraints on resource limitations of switches and SFCs attributes in the distributed data plane. Additionally, an SFCs migration algorithm is designed to optimize the migration cost by deeply analyzing resource allocation in the switch pipeline.Montehas been implemented on both P4 software switches (Bmv2) and hardware switches (Intel Tofino ASIC). Extensive evaluation results show that the migration cost inMonteis 94.03% lower on average than the state-of-the-art deployment scheme, andMontecan effectively save pipeline resources. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Zhetao Li, Weijia Jia 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | FlxVRM: Enabling Online Configuring Memory via Virtualization on Programmable Data PlaneabstractProgrammable data plane (PDP) has emerged as a powerful platform for line-rate packet processing, utilizing on-chip register memory to execute stateful applications. Yet most existing efforts concentrate on static approaches for allocating register memory, necessitating switch restarting and service interruption. Despite the availability of research on sharing memory for concurrent applications, the rigid requirement of limiting memory sharing to the same pipeline stages hampers application flexibility and poses scalability challenges. To address this limitation, we presentFlxVRM,a flexible register memory virtualization layerfor data plane P4 programs which supports high-flexibility sharing of register memory for concurrent applications on PDP.FlxVRMenables memory allocation at any stage and location of the pipeline on PDP for each application at run time. To reduce resource usage during virtualization in the data plane pipeline,FlxVRMfurther merges different tables and actions with similar structures within P4 programs. Additionally,FlxVRMprovides a compiler to generate data plane programs for virtualization as well as the control plane API configuration. A prototype ofFlxVRMis implemented based on P4 hardware switches with Intel Tofino ASIC. Our experiment results show thatFlxVRMsignificantly improves the allocatable memory space for applications by up to 50%, while reducing the resource of the table up to 68%. Mimi Qian, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Zhen Zhang 0017, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | Carlo: Cross-Plane Collaboration for Multiple In-network Computing ApplicationsabstractIn-network computing (INC) is a new paradigm that allows applications to be executed within the network, rather than on dedicated servers. Conventionally, INC applications have been exclusively deployed on the data plane (e.g., programmable ASICs), offering impressive performance capabilities. However, the data plane’s efficiency is hindered by limited resources, which can prevent a comprehensive deployment of applications. On the other hand, offloading compute tasks to the control plane, which is underpinned by general-purpose servers with ample resources, provides greater flexibility. However, this approach comes with the tradeoff of significantly reduced efficiency, especially when the system operates under heavy load. To simultaneously exploit the efficiency of data plane and the flexibility of control plane, we propose Carlo, a cross-plane collaborative optimization framework to support the network-wide deployment of multiple INC applications across both the control and data plane. Carlo first analyzes resource requirements of various INC applications across different planes. It then establishes mathematical models for resource allocation in cross-plane and automatically generates solutions using proposed algorithms. We have implemented the prototype of Carlo on Intel Tofino ASIC switches and DPDK. Experimental results demonstrate that Carlo can compute solutions in a short time while avoiding performance degradation caused by the deployment scheme. Xiaoquan Zhang, Lin Cui 0001, Waiming Lau, Fung Po Tso 0001, Yuhui Deng 0001, Weijia Jia 0001 |
INFOCOM | 4 |
| 2024 | DNN acceleration in vehicle edge computing with mobility-awareness: A synergistic vehicle-edge and edge-edge framework
Lin Cui 0001, Fung Po Tso 0001, Zhetao Li, Weijia Jia 0001 |
Comput. Networks | 3 |
| 2024 | OffsetINT: Achieving High Accuracy and Low Bandwidth for In-Band Network TelemetryabstractNetwork measurement is essential for efficient network management and operations. In-band network telemetry (INT) offers fine-grained per-device per-packet information which could provide full-visibility for networks. However, the existing solutions fall short in achieving high accuracy, generality, and low overhead simultaneously. To address this limitation, we introduceOffsetINTto meet these three criteria. The key idea ofOffsetINTis to use minimal bits to carry collected states during monitoring, which is based on our observation that the value of telemetry states are usually very close (e.g., the time of adjacent arrival packets) or small (e.g., only a few tens of microseconds for processing latency) for most of the time in real networks. Instead of embedding complete values of state in packets,OffsetINToptimizes bit usage by encoding an offset (using fewer bits), which is carried in-band by passing packets to the end-hosts for recovery and analysis. We theoretically derive the bounds of bandwidth mitigation forOffsetINT. We have implementedOffsetINTin both P4 hardware switches (with Intel Tofino ASIC) and BMv2. Expensive evaluation results show thatOffsetINTcan achieve an accuracy of up to 100% compared to the original INT while reducing INT bandwidth by up to 48%. Mimi Qian, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Compiling Service Function Chains via Fine-Grained Composition in the Programmable Data PlaneabstractService function chains (SFCs) are fundamental services in today's datacenters and ISP networks. Explosive volume of network traffic creates high demands for low latency and high performance. The emergence of programmable data planes has offered a new way to overcome the problem. However, limited by pipeline constraints in hardware architecture, implementing multiple network functions on programmable data planes is challenging. Besides, considering various types of network functions, e.g., stateful network functions, a general model is essential for abstracting distinct network functions. In this article, we proposepSFCwhich provides a fine-grained SFCs deployment scheme in programmable data planes. Control flow graph (CFG) is proposed to abstract and analyze various network functions. Then we model pipeline constraints in the hardware architecture using an ILP (Integer Linear Programming), and model the SFCs deployment in the substrate network as a one big switch (OBS) problem. To reduce deployment cost,pSFCfirst composes multiple SFCs to a compound CFG for eliminating redundant logics within SFCs, further decomposes the compound CFG based on the resource limitation per stage, and finally maps the OBS into the substrate network. We have implementedpSFCin both bmv2 software switch and P4 hardware switch (i.e., Intel Tofino ASIC). Evaluation results show thatpSFCreduces switch costs by 45.7% and decreases average latency by 22% without compromising throughput. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Dapper: Deploying Service Function Chains in the Programmable Data Plane Via Deep Reinforcement LearningabstractNetwork functions perform specific packet processing on network traffic. To meet operators' needs, forming service function chains (SFCs) is a fundamental technique used in today's ISPs and datacenter networks. Implementing SFCs in the programmable data plane with high throughput and low latency is a new approach to satisfy demands of ever-growing network traffic. Previous works have proposed different solutions to solve the problem, but they all inevitably have to make trade-offs between running time and performance. For example, an ILP (Integer Linear Programming) can optimize cost but suffers from long running time in large-scale network topologies. Heuristic algorithms depend strongly on manual designs and usually have a performance gap with the optimal solution. In this paper, we proposeDapper, a framework for deploying SFCs in the programmable data plane using DRL (Deep Reinforcement Learning) with graph convolutional network. In order to expand the searching space to prevent the optimal value from being missed,Dapperallows the RL (Reinforcement Learning) agent to simultaneously extract features from both the substrate network and the hardware pipeline, and exploit a graph convolutional network to enhance performance. Moreover, a mask mechanism is also designed to accelerateDapperand improve its scalability.Dapperhas been implemented and extensively evaluated on both P4 hardware switches (equipped with Intel Tofino ASIC) and software switches (i.e., bmv2). Experimental results show thatDappercan automatically generate deployment solutions in a few seconds of running time after training. They also demonstrate thatDapperreduces hardware stage usage and the latency of SFCs by up to 17.8% and 50$\sim$73% respectively on average when compared with heuristics. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Zhetao Li, Weijia Jia 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | pSFC: Fine-grained Composition of Service Function Chains in the Programmable Data PlaneabstractDynamic service function chains (SFC) are enabled by network function virtualization on general purpose servers. The emergence of programmable data planes (PDP) has offered a new way for the deployment of SFC. However, the implementation of network functions is constrained by resource limitations in PDPs (e.g., compute and memory resource). Moreover, most of existing works do not consider the optimization of state information (e.g., registers), which is essential for stateful network functions. In this paper, we propose pSFC which provides a fine-grained SFC deployment scheme in the PDP to tackle the problem. We first model network functions as control flow graphs (CFG) and the process of deployment as a one big switch (OBS) problem, and then propose an ILP (Integer Linear Programming) model for resource optimization for the OBS problem, which is NP-hard. To solve this problem efficiently, pSFC first composes multiple SFCs for eliminating redundant resources, decomposes the compound CFG based on the resource limitation per stage, and finally maps OBS into the substrate network. We have implemented pSFC in both bmv2 software switch and P4 hardware switch (i.e., Intel Tofino). Evaluation shows that pSFC reduces switch costs 45.7% and average latency 15% while providing the correctness of the process of SFC. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001 |
CCGRID | 3 |
| 2022 | Mitigating cyber threats at the network edgeabstractThe easy exploitation of IoT devices with limited security, compute and processing power has enabled hackers to carry out sophisticated attacks. Many research studies have highlighted the benefits of utilising artificial-intelligence based models in DDoS detection, but emphasis has not been placed on quantitative measurements of compute requirements for Machine Learning and Deep Learning algorithms used for DDoS detection, especially in the inference or detection stage. This research aims to fill the gap by performing quantitative measurement and comparison of various lightweight ML and DL algorithms, as well as design a lightweight collaborative framework capable of DDoS detection close to the source of the attack. Toyin Sofoluwe, Fung Po Tso 0001, Iain Phillips 0002 |
IMC | 2 |
| 2022 | B-Scale: Bottleneck-aware VNF Scaling and Flow Routing in Edge CloudsabstractWith the ever-growing demand for low-latency network applications, edge computing emerges as a new paradigm that provides computation and storage resources in close proximity to end-users. Many research efforts have resorted to network function virtualization, wherein network applications are provisioned as service function chains at edge clouds. However, due to the traffic dynamics and limited resource capacity at the network edge, how to efficiently embed service chains with latency optimization and resource efficiency remains as a challenging problem. As most existing research efforts largely overlook the bottle-necked resources of VNFs in the VNF scaling, we seek a more realistic approach to provisioning VNF instances across multiple edge clouds. Also, given the limited resources at the edge, it is of significant importance to improve the VNF utilization rate. Specifically, we formulate the VNF scaling problem as an integer linear programming (ILP) problem, aiming to minimize the end-to-end latency for service function chains. To solve this problem, we devise a novel bottleneck-aware algorithm that manages the number and deployment of newly created instances. After that, we propose an online algorithm for traffic steering to improve the utilization rates of VNF instances and avoid congestion on hotspot links. The proposed algorithm is shown to provide good performance by trace-driven simulation in real-world topologies. Chen Chen 0073, Lars Nagel 0001, Lin Cui 0001, Fung Po Tso 0001 |
ISCC | 4 |
| 2022 | Distributed federated service chaining: A scalable and cost-aware approach for multi-domain networksabstractFuture networks are expected to support cross-domain, cost-aware and fine-grained services in an efficient and flexible manner. Service Function Chaining (SFC) has been introduced as a promising approach to deliver these services. In the literature, centralized resource orchestration is usually employed to process SFC requests and manage computing and network resources. However, centralized approaches inhibit the scalability and domain autonomy in multi-domain networks. They also neglect location and hardware dependencies of service chains. In this paper, we propose Distributed Federated Service Chaining (DFSC), a framework for orchestrating and maintaining SFC placement in a distributed fashion while sharing only a minimal amount of domain information and control. First, a deployment cost minimization problem is formulated as an Integer Linear Programming (ILP) problem with fine-grained constraints for location and hardware dependencies. We show that this problem is NP-hard. Then, a placement algorithm is devised to use information only on inter-domain paths and border nodes. Our extensive experimental results demonstrate that DFSC efficiently optimizes the deployment cost, supports domain autonomy and enables faster decision-making. The results also show that DFSC finds solutions within a factor 1.15 of the optimal solution on average. Compared to a centralized approach in the literature, DFSC reduces the deployment cost by up to 20% and uses 70% less decision-making time. Chen Chen 0073, Lars Nagel 0001, Lin Cui 0001, Fung Po Tso 0001 |
Comput. Networks | 4 |
| 2022 | dDrops: Detecting silent packet drops on programmable data plane
Mimi Qian, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001 |
Comput. Networks | 4 |
| 2022 | Optimizing multipath QUIC transmission over heterogeneous paths
Hongxin Zeng, Lin Cui 0001, Fung Po Tso 0001, Zhen Zhang 0017 |
Comput. Networks | 3 |
| 2022 | Low-latency service function chain migration in edge-core networks based on open Jackson networks
Lin Cui 0001, Fung Po Tso 0001 |
J. Syst. Archit. | 3 |
| 2021 | A survey on stateful data plane in software defined networks
Xiaoquan Zhang, Lin Cui 0001, Kaimin Wei, Fung Po Tso 0001, Yangyang Ji, Weijia Jia 0001 |
Comput. Networks | 4 |
| 2021 | pHeavy: Predicting Heavy Flows in the Programmable Data PlaneabstractSince heavy flows account for a significant fraction of network traffic, being able to predict heavy flows has benefited many network management applications for mitigating link congestion, scheduling of network capacity, exposing network attacks and so on. Existing machine learning based predictors are largely implemented on the control plane of Software Defined Networking (SDN) paradigm. As a result, frequent communication between the control and data planes can cause unnecessary overhead and additional delay in decision making. In this paper, we presentpHeavy, a machine learning based scheme for predicting heavy flows directly on the programmable data plane, thus eliminating network overhead and latency to SDN controller. Considering the scarce memory and limited computation capability in the programmable data plane,pHeavyincludes a packet processing pipeline which deploys pre-trained decision tree models for in-network prediction. We have implementedpHeavyin both bmv2 software switch and P4 hardware switch (i.e., Barefoot Tofino). Evaluation results demonstrate thatpHeavyhas achieved 85% and 98% accuracy after receiving the first 5 and 20 packets of a flow respectively, while being able to reduce the size of decision tree by 5.4x on average. More importantly,pHeavycan predict heavy flows at line rate on the P4 hardware switch. Xiaoquan Zhang, Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2020 | Performance analysis of single board computer clustersabstractThe past few years have seen significant developments in Single Board Computer (SBC) hardware capabilities. These advances in SBCs translate directly into improvements in SBC clusters. In 2018 an individual SBC has more than four times the performance of a 64-node SBC cluster from 2013. This increase in performance has been accompanied by increases in energy efficiency (GFLOPS/W) and value for money (GFLOPS/$). We present systematic analysis of these metrics for three different SBC clusters composed of Raspberry Pi 3 Model B, Raspberry Pi 3 Model B+ and Odroid C2 nodes respectively. A 16-node SBC cluster can achieve up to 60 GFLOPS, running at 80 W. We believe that these improvements open new computational opportunities, whether this derives from a decrease in the physical volume required to provide a fixed amount of computation power for a portable cluster; or the amount of compute power that can be installed given a fixed budget in expendable compute scenarios. We also present a new SBC cluster construction form factor named Pi Stack; this has been designed to support edge compute applications rather than the educational use-cases favoured by previous methods. The improvements in SBC cluster performance and construction techniques mean that these SBC clusters are realising their potential as valuable developmental edge compute devices rather than just educational curiosities. Philip James Basford, Steven J. Ossont, Colin Perkins, Tony Garnock-Jones, Fung Po Tso 0001, Dimitrios P. Pezaros, Robert Mullins 0001, Eiko Yoneki, Jeremy Singer, Simon J. Cox 0001 |
Future Gener. Comput. Syst. | 5 |
| 2019 | Mobility-Aware Probabilistic Caching in UAV-Assisted Wireless D2D NetworksabstractThis paper investigates the problem of cache node placement and selection with the coexistence of unmanned aerial vehicles (UAVs) cache and device- to-device (D2D) cache in mobile networks. In recent years, caching popular content in UAV base stations has received growing interests as a promising solution to improve communication performances. With the agility and mobility features, the dynamic movement of cache-enabled UAV should be further designed to increase the cache-aided throughput. Different from the conventional caching approaches assuming ground users remain static, we consider the dynamic movement design of UAV to maximize the cache- aided throughput taking into account the movement of ground users. As the formulated optimization problem is NP-hard, we propose a mobility-aware probabilistic caching algorithm in which K-means clustering is utilized to obtain the partition of ground users. Simulation results show that the proposed algorithm notably outperforms the pure D2D cache scheme (without UAV caching) in different cases. Yu-Jia Chen, Kai-Min Liao, Meng-Lin Ku, Fung Po Tso 0001 |
GLOBECOM | 4 |
| 2019 | Autonomous Flying WiFi Access PointabstractUnmanned aerial vehicles (UAVs), aka drones, are widely used civil and commercial applications. A promising one is to use the drones as relying nodes to extend the wireless coverage. However, existing solutions only focus on deploying them to predefined locations. After that, they either remain stationary or only move in predefined trajectories throughout the whole deployment. In the open outdoor scenarios such as search and rescue or large music events, etc., users can move and cluster dynamically. As a result, network demand will change constantly over time and hence will require the drones to adapt dynamically. In this paper, we present a proof of concept implementation of an UAV access point (AP) which can dynamically reposition itself depends on the users movement on the ground. Our solution is to continuously keeping track of the received signal strength from the user devices for estimating the distance between users devices and the drone, followed by trilateration to localise them. This process is challenging because our on-site measurements show that the heterogeneity of user devices means that change of their signal strengths reacts very differently to the change of distance to the drone AP. Our initial results demonstrate that our drone is able to effectively localise users and autonomously moving to a position closer to them. Gareth J. Nunns, Yu-Jia Chen, Deng-Kai Chang, Kai-Min Liao, Fung Po Tso 0001, Lin Cui 0001 |
ISCC | 5 |
| 2019 | Extensive evaluation on the performance and behaviour of TCP congestion control protocols under varied network scenarios
Jinting Lin, Lin Cui 0001, Yuxiang Zhang 0007, Fung Po Tso 0001, Quanlong Guan |
Comput. Networks | 4 |
| 2019 | Mystique: A Fine-Grained and Transparent Congestion Control Enforcement SchemeabstractTCP congestion control is a vital component for the latency of Web services. In practice, a single congestion control mechanism is often used to handle all TCP connections on a Web server, e.g., Cubic for Linux by default. Considering complex and ever-changing networking environment, the default congestion control may not always be the most suitable one. Adjusting congestion control to meet different networking scenarios usually requires modification of TCP stacks on a server. This is difficult, if not impossible, due to various operating system and application configurations on production servers. In this paper, we propose Mystique, a light-weight, flexible, and dynamic congestion control switching scheme that allows network or server administrators to deploy any congestion control schemes transparently without modifying existing TCP stacks on servers. We have implemented Mystique in Open vSwitch (OVS) and conducted extensive test-bed experiments in both public and private cloud environments. Experiment results have demonstrated that Mystique is able to effectively adapt to varying network conditions, and can always employ the most suitable congestion control for each TCP connection. More specifically, Mystique can significantly reduce latency by 18.13% on average when compared with individual congestion controls. Yuxiang Zhang 0007, Lin Cui 0001, Fung Po Tso 0001, Quanlong Guan, Weijia Jia 0001, Jipeng Zhou |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2019 | Enabling Heterogeneous Network Function ChainingabstractToday's data center operators deploy network policies in both physical (e.g., middleboxes, switches) and virtualized (e.g., virtual machines on general purpose servers) network function boxes (NFBs), which reside in different points of the network, to exploit their efficiency and agility respectively. Nevertheless, such heterogeneity has resulted in a great number of independent network nodes that can dynamically generate and implement inconsistent and conflicting network policies, making correct policy implementation a difficult problem to solve. Since these nodes have varying capabilities, services running atop are also faced with profound performance unpredictability. In this paper, we propose a Heterogeneous netwOrk Policy Enforcement (HOPE) scheme to overcome these challenges. HOPE guarantees that network functions (NFs) that implement a policy chain are optimally placed onto heterogeneous NFBs such that the network cost of the policy is minimized. We first experimentally demonstrate that the processing capacity of NFBs is the dominant performance factor. This observation is then used to formulate the Heterogeneous Network Policy Placement problem, which is shown to be NP-Hard. To solve the problem efficiently, an online algorithm is proposed. Our experimental results demonstrate that HOPE achieves the same optimality as Branch-and-bound optimization but is 3 orders of magnitude more efficient. Lin Cui 0001, Fung Po Tso 0001, Song Guo 0001, Weijia Jia 0001, Kaimin Wei, Wei Zhao 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Live migration on ARM-based micro-datacentresabstractLive migration, underpinned by virtualisation technologies, has enabled improved manageability and fault tolerance for servers. However, virtualised server infrastructures suffer from significant processing overheads, system inconsistencies, security issues and unpredictable performance which makes them unsuitable for low-power and resource-constraint computing devices that processing latency-sensitive, “Big-data”-type data. Consequently, we ask: “How do we eliminate the overhead of virtualisation whilst still retaining its benefits?” Motivated by this question, we investigate a practical approach for a bare-metal live migration scheme for ARM-based instances low-power servers and edge devices. In this paper, we position ARM-based bare-metal live migration as a technique that will underpin the efficiency on edge-computing and on Micro-datacentres. We also introduce our early work on identifying three key technical challenges and discuss their solutions. Ilias Avramidis, Michael Mackay 0001, Fung Po Tso 0001, Takaaki Fukai, Takahiro Shinagawa |
CCNC | 3 |
| 2018 | Latency-aware joint virtual machine and policy consolidation for mobile edge computingabstractTo guarantee an efficient and high-performance environment for mobile devices to perform offloading with low end-to-end delay, it is important to ensure no network policies are violated. In this paper, we explore the simultaneous, dynamic virtual machine (VM) and policy consolidation, and formulate the Policy-VM Latency-aware Consolidation problem for Mobile Edge Computing, which is shown to be NP-Hard. We propose the PL-Edge, an efficient scheme to jointly consolidate network policies and virtual machines for mobile edge computing to reduce communication end-to-end delays among devices and virtual machines. Our simulation results demonstrate that the proposed PL-Edge can significantly reduces policy-flows end-to-end delay by nearly 45% while adhering strictly to the requirements of network policies. Thiago A. L. Genez, Fung Po Tso 0001, Lin Cui 0001 |
CCNC | 2 |
| 2018 | Dynamic Network Function Chain Composition for Mitigating Network LatencyabstractNetwork Function Virtualisation (NFV) enables rapid deployment of new services in networks on an on-demand basis using general purpose servers. Multiple virtual network functions (VNFs) can be dynamically chained in an ordered sequence for the delivery of end-to-end services. Nevertheless, network latency caused by the sequential order of packet processing on every VNF can hurt the performance of latency-sensitive applications. To reduce such network latency, existing solutions only consider the maximum capacity of individual virtual network functions (VNFs) and do not take into account the fact that performance of VNFs, as with any software applications, is bottlenecked by either CPU or I/O peripheral capacity of the server they run on and their underneath implementation such as singleor multi-threaded.By exploiting this knowledge, we can better determine the number of required VNF instances and distribute the network traffic among them for any given VNF chain. In this paper, we formulate the VNF Scaling and Traffic Distribution problem and prove that it is NP-hard. We then present the design and implementation of Natif, an efficient VNF-Aware VNF insTantIation and traFfic distribution scheme. Through our OpenStack-based testbed evaluations, we demonstrate that Natif can significantly improve the network latency by 188% on average as compared to other approaches. As a chain composition scheme, Natif can effectively work with any VNF chaining algorithms. Wajdi Hajji, Thiago A. L. Genez, Fung Po Tso 0001, Lin Cui 0001, Iain Phillips 0002 |
ISCC | 3 |
| 2018 | Modest BBR: Enabling Better Fairness for BBR Congestion ControlabstractAs a vital component of TCP, congestion control defines TCP's performance characteristics. Hence, it is important for congestion control to provide high link utilization and low queuing delay. Recent BBR tries to estimate available bottleneck capacity to achieve this goal. However, its aggressiveness characteristics generate a massive amount of packet retransmission which harms loss-based congestion control protocol such as Cubic. In this paper, we first dive into this issue and reveal that the aggressiveness of BBR can degrade the performance of Cubic, as well as the overall Internet transmission. Then we present Modest BBR, a simple yet effective solution based on BBR, by responding to retransmission less aggressively. Through extensive testbed experiments and Mininet simulation, we show Modest BBR can preserve high throughput and short convergence time while improve the overall performance when coexisting with Cubic. For example, Modest BBR gets similar throughput compared to BBR, while it improves 7.1% of the overall throughput and achieves better fairness to loss-based schemes. Yuxiang Zhang 0007, Lin Cui 0001, Fung Po Tso 0001 |
ISCC | 3 |
| 2018 | Next generation single board clustersabstractUntil recently, cluster computing was too expensive and too complex for commodity users. However the phenomenal popularity of single board computers like the Raspberry Pi has caused the emergence of the single board computer cluster. This demonstration will present a cheap, practical and portable Raspberry Pi cluster called Pi Stack. We will show pragmatic custom solutions to hardware issues, such as power distribution, and software issues, such as remote updating. We also sketch potential use cases for Pi Stack and other commodity single board computer cluster architectures. Jeremy Singer, Herry Herry, Philip James Basford, Wajdi Hajji, Colin Perkins, Fung Po Tso 0001, Dimitrios P. Pezaros, Robert Mullins 0001, Eiko Yoneki, Simon J. Cox 0001, Steven J. Ossont |
NOMS | 6 |
| 2018 | Enforcing network policy in heterogeneous network function box environment
Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
Comput. Networks | 2 |
| 2018 | Commodity single board computer clusters and their applicationsabstractCurrent commodity Single Board Computers (SBCs) are sufficiently powerful to run mainstream operating systems and workloads. Many of these boards may be linked together, to create small, low-cost clusters that replicate some features of large data center clusters. The Raspberry Pi Foundation produces a series of SBCs with a price/performance ratio that makes SBC clusters viable, perhaps even expendable. These clusters are an enabler for Edge/Fog Compute, where processing is pushed out towards data sources, reducing bandwidth requirements and decentralizing the architecture. In this paper we investigate use cases driving the growth of SBC clusters, we examine the trends in future hardware developments, and discuss the potential of SBC clusters as a disruptive technology. Compared to traditional clusters, SBC clusters have a reduced footprint, are low-cost, and have low power requirements. This enables different models of deployment—particularly outside traditional data center environments. We discuss the applicability of existing software and management infrastructure to support exotic deployment scenarios and anticipate the next generation of SBC. We conclude that the SBC cluster is a new and distinct computational deployment paradigm, which is applicable to a wider range of scenarios than current clusters. It facilitates Internet of Things and Smart City systems and is potentially a game changer in pushing application logic out towards the network edge. Steven J. Ossont, Philip James Basford, Colin Perkins, Herry Herry, Fung Po Tso 0001, Dimitrios P. Pezaros, Robert Mullins 0001, Eiko Yoneki, Simon J. Cox 0001, Jeremy Singer |
Future Gener. Comput. Syst. | 5 |
| 2017 | Heterogeneous NetwOrk Policy Enforcement in data centersabstractWith the emergence of network function virtualization, data center start to deploy a variety of network function boxes (NFBs) in both physical and virtual form factors in order to combines inherent efficiency offered by physical NFBs with the agility and flexibility of virtual ones. However, existing schemes are limited to exclusively consider physical or virtual NFBs, which may reduce the performance efficiency of services running atop. In this paper, we propose a Heterogeneous NetwOrk Policy Enforcement scheme (HOPE) to overcome these challenges. An efficient algorithm that can closely approximate optimal latency-wise NF service chaining is proposed. The experimental results have also shown that HOPE can outperform greedy algorithm by 25% in terms of network latency and is 56× more efficient than naive depth-first search algorithm. Lin Cui 0001, Fung Po Tso 0001, Weijia Jia 0001 |
IM | 2 |
| 2017 | Experimental evaluation of SDN-controlled, joint consolidation of policies and virtual machinesabstractMiddleboxes (MBs) are ubiquitous in modern data centre (DC) due to their crucial role in implementing network security, management and optimisation. In order to meet network policy's requirement on correct traversal of an ordered sequence of MBs, network administrators rely on static policy based routing or VLAN stitching to steer traffic flows. However, dynamic virtual server migration in virtual environment has greatly challenged such static traffic steering. In this paper, we design and implement Sync, an efficient and synergistic scheme to jointly consolidate network policies and virtual machines (VMs), in a readily deployable Mininet environment. We present the architecture of Sync framework and open source its code. We also extensively evaluate Sync over diverse workload and policies. Our results show that in an emulated DC of 686 servers, 10k VMs, 8k policies, and 100k flows, Sync processes a group of 900 VMs and 10 VMs in 634 seconds and 4 seconds respectively. Wajdi Hajji, Fung Po Tso 0001, Lin Cui 0001, Dimitrios P. Pezaros |
ISCC | 2 |
| 2017 | TCon: A Transparent Congestion Control Deployment Platform for Optimizing WAN Transfers
Yuxiang Zhang 0007, Lin Cui 0001, Fung Po Tso 0001, Quanlong Guan, Weijia Jia 0001 |
NPC | 3 |
| 2017 | Modelling Low Power Compute Clusters for Cloud SimulationabstractIn order to minimise their energy use, data centre operators are constantly exploring new ways to construct computing infrastructures. As low power CPUs, exemplified by ARM-based devices, are becoming increasingly popular, there is a growing trend for the large scale deployment of low power servers in data centres. For example, recent research has shown promising results on constructing small scale data centres using Raspberry Pi (RPi) single-board computers as their building blocks. To enable larger scale experimentation and feasibility studies, cloud simulators could be utilised. Unfortunately, state-of-the-art simulators often need significant modification to include such low power devices as core data centre components. In this paper, we introduce models and extensions to estimate the behaviour of these new components in the DISSECT-CF cloud computing simulator. We show that how a RPi based cloud could be simulated with the use of the new models. We evaluate the precision and behaviour of the implemented models using a Hadoop-based application scenario executed both in real life and simulated clouds. Gabor Kecskemeti, Wajdi Hajji, Fung Po Tso 0001 |
PDP | 3 |
| 2017 | Machine learning approaches to the application of disease modifying therapy for sickle cell using classification models
Mohammed Khalaf 0001, Abir Jaafar Hussain, Robert Keight, Dhiya Al-Jumeily, Paul Fergus, Russell Keenan, Fung Po Tso 0001 |
Neurocomputing | 7 |
| 2017 | PLAN: Joint Policy- and Network-Aware VM Management for Cloud Data CentersabstractPolicies play an important role in network configuration and therefore in offering secure and high performance services especially over multi-tenant Cloud Data Center (DC) environments. At the same time, elastic resource provisioning through virtualization often disregards policy requirements, assuming that the policy implementation is handled by the underlying network infrastructure. This can result in policy violations, performance degradation and security vulnerabilities. In this paper, we define PLAN, a PoLicy-Aware and Network-aware VM management scheme to jointly consider DC communication cost reduction through Virtual Machine (VM) migration while meeting network policy requirements. We show that the problem is NP-hard and derive an efficient approximate algorithm to reduce communication cost while adhering to policy constraints. Through extensive evaluation, we show that PLAN can reduce topology-wide communication cost by 38 percent over diverse aggregate traffic and configuration policies. Lin Cui 0001, Fung Po Tso 0001, Dimitrios P. Pezaros, Weijia Jia 0001, Wei Zhao 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Build Trust in the Cloud Computing - Isolation in Container Based VirtualisationabstractCloud computing is revolutionizing many IT ecosystems through offering scalable computing resources that are easy to configure, use and inter-connect. However, this model has always been viewed with some suspicion as it raises a wide range of security and privacy issues that need to be negotiated. This research focuses on the construction of a trust layer in cloud computing to build a trust relationship between cloud service providers and cloud users. In particular, we address the rise of container-based virtualisation has a weak isolation compared to traditional VMs because of the shared use of the OS kernel and system components. Therefore, we will build a trust layer to solve the issues of weaker isolation whilst maintaining the performance and scalability of the approach. This paper has two objectives. Firstly, we propose a security system to protect containers from other guests through the addition of a Role-based Access Control (RBAC) model and the provision of strict data protection and security. Secondly, we provide a stress test using isolation benchmarking tools to evaluate the isolation in containers in term of performance. Ibrahim Alobaidan, Michael Mackay 0001, Fung Po Tso 0001 |
DeSE | 3 |
| 2016 | Synergistic policy and virtual machine consolidation in cloud data centersabstractIn modern Cloud Data Centers (DC)s, correct implementation of network policies is crucial to provide secure, efficient and high performance services for tenants. It is reported that the inefficient management of network policies accounts for 78% of DC downtime, challenged by the dynamically changing network characteristics and by the effects of dynamic Virtual Machine (VM) consolidation. While there has been significant research in policy and VM management, they have so far been treated as disjoint research problems. In this paper, we explore the simultaneous, dynamic VM and policy consolidation, and formulate the Policy-VM Consolidation (PVC) problem, which is shown to be NP-Hard. We then propose Sync, an efficient and synergistic scheme to jointly consolidate network policies and virtual machines. Extensive evaluation results and a testbed implementation of our controller show that policy and VM migration under Sync significantly reduces flow end-to-end delay by nearly 40%, and network-wide communication cost by 50% within few seconds, while adhering strictly to the requirements of network policies. Lin Cui 0001, Richard Cziva, Fung Po Tso 0001, Dimitrios P. Pezaros |
INFOCOM | 3 |
| 2016 | Network and server resource management strategies for data centre infrastructures: A surveyabstractThe advent of virtualisation and the increasing demand for outsourced, elastic compute charged on a pay-as-you-use basis has stimulated the development of large-scale Cloud Data Centres (DCs) housing tens of thousands of computer clusters. Of the significant capital outlay required for building and operating such infrastructures, server and network equipment account for 45 and 15% of the total cost, respectively, making resource utilisation efficiency paramount in order to increase the operators’ Return-on-Investment (RoI). In this paper, we present an extensive survey on the management of server and network resources over virtualised Cloud DC infrastructures, highlighting key concepts and results, and critically discussing their limitations and implications for future research opportunities. We highlight the need for and benefits of adaptive resource provisioning that alleviates reliance on static utilisation prediction models and exploits direct measurement of resource utilisation on servers and network nodes. Coupling such distributed measurement with logically centralised Software Defined Networking (SDN) principles, we subsequently discuss the challenges and opportunities for converged resource management over converged ICT environments, through unifying control loops to globally orchestrate adaptive and load-sensitive resource provisioning. Fung Po Tso 0001, Simon Jouet, Dimitrios P. Pezaros |
Comput. Networks | 1 |
| 2016 | SDN-Based Virtual Machine Management for Cloud Data CentersabstractSoftware-defined networking (SDN) is an emerging paradigm to logically centralize the network control plane and automate the configuration of individual network elements. At the same time, in cloud data centers (DCs), although network and server resources are collocated and managed by a single administrative entity, disjoint control mechanisms are used for their respective management. In this paper, we propose a unified server-network resource management for such converged information and communication technology (ICT) environments. We present a SDN-based orchestration framework for live virtual machine (VM) management that exploits temporal network information to migrate VMs and minimize the network-wide communication cost of the resulting traffic dynamics. A prototype implementation is presented, and a cloud DC testbed is used to evaluate the impact of diverse orchestration algorithms. Our live VM management has been shown to reduce the network-wide communication cost, especially for the high-cost and congestion-prone core and aggregation layers of the DC. Our results show an increase in network-wide throughput by over six times, as well as over 70% communication cost reduction by migrating less than 50% of the VMs. Richard Cziva, Simon Jouet, David Stapleton, Fung Po Tso 0001, Dimitrios P. Pezaros |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2015 | Policy-Aware Virtual Machine Management in Data Center NetworksabstractPolicies play an important role in network configuration and, therefore, in offering secure and high performance services, especially over multi-tenant Cloud Data Center (DC) environments. At the same time, elastic resource provisioning through virtualization often disregards policy requirements, assuming that the policy implementation is handled by the underlying network infrastructure. In this paper, we define PLAN, a Policy-Aware virtual machine management scheme to jointly consider DC communication cost reduction through Virtual Machine (VM) migration while meeting network policy requirements. Lin Cui 0001, Fung Po Tso 0001, Dimitrios P. Pezaros, Weijia Jia 0001, Wei Zhao 0001 |
ICDCS | 2 |
| 2014 | Scalable Traffic-Aware Virtual Machine Management for Cloud Data CentersabstractVirtual Machine (VM) management is a powerful mechanism for providing elastic services over Cloud Data Centers (DC)s. At the same time, the resulting network congestion has been repeatedly reported as the main bottleneck in DCs, even when the overall resource utilization of the infrastructure remains low. However, most current VM management strategies are traffic-agnostic, while the few that are traffic-aware only concern a static initial allocation, ignore bandwidth oversubscription, or do not scale. In this paper we present S-CORE, a scalable VM migration algorithm to dynamically reallocate VMs to servers while minimizing the overall communication footprint of active traffic flows. We formulate the aggregate VM communication as an optimization problem and we then define a novel distributed migration scheme that iteratively adapts to dynamic traffic changes. Through extensive simulation and implementation results, we show that S-CORE achieves significant (up to 87%) communication cost reduction while incurring minimal overhead and downtime. Fung Po Tso 0001, Eleni Kavvadia, Dimitrios P. Pezaros |
ICDCS | 1 |
| 2013 | Implementing Scalable, Network-Aware Virtual Machine Migration for Cloud Data CentersabstractVirtualization has been key to the success of Cloud Computing through the on-demand allocation of shared hardware resources to Virtual Machines (VM)s. However, the network-agnostic placement of VMs over the underlying network topology can itself be a factor of performance degradation by causing congestion at the core layers of the infrastructure where bandwidth is heavily oversubscribed. In this paper, we design and implement S-CORE, a scalable live VM migration scheme to dynamically reallocate VMs to servers while minimizing the overall communication footprint of active traffic flows. We evaluate S- CORE over diverse aggregate load and coordination policies. Our results show that it can achieve up to a 87% communication cost reduction with a limited number of migration rounds, and can be easily accommodated within commodity hardware and hypervisor architectures. The associated memory, CPU, and network overhead are also minimum under typical Cloud Data Center workloads. Fung Po Tso 0001, Gregg Hamilton, Dimitrios P. Pezaros |
IEEE CLOUD | 1 |
| 2013 | Longer Is Better: Exploiting Path Diversity in Data Center NetworksabstractData Center (DC) networks exhibit much more centralized characteristics than the legacy Internet, yet they are operated by similar distributed routing and control algorithms that fail to exploit topological redundancy to deliver better and more sustainable performance. Multipath protocols, for example, use node-local and heuristic information to only exploit path diversity between shortest paths. In this paper, we use a measurement-based approach to schedule flows over both shortest and non-shortest paths based on temporal network-wide utilization. We present the Baatdaat flow scheduling algorithm which uses spare DC network capacity to mitigate the performance degradation of heavily utilized links. Results show that Baatdaat achieves close to optimal Traffic Engineering by reducing network-wide maximum link utilization by up to 18% over Equal-Cost Multi-Path (ECMP) routing, while at the same time improving flow completion time by 41% - 95%. Fung Po Tso 0001, Gregg Hamilton, Rene Weber, Colin Perkins, Dimitrios P. Pezaros |
ICDCS | 1 |
| 2013 | Baatdaat: Measurement-based flow scheduling for cloud data centersabstractSoftware-Defined Networking (SDN) allows for efficient network-wide Traffic Engineering through the logical centralization of the control plane over individual switches that perform packet forwarding independently. Such abstraction is particularly suitable for Data Center (DC) networks that need to react to fluctuating traffic dynamics over short timescales. In this paper, we propose a low-cost, SDN-based system that exposes the temporal network-wide utilization through direct measurement, rather than estimation. We then present the Baatdaat1flow scheduling algorithm which uses spare DC network capacity to mitigate the performance degradation of heavily utilized links. Results show that Baatdaat achieves close to optimal Traffic Engineering by reducing network-wide maximum link utilization by up to 18% over ECMP, while at the same time improving flow completion time by as much as 41% - 95% for different types of flows. Fung Po Tso 0001, Dimitrios P. Pezaros |
ISCC | 1 |
| 2013 | Blind detection of spread spectrum flow watermarksabstractABSTRACT Recently, the direct sequence spread spectrum (DSSS)‐based technique has been proposed to trace anonymous network flows. In this technique, homogeneous pseudo‐noise (PN) codes are used to modulate multiple bit signals that are embedded into the target flow as watermarks. This technique could be maliciously used to degrade an anonymous communication network. In this paper, we propose an effective single flow‐based scheme to detect the existence of these watermarks. Our investigation shows that, even if we have no knowledge of the applied PN code, we are still able to detect malicious DSSS watermarks via mean‐square autocorrelation (MSAC) of a single modulated flow's traffic rate time series. MSAC shows periodic peaks because of self‐similarity in the modulated traffic caused by homogeneous PN codes that are used in modulating multiple bit signals. Our scheme has low complexity and does not require any PN code synchronization. We evaluate this detection scheme's effectiveness via simulations. Our results demonstrate a high detection rate with a low false positive rate. Real‐world experiments on Tor also validate the feasibility of the detection scheme. Our scheme is more flexible and accurate than the existing multiflow‐based approach in DSSS watermark detection. We also present a theory for reconstructing the DSSS code once the DSSS code length is known and simulations validate the feasibility. Copyright © 2012 John Wiley & Sons, Ltd. Weijia Jia 0001, Fung Po Tso 0001, Zhen Ling 0001, Xinwen Fu, Dong Xuan, Wei Yu 0002 |
Secur. Commun. Networks | 2 |
| 2013 | DragonNet: A Robust Mobile Internet Service System for Long-Distance TrainsabstractAbstract—Wide range wireless networks often suffer from annoying service deterioration due to fickle wireless environment. This is especially the case with passengers on long distance train (LDT) to connect onto the Internet. To improve the service quality of wide range wireless networks, we present the DragonNet protocol with its implementation. The DragonNet system is a chained gateway which consists of a group of interlinked DragonNet routers working specifically for mobile chain transport systems. The protocol makes use of the spatial diversity of wireless signals that not all spots on a surface see the same level of radio frequency radiation. In the case of a LDT of around 500 meters, it is highly possible that some of the spanning routers still see sound signal quality, when the LDT is partially blocked from wireless Internet. DragonNet protocol fully utilizes this feature to amortize single point router failure over the whole router chain by intelligently rerouting traffics on failed ones to sound ones. We have implemented the DragonNet system and tested it in real railways over a period of three months. Our results have pinpointed two fundamental contributions of DragonNet protocol. First, DragonNet significantly reduces average temporary communication blackout (i.e. no Internet connection) to 1.5 seconds compared with 6 seconds that without DragonNet protocol. Second, DragonNet efficiently doubles the aggregate throughput on average. Fung Po Tso 0001, Lin Cui 0001, Lizhuo Zhang, Weijia Jia 0001, Di Yao 0006, Jin Teng, Dong Xuan |
IEEE Trans. Mob. Comput. | 1 |
| 2013 | Improving Data Center Network Utilization Using Near-Optimal Traffic EngineeringabstractEqual cost multiple path (ECMP) forwarding is the most prevalent multipath routing used in data center (DC) networks today. However, it fails to exploit increased path diversity that can be provided by traffic engineering techniques through the assignment of nonuniform link weights to optimize network resource usage. To this extent, constructing a routing algorithm that provides path diversity over nonuniform link weights (i.e., unequal cost links), simplicity in path discovery and optimality in minimizing maximum link utilization (MLU) is nontrivial. In this paper, we have implemented and evaluated the Penalizing Exponential Flow-spliTing (PEFT) algorithm in a cloud DC environment based on two dominant topologies, canonical and fat tree. In addition, we have proposed a new cloud DC topology which, with only a marginal modification of the current canonical tree DC architecture, can further reduce MLU and increase overall network capacity utilization through PEFT routing. Fung Po Tso 0001, Dimitrios P. Pezaros |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | User-level data center tomographyabstractMeasurement and inference in data centers present a set of opportunities and challenges distinct from the Internet domain. Existing toolsets may be perturbed or be mislead by issues related to virtualization. Yet, while equally confronted by scale, data centers are relatively homogenous and symmetric. We believe these may be attributes to be exploited. However, data is required to better evaluate our hypotheses. Therefore, we introduce our efforts to gather data using a single framework from which we can launch tests of our choosing. Our observations reinforce recent claims, but indicate changes in the network. They also reveal additional obfuscations stemming from virtualization. Neil Alexander Twigg, Marwan Fayed, Colin Perkins, Dimitrios P. Pezaros, Fung Po Tso 0001 |
SIGCOMM | 5 |
| 2012 | Mobility: A Double-Edged Sword for HSPA Networks: A Large-Scale Test on Hong Kong Mobile HSPA NetworksabstractThis paper presents an empirical study on the performance of mobile High Speed Packet Access (a 3.5G cellular standard usually abbreviated as HSPA) networks in Hong Kong via extensive field tests. Our study, from the viewpoint of end users, covers virtually all possible mobile scenarios in urban areas, including subways, trains, off-shore ferries, and city buses. We have confirmed that mobility has largely negative impacts on the performance of HSPA networks, as fast-changing wireless environment causes serious service deterioration or even interruption. Meanwhile, our field experiment results have shown unexpected new findings and thereby exposed new features of the mobile HSPA networks, which contradict commonly held views. We surprisingly find out that mobility can improve fairness of bandwidth sharing among users and traffic flows. Also, the triggering and final results of handoffs in mobile HSPA networks are unpredictable and often inappropriate, thus calling for fast reacting fallover mechanisms. Moreover, we find that throughput performance does not monotonically decrease with increased mobility level. We have conducted in-depth research to furnish detailed analysis and explanations to what we have observed. We conclude that mobility is a double-edged sword for HSPA networks. To the best of our knowledge, this is the first public report on a large-scale empirical study on the performance of commercial mobile HSPA networks. Fung Po Tso 0001, Jin Teng, Weijia Jia 0001, Dong Xuan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | DragonNet: A robust mobile Internet service system for long distance trainsabstractWide range wireless networks often suffer from annoying service deterioration due to fickle wireless environment. This is especially the case with passengers on long distance train (LDT) to connect onto the Internet. To improve the service quality of wide range wireless networks, we present the DragonNet protocol with its implementation. The DragonNet system is a chained gateway which consists of a group of interlinked DragonNet routers working specifically for mobile chain transport systems. The protocol makes use of the spatial diversity of wireless signals that not all spots on a surface see the same level of radio frequency radiation. In the case of a LDT of around 500 meters, it is highly possible that some of the spanning routers still see sound signal quality, when the LDT is partially blocked from wireless Internet. DragonNet protocol fully utilizes this feature to amortize single point router failure over the whole router chain by intelligently rerouting traffics on failed ones to sound ones. We have implemented the DragonNet system and tested it in real railways over a period of three months. Our results have pinpointed two fundamental contributions of DragonNet protocol. First, DragonNet significantly reduces average temporary communication blackout (i.e. no Internet connection) to 1.5 seconds compared with 6 seconds that without DragonNet protocol. Second, DragonNet efficiently doubles the aggregate throughput on average. Fung Po Tso 0001, Lin Cui 0001, Lizhuo Zhang, Weijia Jia 0001, Di Yao 0006, Jin Teng, Dong Xuan |
INFOCOM | 1 |
| 2010 | Mobility: a double-edged sword for HSPA networks: a large-scale test on Hong Kong mobile HSPA networksabstractThis paper presents an empirical study on the performance of mobile High Speed Packet Access (HSPA, a 3.5G cellular standard) networks in Hong Kong via extensive field tests. Our study, from the viewpoint of end users, covers virtually all possible mobile scenarios in urban areas, including subways, trains, off-shore ferries and city buses. We have confirmed that mobility has largely negative impacts on the performance of HSPA networks, as fast-changing wireless environment causes serious service deterioration or even interruption. Meanwhile our field experiment results have shown unexpected new findings and thereby exposed new features of the mobile HSPA networks, which contradict commonly held views. We surprisingly find out that mobility can improve fairness of bandwidth sharing among users and traffic flows. Also the triggering and final results of handoffs in mobile HSPA networks are unpredictable and often inappropriate, thus calling for fast reacting fallover mechanisms. We have conducted in-depth research to furnish detailed analysis and explanations to what we have observed. We conclude that mobility is a double-edged sword for HSPA networks. To the best of our knowledge, this is the first public report on a large scale empirical study on the performance of commercial mobile HSPA networks. Fung Po Tso 0001, Jin Teng, Weijia Jia 0001, Dong Xuan |
MobiHoc | 1 |
| 2009 | Blind Detection of Spread Spectrum Flow WatermarksabstractRecently, the direct sequence spread-spectrum (DSSS)-based technique has been proposed to trace anonymous network flows. In this technique, homogeneous pseudo-noise (PN) codes are used to modulate multiple-bit signals that are embedded into the target flow as watermarks. This technique could be maliciously used to degrade an anonymous communication network. In this paper, we propose a simple single flow-based scheme to detect the existence of these watermarks. Our investigation shows that even if we have no knowledge of the applied PN code, we are still able to detect malicious DSSS watermarks via mean-square autocorrelation (MSAC) of a single modulated flow's traffic rate time series. MSAC shows periodic peaks due to self-similarity in the modulated traffic caused by homogeneous PN codes that are used in modulating multiple-bit signals. Our scheme has low complexity and does not require any PN-code synchronization. We evaluate this detection scheme's effectiveness via simulations and real-world experiments on Tor. Our results demonstrate a high detection rate with a low false positive rate. Our scheme is more flexible and accurate than an existing multi-flow-based approach in DSSS watermark detection. Weijia Jia 0001, Fung Po Tso 0001, Zhen Ling 0001, Xinwen Fu, Dong Xuan, Wei Yu 0002 |
INFOCOM | 2 |
| 2008 | Toward ubiquitous Video-based Cyber-Physical SystemsabstractCyber-physical systems (CPS) is a new generation of engineered systems that integrate physical systems with the capability of networked computing and control. Real-time video capture and communication is expected to be an important function in many cyber-physical systems that involve camera-equipped mobile phones. In this paper, we present AnySense, a network architecture that supports video communication between 3G phones and Internet hosts in cyber-physical systems. AnySense implements transcoding of video streams between the Internet and circuit-switched 3G cellular networks, and is transparent to 3G service providers. AnySense can support a class of ubiquitous cyber-physical systems that require video-based information collection and sharing. A prototype of AnySense has been built and a video demo is available at http://www.anyserver.org/. Guoliang Xing, Weijia Jia 0001, Yufei Du, Fung Po Tso 0001, Mo Sha 0001, Xue (Steve) Liu |
SMC | 4 |
| 2007 | Video surveillance patrol robot system in 3G, Internet and sensor networksabstractWe propose to demo a ubiquitous surveillance patrol robot system which can patrol in a candidate site to perform events detection where a wireless sensor network may be deployed. We have enabled the 3G phone controlled patrol robot (over 3G circuit switched network) with integrated access to the WiFi/Internet. Internet is used to provide sensor query, to send control signal to the robot and to request the real time audiovisual data from the robot. The robot can receive the movement instructions from and pull the real-time multimedia data stream to a remote user via WiFi laptop or 3G terminal. We also implemented a gateway which is a key component for the platform in responsible for the interconnection and heterogeneous communication of the networks. Fung Po Tso 0001, Lizhuo Zhang, Weijia Jia 0001 |
SenSys | 1 |
| 2006 | Performance Evaluation of Scheduling in IEEE 802.16 Based Wireless Mesh NetworksabstractWe propose an efficient centralized scheduling algorithm in IEEE 802.16 based wireless mesh networks (WMN) to provide high qualified wireless multimedia services. Our algorithm takes special attention on the relay function of the mesh nodes in a transmission tree which is seldom studied in previous research. Some important design metrics, such as fairness, channel utilization and transmission delay are considered in this scheduling algorithm. IEEE 802.16 employs TDMA and the selection policy for scheduled links in a time slot will definitely impact the system performance. We evaluated the proposed algorithm with four selection criteria through extensive simulations and the results are instrumental for improving the performance of IEEE 802.16 based WMNs in terms of link scheduling Bo Han 0001, Fung Po Tso 0001, Lidong Lin, Weijia Jia 0001 |
MASS | 2 |