VLDB 2026 Research / reviewers in the wild / expert
Tao Huang 0005
dblp:34/808-5
· DBLP profile ↗
209ranked-venue papers
9as first author
149since 2021 · last 2026
0000-0002-3545-1122ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 175 · 6 first-author · 130 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 3 since 2021Systems, architecture and hardware · 8 · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Argus: Scalable and Deterministic Network Fault Localization for AI Training ClustersabstractNetwork failures in Artificial Intelligence (AI) training clusters can degrade entire jobs, making fast and accurate fault localization critical. Existing active probing systems suffer from two fundamental limitations: probabilistic path coverage that cannot guarantee complete link observability, and binary anomaly detection that fails to distinguish concurrent failures or localize gray failures. Equal-Cost Multi-Path (ECMP) routing is deterministic given the same 5-tuple, and ECMP configurations are accessible in operator-controlled clusters. We exploit this property to derive exact probe paths through offline hash computation without network measurement. Based on this approach, we design a hash-aware probing system that constructs a deterministic coverage matrix to select minimal probes guaranteeing complete link coverage. We introduce edge signatures to ensure fault distinguishability and a tiered diagnosis approach where lightweight iterative localization handles hard failures while sparse regression localizes gray failures. Preliminary evaluations on fat-tree topologies with up to 10,240 hosts show that our system achieves 100% link coverage with 39× fewer probes than R-Pingmesh, F1 score of 0.75–0.92 for multi-link failures, and 0.67 F1 for gray failures where existing methods fail entirely. Yuxiang Wang 0011, Jiao Zhang 0002, Xianyu Huang, Yubo Ruan, Yingjie Duan, Shoushou Ren, Xianjun He, Tao Huang 0005 |
APNet | 11 |
| 2026 | Lossless-SR: Towards Non-Disruptive Source Routing for Topology-Varying LEO Satellite Networks
Tian Pan 0001, Guohao Ruan, Zijia Xu, Yuehui Tan, Jiao Zhang 0002, Tao Huang 0005 |
ICC | 9 |
| 2026 | Tlaloc: A Generic Multipath Load Balancing for RoCE
Huimin Luo, Jiao Zhang 0002, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
INFOCOM | 5 |
| 2026 | Zephyr: Switch-Aided Weighted Congestion Control for Differentiated Inference Workloads
Zijia Xu, Tian Pan 0001, Chenlin Ge, Hao Ouyang, Guohao Ruan, Tao Huang 0005 |
IWQoS | 8 |
| 2026 | CStar Gateway: Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Tian Pan 0001, Jin Ke 0005, Baohai Hu, Changgang Zheng, Enge Song, Donglin Lai, Yisong Qiao, Bengbeng Xue, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Yang Song 0031, Xionglie Wei, Biao Lyu, Rong Wen, Zhigang Zong, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu |
NSDI | 25 |
| 2026 | Euler: An Out-of-Order-Aware Load Balancing with Adaptive Granularity for AI Clusters
Jiafeng Jiang, Jiao Zhang 0002, Huimin Luo, Shuo Wang 0006, Tao Huang 0005 |
WCNC | 5 |
| 2026 | Security-aware online task offloading for edge computing based on deterministic networking
Weiqian Tan, Binwei Wu, Shuo Wang 0006, Tao Huang 0005 |
Comput. Commun. | 4 |
| 2026 | Flex-LDN: Toward Flexible Resource Reservation in Large-Scale Deterministic NetworkingabstractThe rapid growth of large-scale Industrial Internet of Things (IIoT) drives the deployment of time-sensitive applications requiring stringent deterministic latency guarantees and high-throughput communication. Consequently, Large-Scale Deterministic Networking (LDN) and its variant, Advanced-LDN (A-LDN), have been proposed to achieve deterministic transmission by partitioning the timeline into fixed-length cycles and allocating sufficient resources for time-sensitive (TS) flows within each cycle. To reduce the complexity of resource reservation decisions, LDN adopts a uniform resource reservation pattern (RRP) that allocates identical resources in each cycle, while A-LDN uses a periodic RRP that reserves resources at fixed cycle intervals. However, these simplified RRPs lead to excessive over-provisioning and fail to meet the high-throughput demands of IIoT applications. To address this, we propose a flexible LDN (Flex-LDN) mechanism that supports arbitrary RRPs and minimizes resource reservation for TS flows. First, due to the lack of an analytical method to establish upper bounds of shaping delays for arbitrary RRPs, Flex-LDN constructs a service curve model and calculates delay bounds using network calculus. Furthermore, since shaping delay budgets vary for different paths of a TS flow, the required resource reservation to ensure deterministic latency and jitter bounds also differs. Therefore, we design a load-balancing-based RRP decision algorithm and a transmission strategy space generation algorithm to assign RRPs to each path, thereby constructing a transmission strategy space that consists of RRPs and paths for each TS flow. The RRPs generated by these algorithms minimize resource reservation while satisfying the shaping delay budgets. This approach ensures minimal resource allocation during TS flow scheduling. Finally, to efficiently solve the TS flow scheduling problem, we model it as an Exact Potential Game (EPG) and propose a distributed Ordered Best-Response (OBR) algorithm that can converge to a Nash equilibrium in polynomial time. Simulation results show that Flex-LDN improves the TS flow scheduling success ratio by 62.2% and 24.5% over LDN and A-LDN, respectively, while reducing scheduling time by 72.5% compared to A-LDN. Weiqian Tan, Binwei Wu, Shuo Wang 0006, Tao Huang 0005 |
IEEE Internet Things J. | 4 |
| 2026 | STSR: A Satellite-Tailored Segment Routing Method for Satellite-Terrestrial Integrated NetworkabstractSegment Routing (SR) provides an effective approach to path control with minimal control-plane signaling. This makes SR a strong candidate for routing in the Satellite-Terrestrial Integrated Network (STIN), the core infrastructure enabling ubiquitous Internet of Things (IoT) connectivity. However, directly applying existing SR solutions to satellite networks presents significant challenges, which include limited bandwidth and constrained onboard processing capabilities, hindering efficient IoT data transmission. To enable reliable and efficient IoT services, we propose Satellite-Tailored Segment Routing (STSR), a novel framework designed specifically for satellite networks in STIN. STSR is built as a lightweight and Segment Routing over IPv6 (SRv6)-compatible extension. It deploys a customized data plane that enables efficient source routing through bit-based encoding and streamlined processing, which is vital for resource-constrained satellites. Furthermore, we develop a quality of service (QoS)-aware route compression scheme designed to meet diverse IoT service demands. This scheme leverages the computational resources of terrestrial controllers to generate compact, STSR-encoded paths. By accounting for QoS requirements and satellite-specific dynamics, the embedded algorithm enhances routing performance for heterogeneous IoT flows within the satellite network. Simulation and evaluation results demonstrate that STSR outperforms existing SRv6-based approaches in path encoding efficiency, payload transmission efficiency, processing overhead, and traffic engineering performance. Jiang Liu 0010, Weihong Wu, Yingsheng Geng, Ran Zhang 0004, Tao Huang 0005 |
IEEE Internet Things J. | 6 |
| 2026 | Two-Timescales Optimization of Content Placement and Delivery in Satellite-Terrestrial Edge Computing NetworksabstractIn this paper, we establish a two-timescale framework for the joint optimization for the content placement and content delivery problem in satellite-terrestrial edge computing networks (STECN). Our goal is to optimize content placement to improve network performance while ensuring diverse quality of service (QoS) for content delivery. We decouple the problem into two timescales to balance real-time responsiveness and long-term efficiency. Specifically, considering frequent content placement incurs huge traffic cost, we optimize the content placement in order to reduce resource expenses in large timescales. The optimization problem is formulated as an integer linear programming (ILP) problem to improve both traffic efficiency and cache resource utilization. We leverage a heuristic atom search optimization (ASO) approach to address the problem, which yields an optimal strategy with low computational complexity. In small timescales, we model content delivery as a Markov decision process (MDP) to minimize content delivery delays at small timescales while maintaining smooth network traffic. A deep reinforcement learning (DRL) framework is used for policy learning to dynamically adapt to varying network conditions. By considering the correlation between the small and large timescale optimization, we propose a hierarchical solution to jointly address both issues. Finally, extensive simulations confirm the effectiveness and superiority of the proposed scheme. Renchao Xie, Qinqin Tang, Zeru Fang, Tao Huang 0005, Zehui Xiong |
IEEE Trans. Commun. | 5 |
| 2026 | P2TS: A Preemptive Approach for Priority-Aware Task Scheduling in Computing Power NetworksabstractAs an emerging computing paradigm, Computing Power Networks (CPNs) are dedicated to coordinating and managing network resources and computing resources to achieve interconnectivity in computing power perception. Efficient collaborative computing of massive data can be achieved through the scheduling function of CPNs. However, existing scheduling research mainly focuses on selecting network links and computing nodes, lacking consideration for task execution after scheduling, which may degrade the Quality of Service (QoS), leading to widespread failures and significant losses. To address this issue, we design a priority-aware preemptive task scheduling (P2TS) strategy for CPNs to jointly optimize task scheduling and execution in terms of success rate, average processing delay, and load balancing. Specifically, at the execution level, we propose a priority-aware preemptive mechanism (P2M) to optimize post-scheduling task execution. Then, at the scheduling level, we apply deep reinforcement learning (DRL) to optimize the scheduling process supporting the P2M in CPNs. A series of simulations are conducted to demonstrate the superiority of our strategy. Tao Huang 0005, Haoxiang Qiu, Qinqin Tang, Renchao Xie, Tianjiao Chen, Zehui Xiong |
IEEE Trans. Mob. Comput. | 1 |
| 2026 | A Cluster-Based Data Transmission Strategy for Blockchain Network in the Industrial Internet of ThingsabstractThe proliferation of devices and data in the Industrial Internet of Things (IIoT) has rendered the traditional centralized cloud model unable to meet the stringent requirements of wide-scale and low latency in these IIoT scenarios. As emerging technologies, edge computing enables real-time processing and analysis on devices situated closer to the data source while reducing bandwidth requirements. Blockchain, being decentralized, could enhance data security. Therefore, edge computing and blockchain are integrated in IIoT to reduce latency and improve security. However, the inefficient data transmission of blockchain leads to increased transmission latency in the IIoT. To address this issue, we propose a cluster-based data transmission strategy (CDTS) for blockchain network. Initially, an improved weighted label propagation algorithm (WLPA) is proposed for clustering blockchain nodes. Subsequently, a spanning tree topology construction (STTC) is designed to simplify the blockchain network topology, based on the above node clustering results. Additionally, leveraging clustered nodes and tree topology, we propose a data transmission strategy to speed up data transmission. Simulation experiments show that CDTS effectively reduces data transmission time and better supports large-scale IIoT scenarios. Ru Huo, Xiangfeng Cheng, Chuang Sun 0002, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2026 | Mercury: Multipath Spraying for Joint Congestion and Reordering Control in RDMAabstractDue to the low entropy traffic characteristics of LLM (Large Language Model) training, existing load balancing mechanisms such as Equal-Cost Multi-Path (ECMP) fail to fully utilize the redundant bandwidth between computing nodes in RDMA over Converged Ethernet (RoCE). Packet spraying mechanism has become a typical solution to the load balancing problem in RoCEs. However, it has a negative effect on congestion control mechanisms and suffers severe out-of-order problems. In this paper, we propose Mercury, a host-driven spraying scheme that synergizes congestion feedback and reordering control. Mercury selects paths by leveraging ECN, RTT, and reordering metrics, adjusts rates via multi-metric window. It also employs receiver-side buffers with priority-based dropping to mitigate out-of-order penalties. Evaluations in ns-3 under AllReduce and All-to-All traffic show that Mercury consistently outperforms the ECMP-based baselines, including DCQCN, TIMELY, HPCC, SWIFT, and BOLT, with the largest reduction in Max FCT reaching 63%. Under multi-path load balancing, Mercury delivers the lowest Max FCT for large messages in AllReduce and for most message sizes in All-to-All. It outperforms STRACK and MP-RDMA by up to 28% and 35% in AllReduce, and by up to 25% and 30% in All-to-All. Yuxiang Wang 0011, Jiao Zhang 0002, Leixin Cai, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2026 | LAPS: Latency Aware Packet Spraying on Unequal-Cost Multi-Path Data Center Networks
Ying Wan 0001, Jinyu Xiao, Haoyu Song 0001, Zhikang Chen, Yunhui Yang, Bin Liu 0001, Tao Huang 0005 |
IEEE Trans. Netw. | 8 |
| 2026 | HierCC: Taming Traffic Uncertainty in RDMA Data Centers With Hierarchical Congestion ControlabstractExisting congestion control schemes for RDMA resolve the dilemma of guaranteeing high throughput and ultra-low latency to some extent from a variety of perspectives. However, they are inefficient in addressing transient large queue build-up and under-utilized bandwidth caused by frequent traffic bursts. In this paper, we argue that traffic uncertainty is the fundamental challenge that limits these schemes from addressing the aforementioned dilemma. Inspired by the investigation that aggregated flows within the same rack are relatively long-lived, we propose HierCC, which aggregates flows destined to the same IP in a rack to ease traffic uncertainty and further provides hierarchically control within the first-hop ToR and between racks. Specifically, the inter-rack rates of aggregate flows are controlled by a credit-based mechanism. Then the bandwidth obtained by the aggregated flow is allocated to the corresponding intra-rack individual flows promptly and accurately. We implement HierCC in a testbed that consists of DPDK-based end-hosts and P4-based Tofino switches. The performance of HierCC is evaluated by comprehensive testbed experiments and SystemC/NS3 simulations. Results indicate that, compared with state-of-the-art, HierCC can mitigate buffer usage by up to$10\times $and reduce the average and 99th percentile FCT by up to 84% and 80%, respectively. Zirui Wan, Jiao Zhang 0002, Xiaolong Zhong, Zixuan Guan, Haoyu Pan, Tian Pan 0001, Tao Huang 0005 |
IEEE Trans. Netw. | 8 |
| 2026 | Weir: Scalable RDMA With Delay-Based RNIC Cache Control Software Middleware for Data Center NetworksabstractRemote Direct Memory Access (RDMA) is widely used in distributed services in Data Center Networks (DCNs) due to its high performance. As DCNs expand in scale, RDMA faces scalability issues. The reason is that the high concurrency Queue Pairs (QPs) lead to cache misses on RDMA Network Interface Card (RNIC) and frequent evictions, and the behaviour of fetching the cache via PCIe leads to performance degradation of RDMA. In this paper, we model the behaviour of Work Queue Element (WQE) on RNIC as a producer-consumer model and investigate that the root cause of WQE cache misses is the mismatch between the production rate of the CPU and the consumption rate of the RNIC. We design Weir from the perspective of WQE cache control to avoid cache misses and improve throughput under high concurrent QPs. Weir determines the cache occupancy on the RNIC by monitoring the number of active QPs and the increase/decrease in the life cycle of WQEs, and calculates the production rate and pacing by credit. The implementation of Weir exhibits minimal CPU overhead. Evaluation results show that Weir can maintain 97Gbps throughput without degradation even with up to 16K concurrent QPs, and effectively reduces various observable cache misses by$5\times $to$10\times $compared to commercial RNICs. Additionally, experiments show that Weir has better connection scalability than XRC and DCT. Jiao Zhang 0002, Yongchen Pan, Dexuan Liao, Huimin Luo, Tao Huang 0005, Haipeng Yao |
IEEE Trans. Netw. | 6 |
| 2026 | Access Resource Allocation With ISL-Based Backhaul Awareness in LEO Satellite NetworksabstractLow Earth Orbit satellite networks play a significant role in providing global ubiquitous services. The application of Inter-Satellite Links (ISLs) has accelerated development in Non-Terrestrial Networks with satellite backhaul. However, ISL-based backhaul does not match the performance of terrestrial fiber links, which makes its impact on end-to-end performance non-negligible. Existing radio resource allocation solutions usually neglect the backhaul performance, potentially failing to meet end-to-end Quality of Service requirements. In this work, we propose ARA-IBA, an access resource allocation scheme with ISL-based backhaul awareness. The satellite network’s backhaul path states, including delay and packet loss rate, are exposed to on-board base stations to enable dynamic resource scheduling. In ARA-IBA, resources are jointly scheduled for Guaranteed Bit Rate (GBR), Delay-critical GBR, and Non-GBR users. For GBR and Delay-critical GBR users, backhaul delay is utilized to optimize end-to-end delay satisfaction. A clustering game is employed to perform fine-grained allocation adjustments. Overloaded Non-GBR users are then scheduled using a pointer network. ARA-IBA optimizes backhaul packet loss rate while maintaining scalability to accommodate varying user numbers. Simulation results demonstrate that the proposed algorithm outperforms conventional methods in terms of users’ delay satisfaction and backhaul packet loss rates. Ran Zhang 0004, Jiang Liu 0010, Shiran Sun, Xinyue Lu, Qinqin Tang, Tao Huang 0005 |
IEEE Trans. Wirel. Commun. | 7 |
| 2025 | Topology-Adaptive LEO Satellite Network Telemetry via Graph Isomorphism and Topology Partitioning
Yan Zhang 0063, Tian Pan 0001, Guohao Ruan, Yi Liu 0151, Jiang Liu 0010, Tao Huang 0005 |
APNet | 8 |
| 2025 | Efficient Large-scale Model Training with Disaggregated Storage and Computing Architecture
Mengyao Han, Zheng Ruan, Naihan Zhang, Tao Huang 0005, Xiongyan Tang |
APNet | 8 |
| 2025 | Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Yang Song 0031, Tian Pan 0001, Zhigang Zong, Bengbeng Xue, Xionglie Wei, Yisong Qiao, Donglin Lai, Baohai Hu, Jin Ke 0005, Enge Song, Jianyuan Lu, Xing Li 0007, Biao Lyu, Rong Wen, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu |
APNet | 19 |
| 2025 | NSDocker: A Lightweight and Realistic Satellite Network Emulator Integrating NS-3 and Docker
Guohao Ruan, Tian Pan 0001, Haibin Song, Qiang Fu 0011, Yi Liu 0151, Tao Huang 0005 |
APNet | 9 |
| 2025 | CFseq: A Framework for Constructing Compression-Friendly Field Sequences for Network LogsabstractThe rapid growth of network traffic has resulted in a substantial increase in log data, creating significant challenges for storage and processing. Although general-purpose compression algorithms are widely used, they often underperform on network logs due to their inability to exploit inherent structural characteristics. While advanced compression techniques can offer better performance, they typically require extensive system modifications and add deployment complexity. This paper presents CFseq, a lightweight and efficient framework designed to construct compression-friendly field sequences that improve the compressibility of network logs. CFseq is founded on two key observations: first, some fields exhibit high redundancy; second, others contain shared prefixes or suffixes that are well suited to compression algorithms. The framework comprises two modules: the Text Similarity Enhancement module, which ranks fields based on information entropy, and the Brute-Force Search module, which identifies the optimal field order for compression. CFseq operates without modifying existing compression or decompression pipelines, allowing for seamless and low-cost integration. Experimental results show that CFseq improves the compression ratios of general-purpose compressors by up to 32 % and enhances the performance of the state-of-theart advanced compressor Denum by up to 20 %. Yunwei Dai, Tao Huang 0005, Shuo Wang 0006 |
CLUSTER | 2 |
| 2025 | Modular State Channels Enable Efficient Blockchain-based Web 3.0
Wei Chen 0131, Ru Huo, Yang Liu 0171, Tao Huang 0005, Jiaheng Zhang |
GLOBECOM | 4 |
| 2025 | Green Digital Twin-Enabled IIoT: Jointly Optimizing Service Freshness and Carbon EmissionabstractDigital Twin (DT) technology is a key enabler of the Industrial Internet of Things (IIoT), facilitating predictive control, fault detection, and simulation through high-fidelity virtual replicas of physical assets. To alleviate the significant computational burden on the Cloud Server (CS) of centralized DT services, existing solutions commonly deploy DT modules on edge servers (ESs). However, deploying each DT module at the edge requires intensive resources to support frequent data updates, processing, and analysis; thus, unrestricted module deployment can pose substantial sustainability challenges, particularly when renewable energy availability is limited. To tackle this challenge, we propose a green DT-IIoT architecture and formulate a DT module placement problem that jointly minimizes DT service freshness and carbon emissions, subject to constraints on data synchronization accuracy, computing resources, and storage capacity. Among these, DT service freshness, which indicates real-time responsiveness, and data synchronization accuracy, which reflects the reliability of real-time input data, are two critical metrics affecting DT service quality. Furthermore, we propose a Dueling Double Deep Q-Network (D3QN) based placement algorithm (DDMP), which achieves high performance with a relatively simple structure that is well-suited for rapidly evolving IIoT scenarios. Simulation results demonstrate that our proposed approach enhances DT service freshness while reducing carbon emissions compared with baseline methods. Renchao Xie, Gaochang Xie, Qinqin Tang, Tao Huang 0005 |
GLOBECOM | 7 |
| 2025 | Joint Popularity-Aware Distributed Layered Service Caching and Application Deployment in Mec NetworksabstractThe exponential increase in connected user devices poses scalability challenges for centralized cloud computing. Mobile Edge Computing (MEC) and Fog Computing alleviate latency by deploying computation and storage resources closer to end-users. However, due to the resource limitations, heterogeneity, and dispersed nature of edge servers, there is a need to jointly optimize service caching and application placement strategies to enhance service quality. Given the widespread use of containerized services at the edge, we propose a distributed caching scheme that allows all edge nodes to cache services at the granularity of container image layers. This collaborative caching approach reduces the real-time latency, bandwidth consumption, and caching costs associated with retrieving and initializing applications. Additionally, to address the variability in application popularity across different edge regions, we model application popularity using a Zipf distribution and construct a multi-slot joint optimization model for caching and deployment decisions based on deployment cost, application startup time, and average delay. We then propose a two-stage optimization method to solve this model, demonstrating through comparison with centralized and P2P models the effectiveness of the proposed approach. Renchao Xie, Qinqin Tang, Tao Huang 0005, Tianjiao Chen, Gaochang Xie, Zehui Xiong |
ICC | 4 |
| 2025 | Intelligent Control Integrating Sensing, Communication and Computing in Industrial Internet of ThingsabstractIn recent years, the rapid development of the industrial Internet of things (IIoT) has brought new innovation opportunities to the manufacturing industry, but it also faces some major challenges. Currently, the IIoT systems often lack flexibility in sensing capabilities and have rigid communication architectures, resulting in insufficient coordination between different control tasks. In addition, the disconnect between sensing, communication, and computing further limits the system's ability to achieve optimal control, and the lack of intelligent data processing in the system also increases the control cost. In response to these challenges, this paper proposes an intelligent control framework, CISCC, which combines industrial edge computing technology with artificial intelligence (AI) models to achieve a deep integration of sensing, communication, and computing resources in IIoT systems, aiming to jointly optimize the configuration of these resources to improve the overall system performance and reduce control costs. Simulation results show that the CISCC framework can effectively handle resource allocation problems in IIoT systems, thereby better supporting the development of smart manufacturing applications. Yutian Yang, Zihang Yin, Qinqin Tang, Yang Liu 0171, Jiayi Cui, Renchao Xie, Tao Huang 0005 |
ICC | 7 |
| 2025 | Accelerating Distributed Training on Parameter Server Architecture With Path-Aware MulticastabstractIt is observed that the bottleneck in distributed training has shifted from computation to communication due to contention in concurrent transmissions and substantial redundant traffic. In the Parameter Server (PS) architecture, the server aggregates gradients from multiple workers and then distributes updated model parameters back to the workers in a one-to-many manner. Currently, model parameters are distributed via unicast, sending multiple identical copies of the data, which leads to significant bandwidth waste. Although multicast can save bandwidth, current approaches have two main drawbacks: on one hand, many protocols require maintaining excessive multicast state inside the network; on the other hand, the lack of coordination among multiple multicast trees can still lead to path conflicts. In this work, we propose path-aware multicast, which includes innetwork multicast tree reservation and per-hop control multicast. Specifically, before each round of model parameter distribution, the server queries the network for a multicast tree that satisfies the bandwidth requirement. The calculated multicast tree is then returned with bandwidth reserved at its tree nodes. Next, model parameters are forwarded with hop-by-hop control along the multicast tree. After the multicast is completed, the reserved network resources are released. Our evaluation shows that in an$8 \times 8$spine-leaf topology, path-aware multicast improves link load balancing by 32.6 % compared to random multicast and accelerates model parameter distribution by up to nearly$N \times$compared to unicast, where$N$is the number of workers. Chuanying Yuan, Tian Pan 0001, Guohao Ruan, Hao Li 0011, Yan Zou, Jiao Zhang 0002, Tao Huang 0005 |
ICC | 10 |
| 2025 | Achieving Adaptive Multi-Path Routing and Order-Preserving Time Slot Planning in TSNabstractWith the rise of autonomous driving, the performance requirements for In-vehicular networks are continuously increasing. Existing research leverages Frame Replication and Elimination for Reliability (FRER) and Time-Aware Shaper (TAS) mechanisms in Time-Sensitive Networking (TSN) to achieve deterministic transmission. FRER requires transmitting flows over multiple disjoint paths. However, FRER lacks redundancy degree selection strategies, and the delay differences between redundant paths lead to packet disorder, which increases network resource overhead and compromises traffic QoS. In this paper, we propose a Bandwidth-Aware with Frame Replication and Elimination for Reliability (BA-FRER) algorithm dynamically selects redundancy degree based on the current network bandwidth resources, delay, and reliability utility function. Additionally, we propose a Redundant-Aware Order Preservation (RAOP) algorithm configures time slots for each flow based on the TAS mechanism to align the delays of redundant paths. The evaluation results show that the BA-FRER algorithm improves the utilization of network bandwidth resources, the flow access rate, and the reliability, while the RAOP algorithm reduces the probability of packet disorder. Yanke Li, Shuo Wang 0006, Guoyu Peng, Guizhen Li, Jiao Zhang 0002, Tao Huang 0005 |
ICCCN | 8 |
| 2025 | Mercury: A Dynamic Multi-path Packet Spraying Scheme for RDMA NetworksabstractDue to the low entropy traffic characteristics of LLM (Large Language Model) training, existing load balancing mechanisms such as Equal-Cost Multi-Path (ECMP) fail to fully utilize the redundant bandwidth between computing nodes in RDMA over Converged Ethernet (RoCE). Packet spraying mechanism has become a typical solution to the load balancing problem in RoCEs. However, it has a negative effect on congestion control mechanisms and suffers severe out-of-order problems.In this paper, we propose Mercury, an host-driven spraying scheme that synergizes congestion feedback and reordering control. Mercury selects paths by leveraging ECN, RTT, and reordering metrics, adjusts rates via multi-metric window. It also employs receiver-side buffers with priority-based dropping to mitigate out-of-order penalties. Evaluations in ns-3 under AllReduce/All-to-All traffic show Mercury reduces maximum flow completion time (Max FCT) by 40%-63% compared to ECMP-based DCQCN/TIMELY/HPCC. It also achieves at least 10%-20% improvement against switch-based spraying. Yuxiang Wang 0011, Jiao Zhang 0002, Zirui Wan, Leixin Cai, Shuo Wang 0006, Tao Huang 0005 |
ICCCN | 6 |
| 2025 | Valve: Scalable RDMA with Gap-based Cache Control Middleware for Data Center NetworksabstractRemote Direct Memory Access (RDMA) has become a cornerstone technology in Data Center Networks (DCNs). However, DCNs have expanded substantially, leading to severe connection scalability issues for RDMA. The critical reason behind these issues stems from frequent cache misses on RDMA NICs (RNICs) when handling numerous Queue Pair (QP) connections. Cache misses require time-consuming retrievals from the host via PCIe, resulting in a degradation in RDMA performance. Existing software solutions primarily aim to alleviate QP Context (QPC) cache pressure, while hardware solutions incur prohibitive costs. In this paper, we identify Work Queue Element (WQE) cache, rather than QPC cache, as the fundamental bottleneck limiting connection scalability. Hence, we propose a software middleware, Valve, designed to mitigate cache misses through WQE cache control. Valve regulates WQE posting to control WQE cache by monitoring RNIC cache usage and adaptively adjusting the gap of WQE posting. Valve boasts ease of deployment, requiring the addition of approximately 1000 lines of code, and incurs low CPU overhead. Valve maintains peak performance of RNIC throughput regardless of the number of QPs and considerably reduces observable cache misses (such as ICM, MTT, and MPT cache misses) by 2.8× to 3.1× compared to XRC and DCT. Jiao Zhang 0002, Dexuan Liao, Yongchen Pan, Tao Huang 0005 |
ICNP | 6 |
| 2025 | StableRoute: When Dijkstra's Algorithm Meets Topology-Varying Satellite NetworksabstractLow Earth Orbit (LEO) satellite constellations are becoming a viable means for Internet access. However, their topology changes as satellites move towards or away from orbital intersection points, leading to constant link down or up. This may cause routing table entry updates and thus path changes between satellites. A path change during transmission may lead to out-of-order packet delivery and invalidate the current TCP congestion window. While some path changes are inevitable, some are avoidable. Dijkstra's algorithm is a popular choice among the routing protocols proposed for LEO satellite networks. We observe that many next-hop route updates by Dijkstra's algorithm are avoidable. Motivated by this, we propose StableR-oute, which stabilizes routing paths from different perspectives. StableRoute Local (SR_L) leverages equal-cost shortest paths and stays with the current one if it is still valid. StableRoute K-Short (SR_K) allows a path longer than the shortest path. StableRoute Global (SR_G) leverages the predictable satellite trajectories and topology variations, and thus works out a next-hop route selection sequence that minimizes the number of route updates over a time period. The evaluation shows that SR_L, SR_K and SR_G outperform Dijkstra's algorithm, substantially reducing the number of route updates in changing topologies. Tian Pan 0001, Guohao Ruan, Qiang Fu 0011, Zhengjie Luo, Xingshuang Luo, Tao Huang 0005 |
INFOCOM | 7 |
| 2025 | ACC: Addressing Performance Limitations in Datacenters with Atomic Congestion Control
Zirui Wan, Jiao Zhang 0002, Tian Pan 0001, Pingping Lin, Tao Huang 0005 |
INFOCOM | 7 |
| 2025 | HELDR: Packet Loss Detection and Retransmission for Live Streaming Hyper-Edge NetworkabstractLive streaming platforms like Douyin have developed the Live Streaming Hyper-Edge Delivery Network (LSHEDN) to reduce bandwidth cost. In LS-HEDN, the Content Delivery Network (CDN) splits the live streaming into multiple substreams by randomly assigning each frame to them. Hyperedge devices like set-top boxes with cheap and idle bandwidth resources forward a substream from CDN to multiple users. A protocol based on User Datagram Protocol (UDP) is adopted between devices and users, with users detecting packet loss and requesting retransmissions via Negative Acknowledgment (NACK). Given the demand for lower latency and the inherent fluctuations in public network, existing receiver-side packet loss detection and retransmission methods fall short in achieving both timeliness and accuracy simultaneously. This is manifested as frequent rebuffering and excessive redundancy. Notably, when head-of-line blocking(HOL blocking) occurs in the upstream link of the device, these issues become even more pronounced. To address this, we propose Hyper-Edge Loss Detection and Retransmission (HELDR) algorithm. It features a loss detection algorithm tailored to the transmission characteristics in LSHEDN, which improves detection accuracy. Its immediate retransmission mechanism and the backup devices retransmission mechanism enhance timeliness. Large-scale online A/B tests results show that HELDR reduces the average rebuffering rate by 41.2%, reduces the average redundancy rate by 15.7%. Peisheng Guo, Jiao Zhang 0002, Zhichen Xue, Yajie Peng, Xiaofei Pang, Tao Huang 0005, Ruili Fang, Zhenpeng Zhu, Dehui Wei |
IWQoS | 9 |
| 2025 | Achilles: an Enhanced Scheme for Reactive Transport in Datacenters
Zirui Wan, Jiao Zhang 0002, Haoyu Pan, Tao Huang 0005 |
IWQoS | 4 |
| 2025 | Hermes: Enhancing Layer-7 Cloud Load Balancers with Userspace-Directed I/O Event NotificationabstractLayer-7 load balancers (L7 LBs) improve service performance, availability, and scalability in public clouds. They rely on I/O event notification mechanisms such as epoll to dispatch connections from the kernel to userspace workers. However, early epoll versions suffered from the thundering herd problem. Epoll exclusive (available since Linux 4.5) mitigates this but introduces LIFO wakeups, causing connection concentration on a few workers. Reuseport (Linux 3.9) hashes connections across workers but suffers from hash collisions and lacks awareness of worker load. Since each worker serves multi-tenant traffic, inter-worker load balancing is critical to avoid worker overload and preserve tenant performance isolation. Tian Pan 0001, Enge Song, Yueshang Zuo, Shaokai Zhang, Yang Song 0031, Jiangu Zhao, Wengang Hou, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu |
SIGCOMM | 13 |
| 2025 | DNSLogzip: A Novel Approach to Fast and High-Ratio Compression for DNS LogsabstractDomain Name System (DNS) logs capture detailed records of the queries and responses exchanged between DNS servers and clients, playing a crucial role in applications such as cybersecurity monitoring and regulatory compliance, which often require long-term data retention. With the rapid growth of Internet traffic, the volume of DNS logs has surged, presenting significant storage challenges. Although many DNS operators use general-purpose compression algorithms to reduce storage costs, these solutions fail to fully exploit the unique characteristics of DNS data, leading to inefficiencies and rising storage demands. Yunwei Dai, Guyue Liu, Tao Huang 0005, Shuo Wang 0006, Xingli Wu, Heshun Li, Fanglong Hu |
SIGCOMM | 3 |
| 2025 | ByteTracker: An Agentless and Real-time Path-aware Network Probing SystemabstractAs the number of data center servers grows into the millions and due to the demand for more accurate, rapid and powerful network fault detection and location, the existing Pingmesh-centric monitoring and diagnostic system is not efficient enough. In this paper, we propose ByteTracker, the first agentless probing and diagnostic system for large-scale data center networks. It does not need to deploy probe processes or make any configurations on end hosts, and all probes are launched by a small number of centralized Probers. ByteTracker achieves accurate, real-time probe path tracking with packet mirroring on switches. By reducing end-host probe noise, precisely identifying network timeout probes, accurately tracking probe paths, and marking the failed switch with multiple network timeout probes, ByteTracker can locate network failures with nearly 100% accuracy. We have deployed ByteTracker in all of our data centers for over half a year. During deployment, ByteTracker can detect almost all network anomalies and locate them within 5 seconds with 100% accuracy. Shixian Guo, Kefei Liu 0004, Yulin Lai, Yangyang Bai, Jianghang Ning, Yongbin Dong, Sisi Wen, Jiale Feng, Chengcai Yao, Zhuo Jiang, Jiao Zhang 0002, Tao Huang 0005 |
SIGCOMM | 23 |
| 2025 | Fine-Grained Service Scheduling Scheme Based on Application-Aware Network for Internet of VehiclesabstractA novel fine-grained service scheduling scheme based on application-aware network (APN) for Internet of Vehicles (IoV) is proposed to achieve efficient and differentiated network resource orchestration. Application and service flow information is translated to application identifier (ID) and Sub-service ID encapsulated in IPv6 header. In this way, network can perceive various applications and provide suitable network channel resource for them. The APN-based information identifying method for application does not require unpacking of data packets and does not overly rely on network controllers, which can improve service efficiency. The experimental verification has been conducted on IoV experimental network in Xiong'an, China. In the experiment, the end-to-end delay of remote driving of APN autonomous vehicle is decreased by 67.5% compared to that of normal autonomous vehicle. The reason is that the network can identify the critical service flow sent by the APN vehicle and guide it to the dedicated channel to avoid competing for bandwidth resources in the shared channel which has 10% packet loss rate. In addition, the packet loss rate thresholds for all remote driving service flows are obtained through experiment, so that the optimal flow orchestration can be used to realize efficient utilization of network resource while ensuring smooth remote control. Naihan Zhang, Xinxin Yi, Gui Wen, Chong Zheng, Qiangzhou Gao, Mengyao Han, Tao Huang 0005, Xiongyan Tang |
VTC2025-Spring | 9 |
| 2025 | Intelligence Sharing in LEO Satellite Edge Computing Networks: A Coalition-based ApproachabstractIn this paper, we propose an innovative architecture for sharing intelligence in low earth orbit (LEO) satellite edge computing networks. Specifically, we adopt the sharing of intel-ligence to satellites via ground stations to improve the response speed of satellites in processing intelligent services. Considering the burden of frequent transmission of intelligent models over unstable ground-satellite links, the satellites share intelligence with each other in the coalition, which greatly reduces the service response delay. In addition, considering the poor generalization of pre-trained intelligent models transmitted by ground stations, we design a model aggregation scheme with differentiated weights. Each coalition appoints a coalition center satellite, tasked with aggregating models and re-sharing them to individual satellites, thereby enhancing model performance. Then, we propose two low-time complexity algorithms to solve the above problems. Finally, the effectiveness and superiority of the proposed schemes are verified through extensive simulations. Zeru Fang, Qinqin Tang, Renchao Xie, Tao Huang 0005, Tianjiao Chen, Ran Zhang 0004, Sha Tan |
WCNC | 4 |
| 2025 | Torrent: Re-Architecting End-to-End Transmission for Cross-Datacenter RDMA NetworksabstractSustainability is becoming increasingly challenging in today's data centers with limited space, power and connectivity. Large cloud service providers interconnect geographically distributed datacenters for better scalability and availability. Applications running on cross-datacenter network impose great challenges in transport design. In this paper, we identify two inherent limitations of extending the existing transport technology, RDMA, and its Ethernet derivative, RoCE, to long-haul transmission. First, the on-chip resources of commodity RDMA NICs are insufficient for long-haul transmission. Second, applying existing traffic control schemes to inter-datacenter environment exhibits poor performance. Motivated by this, we propose Torrent, a switch-driven transport framework which partitions end-to-end control into three sub-control loops. To achieve the combined goals of fairness and high performance in cross-datacenter scenarios, Torrent employs fast acknowledgment and near-end congestion control on datacenter interconnection (DCI) switches. We implement Torrent prototypes on commodity programmable switches and evaluate it through real-world testbed experiments. Our results show that Torrent can achieve high link utilization over ultra-long distances and quickly converge congested flows to steady rates. Haoyu Pan, Zirui Wan, Jiao Zhang 0002, Tao Huang 0005 |
WCNC | 5 |
| 2025 | LMSR: A Low-Jitter Multiple Slots Routing Algorithm in LEO Satellite NetworksabstractIn recent years, Low Earth Orbit (LEO) constellation based networks have attracted wide attention from both the academia and the industry. Broadband access and backhaul becomes a typical application for LEO satellite networks. However, the high-speed motion of satellites brings periodic fluctuation of delays, which is a great violation to the Quality of Service (QoS) provisioning. In this work, inspired by the success of Software Defined Networking (SDN), and considering the dynamics of LEO satellite networks, we propose a low-jitter multiple slots routing in LEO satellite networks to provide low-jitter end-to-end delay path computing and control. The proposed multiple slots routing optimization mechanism takes account of the topology shift as well as the dynamic propagation delay across multiple topology snapshots, and thus achieves low-jitter performance. The performance improvement of the proposed mechanism is validated by the simulation, which promises the value of jitter under 30 ms. Shiran Sun, Ran Zhang 0004, Zekun Sun, Qinqin Tang, Tao Huang 0005 |
WCNC | 6 |
| 2025 | Service Anycast Forwarding for Software Defined Computing Power NetworkabstractWith the rise of the computing power network (CPN), which integrate edge computing, cloud computing, and network infrastructure, replicated computing services are increasingly distributed to meet user demands for location-independent, reliable, and low latency services. Service anycast forwarding coordinates distributed service instances by binding them to a unified identifier and dynamically routing requests to the optimal instance. However, challenges such as varying user demand distribution, network complexity, and service instance heterogeneity complicate balanced service forwarding. To address these, we propose an SDN-based service anycast forwarding mechanism for CPN (SA-CPN). In the data plane, a cyclic forwarding queue efficiently maps weighted strategies and selects instances for each service request, improving policy performance. In the control plane, an optimal transport model balances network and computation latency based on service instance capabilities. We further design an optimal transport-based service anycast forwarding algorithm (OTSAF) using Sinkhorn iterations. Our implementation of SA-CPN in a real system shows that OTSAF consistently outperforms four baseline methods across various performance metrics. Renchao Xie, Qinqin Tang, Tao Huang 0005, Tianjiao Chen, Zehui Xiong |
WCNC | 4 |
| 2025 | Time-Space-Varying Resource Graph-Based Dependent Task Offloading for Satellite-Terrestrial Integrated Computing Power NetworksabstractWith the continuous advancement of network technologies and hardware devices, computation-intensive and latency-sensitive tasks have emerged worldwide, requiring networks to provide extensive coverage, low latency, and robust computing capabilities. Leveraging the global coverage of LowEarth Orbit (LEO) satellites and the flexible resource invocation capabilities of the Computing Power Network (CPN), we propose a Satellite-Terrestrial Computing Power Network (ST-CPN) framework that integrates both strengths. In this framework, tasks can be offloaded to satellites closer to users for processing, ensuring high-quality services anytime and anywhere. However, due to the dynamic nature of the network and the limited resources of individual nodes, efficiently executing complex dependent tasks presents significant challenges. Therefore, we investigate the dependent task offloading problem in the dynamic ST-CPN environment. Considering dynamic changes of topology and available resources caused by satellite mobility, we propose a Time-Space-Varying Resource Graph (TSVRG) to capture the status of the communication, storage, and computation resources. On this basis, given that individual nodes struggle to process dependent tasks, we offload multiple subtasks of a task to different nodes for collaborative processing. In this paper, we model the task as a Directed Acyclic Graph (DAG) and transform the offloading problem into a mapping problem from the DAG to TSVRG. We then introduce a Delay Predictionbased Graph Mapping Algorithm (DPGMA) to address this problem. Simulation results indicate that our scheme achieves better performance than the benchmark schemes. Renchao Xie, Qinqin Tang, Zehui Xiong, Gaochang Xie, Tao Huang 0005 |
WCNC | 6 |
| 2025 | SeqBalance: Congestion-Aware Load Balancing With No Reordering in Data Center NetworksabstractWith the rapid development of the Internet of Things (IoT), an increasing amount of sensor data generated by IoT applications has been transferred to data center networks for storage and data analysis. Remote Direct Memory Access (RDMA) is widely used in data center networks because of its high performance. However, due to the characteristics of RDMA’s retransmission strategy, current load balancing schemes for data center networks are unsuitable for RDMA. In this paper, we propose SeqBalance, a load balancing framework designed for RDMA. SeqBalance implements fine-grained load balancing for RDMA through a reasonable design and does not cause reordering problems. SeqBalance detects link congestion at the switch by sensing ECN signals and link utilization, and guides routing decisions accordingly. SeqBalance’s designs are all based on existing commercial RNICs and commercial programmable switches, so they are compatible with existing data center networks. We have implemented SeqBalance Shaper for fine-grained sub-flow splitting in Mellanox CX-6 RNIC and implemented routing decisions in Intel Tofino P4 programmable switch. The results of hardware testbed experiments and large-scale simulations show that compared with existing load balancing schemes, SeqBalance improves 24.7% and 15.9% on average FCT and 99th-percentile FCT. Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
IEEE Internet Things J. | 6 |
| 2025 | Achieving Class-Aware Mixed-Flow Scheduling in Hybrid Wired-Wireless Time-Sensitive NetworksabstractThe emergence of new time-sensitive networking (TSN) technologies empowers almost-deterministic ultra-reliable low-latency communications for industrial cloud-fog automation paradigms. However, current heterogeneous networks struggle to balance time-sensitivity and flexibility, particularly in mixed-flow scenarios. Existing scheduling approaches within 5G-TSN integration model either suffer from quality of service (QoS) flow mismatch or microbursts. This paper proposes a class-aware mixed-flow scheduling (CAMFS) strategy to enable domain-specific resource allocation in a hybrid wired-wireless TSN, while meeting the differentiated time-sensitive (TS) requirements. A time-triggered multiple cyclic-queuing (TTMCQ) shaper is designed to effectively align the classified and regulated mixed flows from TSN domains to 5G QoS flows through class-aware mapping. We also introduce an anti-starvation resource optimization method that minimizes the total average idle resources resulting from temporal-spatial resource over-provisioning within reserved TS windows (RTWs). Additionally, we present a CAMFS algorithm aimed at enhancing schedulability and resource utilization by sorting mixed flows based on a combination of flow features. Finally, simulation results show that CAMFS exhibits outstanding scheduling performance in terms of end-to-end service latency, scheduling success ratio, and normalized resource distribution compared to other state-of-the-art methods. Guoyu Peng, Shuo Wang 0006, Tao Huang 0005, Kangzhe Zhao, Guizhen Li |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | SeCo4: Co-Design of Sensing, Communication, and Computing for Intelligent Control in Industrial Cyber-Physical SystemsabstractIndustrial Cyber-Physical Systems (CPS) have made significant strides in recent years, driving the future of manufacturing. However, for further advancement in Cloud-Fog Automation (CFA), several challenges remain: rigid sensor sampling, inflexible communication configurations, insufficient coordination between cloud and fog resources, and a lack of integration between sensing, communication, and computing for effective control. To address these issues, this article presents SeCo4, an intelligent control framework for the co-design of sensing, communication, and computing in industrial CPS. The SeCo4 optimization problem is analyzed and divided into two sub-problems: a multi-controller cloud resource competition problem, formulated with a combinatorial auction to enable multi-controller competition for additional cloud resources and improve control performance; and a joint resource optimization problem for sensing, communication, and computing, modeled using a Mixed Integer Programming (MIP) problem to minimize control costs. Given the interdependence of these sub-problems, a hierarchical solution based on the online matching mechanism and the heuristic approach is developed to iteratively find the optimal solution. Finally, extensive simulations demonstrate the effectiveness and superiority of the proposed approach. Qinqin Tang, Yutian Yang, Jiayi Cui, Renchao Xie, Tao Huang 0005, Tianjiao Chen, Ran Zhang 0004, Zehui Xiong |
IEEE J. Sel. Areas Commun. | 5 |
| 2025 | QoE-Optimized MultiPath Scheduling for Video Services in Large-Scale Peer-to-Peer CDNsabstractVideo content providers such as Douyin implement Peer-to-Peer Content Delivery Networks (PCDNs) to reduce the costs associated with Content Delivery Networks (CDNs) while still maintaining optimal user-perceived quality of experience (QoE). PCDNs rely on the remaining resources of edge devices, such as edge access devices and hosts, to store and distribute data with a Multiple-Server-to-One-Client (MS2OC) communication pattern. MS2OC parallel transmission pattern suffers from severe data out-of-order issues. PCDNs offer significant cost savings by using multiple low-cost edge devices. However, due to its unique characteristics, including pull-based streaming transmission, many heterogeneous paths, and large receiving buffers, directly applying existing schedulers designed for Multipath TCP (MPTCP) to PCDN fails to meet the two goals of high aggregate bandwidth and low end-to-end delivery latency. To tackle this issue, we provide a detailed overview of Douyin’s self-developed PCDN video transmission system and introduce the first QoE-enhanced packet-level scheduler for PCDN systems, named Pscheduler. Pscheduler evaluates path quality with a congestion-control-decoupled algorithm and employs our proposed path-pick-packet method for data distribution, ensuring a smooth video playback experience. Additionally, we propose a redundant transmission algorithm to enhance task download speeds for segmented video transmission. Our extensive online A/B tests, involving 100,000 Douyin users generating tens of millions of video data points, demonstrate that Pscheduler achieves an average improvement of 60% in goodput, a 20% reduction in data delivery waiting time, and a 30% reduction in rebuffering rates. Furthermore, we conducted simulation experiments that further validate the effectiveness of Pscheduler, confirming its improvements in performance metrics under various network conditions. Dehui Wei, Jiao Zhang 0002, Xiang Liu 0017, Zhichen Xue, Tao Huang 0005, Linshan Jiang, Jialin Li 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | Collaborative Video Processing of Multiple Cameras in Smart Transportation: Content Analysis and Resource AllocationabstractIn the context of smart transportation, the collaborative processing of video data sourced from multiple cameras plays a pivotal role in promoting efficient traffic management and augmenting safety measures. Nevertheless, the exponential surge in surveillance cameras deployment has concurrently engendered a rapid increase in the magnitude of video analysis tasks and data volume. To address these challenges, we propose a comprehensive framework for collaborative video processing. Primarily, a collaborative content analysis approach is proposed, and which employs a Transformer-based ReID (Re-identification) algorithm to construct key stickers. These key stickers are optimized with cross-cameras correlations and serve as the foundational structure for subsequent online video compression. Subsequently, we propose a collaborative resource allocation approach, and which involves the formulation of a queue model designed for the orchestration of online camera analysis tasks. In addition, we have devised an enhanced deep reinforcement learning algorithm to fine-tune the task scheduling configuration of multiple cameras, with guidance from the queue model. Extensive experiments and simulations were conducted to evaluate the proposed framework. The results demonstrate its effectiveness in achieving accurate and real-time analysis of video data in smart transportation scenarios. Ru Huo, Chuang Sun 0002, Shuo Wang 0006, Tao Huang 0005 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Optimizing Fault-Tolerant Time-Aware Flow Scheduling in TSN-5G NetworksabstractThe integration of time-sensitive networking (TSN) and fifth-generation (5G) offers a promising solution for real-time and reliable data transmission in the Industrial Internet of Things (IIoT). However, current research focuses on traffic scheduling in TSN-5G networks to support low latency. New challenges arise when TSN-5G networks leverage time-aware shaper (TAS) and frame replication and elimination for reliability (FRER) to achieve low latency and high reliability. Simply combining TAS and FRER (SCTF) requires scheduling all time-triggered (TT) flows and their replica flows, which substantially increases the computational complexity of gate control lists (GCLs) and severely weakens scheduling capabilities. Moreover, the packet elimination function (PEF) in FRER may induce packet misordering. In this paper, we propose an efficient and fault-tolerant time-aware shaper (EF-TAS) mechanism for TSN-5G networks. EF-TAS only allocates timeslots for TT flows, while replica TT (RT) flows are delivered using a best-effort strategy. Due to the potential violation of deadlines in RT flows, we design an adaptive cyclic GCL window (ACGW)-based hybrid scheduling (AHS) algorithm to schedule TT and RT flows differentially. The AHS algorithm utilizes network calculus to ensure the timely arrival of RT flows without affecting the deterministic transmission of TT flows. In particular, we provide upper bounds on the amount of reordering to quantify the disorder caused by PEF and analyze the impact of introducing the packet ordering function (POF) on EF-TAS performance. The evaluation results show that EF-TAS not only meets the reliability and deadline requirements but also significantly reduces the total number of GCL entries and the computation time of GCLs compared to state-of-the-art methods. Guizhen Li, Shuo Wang 0006, Yudong Huang, Tao Huang 0005, Yuanhao Cui, Zehui Xiong |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Incentive Mechanism Design for Trust-Driven Resources Trading in Computing Force Networks: Contract Theory ApproachabstractRecently, Computing Force Networks (CFNs) have emerged to deeply integrate and flexibly schedule multi-layer, multi-domain, distributed, and heterogeneous computing force resources. CFNs build a resources trading platform between consumers and providers, facilitating efficient resource sharing. Therefore, resources trading is an important issue but it faces some challenges. Firstly, because all kinds of large-scale and small-scale resource providers are distributed in a wide area and the number of consumers is larger compared with edge/cloud computing scenarios, the credibility of consumers and providers is hard to guarantee. Secondly, due to market monopolies by large resource providers, fixed pricing strategies, and information asymmetry, both consumers and providers exhibit a low willingness to engage in resources trading. To solve these challenges, the paper proposes an incentive mechanism for trust-driven resources trading to guarantee trusted and efficient resources trading. We first design a trust guarantee scheme based on reputation evaluation, blockchain, and trust threshold setting. Then, the proposed incentive scheme can dynamically adjust prices and enable the platform to provide appropriate rewards based on providers’ classified types and contributions. We formulate an optimization problem aiming at maximizing the trading platform’s utility and obtaining an optimal contract based on individual rationality and incentive compatible constraints. Simulation results verify the feasibility and effectiveness of our scheme, highlighting its potential to reshape the future of computing resource management, increase overall economic efficiency, and foster innovation and competitiveness in the digital economy. Renchao Xie, Wen Wen 0011, Qinqin Tang, Xiaodong Duan, Lu Lu 0016, Tao Sun 0010, Tao Huang 0005, F. Richard Yu |
IEEE Trans. Netw. Serv. Manag. | 8 |
| 2025 | RoCELet: Host-Based Flowlet Load Balancing for RoCEabstractRemote Direct Memory Access (RDMA) is becoming a popular high-speed networking technology. It uses kernel bypass and zero copy to achieve high throughput and low latency with little CPU overhead. However, standard RoCE transmission uses Equal Cost Multipath (ECMP) for load balancing, which can result in lower transmission performance due to hash conflicts. Meanwhile, it has been verified that, unlike TCP, the unique retransmission mode and flow characteristics of RoCE make previous load balancing algorithms not well applied to RoCE. In this paper, we introduce RoCELet, a load balancing algorithm for RoCE. It achieves fine-grained RoCE load balancing by actively generating flowlets, effectively utilizing the rich end-to-end paths in the data center. We implement a prototype based on DPDK and evaluate it through small-scale testbed experiments and large-scale simulations. Our results show that compared to state-of-the-art load balancing algorithms, RoCELet optimizes 48.2% and 16.4% in average FCT and$99^{th}$-ile FCT, respectively. Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Jiafeng Jiang, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
IEEE Trans. Netw. | 7 |
| 2025 | RHCC: Revisiting Intra-Host Congestion Control in RDMA NetworksabstractRDMA has been widely deployed in production datacenters. The conventional wisdom believes that the intra-host network delivers stable and high performance. However, intra-host resources witness a relative stagnation in technology trends compared to the evolving RDMA NIC (RNIC). Thus, the RNIC traffic may not get sufficient intra-host resources when it contends with CPU-to-memory traffic. A line of recent works from large-scale production datacenter operators demonstrates the emergence of intra-host congestion and associated performance collapse, which forces us to revisit the practice of intra-host congestion control. However, the ability to efficiently control RDMA intra-host networks is far less mature than inter-host networks, which brings challenges in congestion monitoring, intra-host resource allocation and RNIC traffic adjustment. In this paper, we propose RDMA intra-Host Congestion Control (RHCC), which combines CPU-to-memory traffic congestion avoidance with sub-RTT granularity and proactive RNIC traffic adjustment. RHCC ensures fast congestion avoidance and can work with different inter-host congestion control methods. We implement RHCC on commodity servers and RNICs and conduct experiments to evaluate the performance. The results show that RHCC can increase/decrease the network throughput/latency by up to 2$\times$and 1.4$\times$, respectively. Zirui Wan, Jiao Zhang 0002, Yuxiang Wang 0011, Kefei Liu 0004, Haoyu Pan, Yongchen Pan, Tao Huang 0005 |
IEEE Trans. Netw. | 7 |
| 2025 | Re-Architecting Traffic Control in Cross-Datacenter RDMA NetworksabstractThe network-intensive applications, like machine learning and cloud storage, are increasingly driving two critical trends:1)RDMA has been widely deployed to provide high-speed networks;2)applications are distributively deployed across multiple regional datacenters to satisfy demands for content providers and customers. To fully utilize the benefits of RDMA, we desire to extend it to support cross-datacenter networks. However, the long-haul transport suffers a considerably long control loop, and thus the hybrid of long-haul and intra-datacenter traffic can easily cause severe congestion. We revisit existing traffic control methods and find they are insufficient to resolve this hybrid traffic congestion. Generally, regional datacenters are connected using dedicated long-haul optical fiber and datacenter interconnection (DCI) switches. In this paper, we propose Approach Traffic Control (ATC), a novel solution focusing on two-side DCI-switches (i.e., the approach point for datacenters) to separately alleviate the hybrid traffic congestion in the local and distal datacenters, as a building block for host-driven control methods. This design principle helps ATC shorten the control loop to a single datacenter scale while aggregating congestion information of the whole datacenter range with minor deployment complexity. We implement ATC on P4-based switches and conduct evaluations using real-world testbeds and large-scale NS3 simulations. The results show that ATC ensures fast congestion avoidance and delivers significant performance. For example, ATC reduces the FCT of intra-datacenter and long-haul traffic by up to 88% and 52%, respectively. Zirui Wan, Jiao Zhang 0002, Yuzhen Su, Haoyu Pan, Mingxuan Yu, Tao Huang 0005 |
IEEE Trans. Netw. | 7 |
| 2024 | Hostmesh: Monitor and Diagnose Networks in Rail-optimized RoCE ClustersabstractRoCE services are sensitive to failures and bottlenecks, which become more common as the RoCE network scales. To effectively detect and locate these problems independent of service traffic, RoCE networks require a monitoring and diagnostic system based on active probing. However, existing active probing schemes typically rely on a controller to design the probing plan for each server, which is difficult to deploy and has high synchronization overhead in multi-tenant clusters. Fortunately, rail-optimized clusters have become more common in recent years to improve network performance. In these clusters, the controller is unnecessary. Kefei Liu 0004, Jiao Zhang 0002, Zhuo Jiang, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Zicheng Wang 0004, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
APNet | 15 |
| 2024 | Rethinking Intra-host Congestion Control in RDMA NetworksabstractRDMA has been widely deployed in production datacenters. The conventional wisdom believes that the intra-host network delivers stable and high performance. However, intra-host resources witness a relative stagnation in technology trends compared to the evolving RDMA NIC (RNIC). Thus, the RNIC traffic may not get sufficient intra-host resources when it contends with intra-host traffic. A line of recent works from large-scale production datacenter operators demonstrates the emergence of intra-host congestion and associated performance collapse, which forces us to rethink the practice of intra-host congestion control. However, the ability to efficiently control RDMA intra-host networks is far less mature than inter-host networks, which brings challenges in congestion monitoring, intra-host resource allocation and RNIC traffic adjustment. In this paper, we propose RDMA intra-Host Congestion Control (RHCC), which combines sub-RTT granularity intra-host traffic congestion avoidance and proactive RNIC traffic adjustment. We implement RHCC on commodity servers and RNICs and conduct experiments to evaluate the performance. The results show that RHCC can increase/decrease the network throughput/latency by up to 2 × and 1.4 ×, respectively. Zirui Wan, Jiao Zhang 0002, Yuxiang Wang 0011, Kefei Liu 0004, Haoyu Pan, Tao Huang 0005 |
APNet | 6 |
| 2024 | Spatiotemporal Task Scheduling for Green Computing in Computing Power NetworksabstractRecently, the advancement of information technologies have accelerated the generation of big data, necessitating substantial computing power. This has spurred the development of Computing Power Networks (CPNs), which can overcome the limitations of computing power isolation. However, CPNs consume significant energy and produce large carbon emissions during big data processing. Therefore, an energy-efficient task scheduling scheme, coupled with the utilization of renewable energy, appears to be particularly necessary. Nevertheless, the interplay between computing power and networks, and the spatiotemporal variations in green CPNs pose a challenge to designing the task scheduling scheme. In this paper, we propose a transferable spatiotemporal task scheduling scheme with a triple selection of CPN nodes, routing paths, and forwarding time of tasks. The scheme can overcome the dynamics of green CPNs, and jointly optimize the energy consumption and carbon emissions with ensuring delay constraints and long-term load balancing. Then, we present a task scheduling algorithm based on improved nondominated sorting genetic algorithm-II (NSGA-II) to solve the problem, and numerical results demonstrate that our scheme is effective in reducing the overall energy consumption and carbon emissions of CPNs. Wen Wen 0011, Renchao Xie, Qinqin Tang, Zehui Xiong, Gaochang Xie, Tao Huang 0005 |
GLOBECOM | 6 |
| 2024 | In-band Network-Wide Telemetry for Topology-Varying LEO Satellite NetworksabstractDriven by technological advances and new business models, we have seen a renewed interest in LEO satellite constellations. The deployment of large-scale LEO satellite networks is becoming a reality. The network topology changes periodically, as satellites orbit the Earth. This imposes a great challenge to network monitoring. Meanwhile, as a new network monitoring method, In-band Network Telemetry (INT) can provide per-hop granular telemetry metadata, needed to tackle the mobile nature of LEO satellite constellations. Given this, we apply INT to LEO satellite networks for real-time fine-grained monitoring. We propose a path planning solution to identify the paths for network-wide telemetry and the paths for disseminating the telemetry data to the ground facilities. By taking advantage of the predictable satellite trajectories and topology variations, the path planning solution is designed to achieve network-wide coverage and minimize telemetry overhead. We take the LEO48 constellation as an example to visually show the detailed paths of the monitoring scheme. We conduct experiments on different sizes of networks to evaluate the original path planning algorithm and the improved balanced algorithm in this paper, demonstrating the timeliness and balance of the telemetry solution. Yan Zhang 0063, Tian Pan 0001, Qiang Fu 0011, Jiang Liu 0010, Haipeng Yao, Tao Huang 0005 |
GLOBECOM | 8 |
| 2024 | RACC: Rapid and Accurate INT-Based Congestion Control in RDMA NetworkabstractWith the rapid growth of network speed, datacenter applications have increasing demands on networks for high throughput and ultra-low latency. RDMA is widely deployed due to its high performance. Most existing RDMA congestion control schemes employ end-to-end architecture with inherent feedback delays of at least one Round-Trip Time (RTT).To overcome these limitations, we propose RACC, a rapid and accurate INT-Based RDMA congestion control scheme. The switches directly provide INT feedback, reducing the feedback signal latency. In addition, the adaptive rate update mechanism is adopted at the sender, which enhances the rapid response to network congestion and the stable control of in-flight bytes. We conduct simulation experiments based on the Fat-Tree topology to analyze the requirements for datacenter performance metrics including convergence, fairness, and dynamic queues. Our evaluations show that the peak queue lengths are reduced by up to 80% compared to HPCC and PowerTCP, and the convergence time after congestion is reduced by half. Yanzhe Zhao, Shuo Wang 0006, Guoyu Peng, Tao Huang 0005 |
GLOBECOM | 6 |
| 2024 | D-Router: Decoupled Content Routers with Remote Content StoreabstractNamed Data Networking (NDN) enables efficient content distribution through in-network caching. However, the additional states of network intermediary nodes make NDN forwarding more burdensome, and the unpredictability of cache hits during forwarding leads to uncertain content retrieval latency. To overcome performance bottlenecks at the router's data plane and enhance network determinism, we propose the decoupled content router with remote content store (D-Router). This novel architecture decouples the local content store (CS) from routers and introduces the remote CS device for pooling important content. When Interest packets arrive at a router whose CS is overloaded, we ensure determinism by forwarding them to the remote CS for processing if the requested content is cached there, preventing blocking before the local CS of routers and potential random cache hits along the forwarding path. The dual-path bypass forwarding is supported through the design of routers and a dual-path routing protocol. D-Router is compatible with traditional NDN. Experiments show notable enhancements in data plane performance, including a 30% reduction in round-trip time (RTT), a 25% increase in throughput, improved determinism, and reduced network jitter. Additionally, the decoupling of CS makes it easier for network administrators to deploy network upgrades. Tian Pan 0001, Chunyang Wu, Guohao Ruan, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 6 |
| 2024 | Dual-timescales Optimization for Resource Slicing and Task Scheduling in Satellite Edge Computing NetworksabstractThis paper establishes a dual-timescale framework for joint resource slicing and task scheduling in satellite edge computing (SEC) networks. Specifically, to capture network dynamics and task stochasticity at small timescales, we formulate the task scheduling problem as a Markov decision process (MDP) to minimize task delay, network energy consumption, and packet loss. We design a deep reinforcement learning-assisted task scheduling (DRTS) algorithm inspired by the soft actor-critic (SAC) algorithm to learn the scheduling policy. Task processing performance is affected by communication and computing re-sources allocated to respective resource slices. Thus, considering that frequent resource slicing has a significant management over-head, we further optimize resource slices on a larger timescale. To obtain a policy with low complexity, we propose a greedy-based heuristic algorithm. A hierarchical solution is constructed to find the optimal solution due to the correlation between the two timescale problems. Finally, to validate the effectiveness and superiority of the proposed scheme, extensive simulations are performed. Zeru Fang, Qinqin Tang, Renchao Xie, Tao Huang 0005, Tianjiao Chen, F. Richard Yu |
ICC | 4 |
| 2024 | Contract Theory-Based Customized Service Scheduling for Predictable QoS in WANsabstractService Customized Networking (SCN) is emerging as an escalating technological trend to address the personalized requirements of services, which are plagued in traditional “best-effort” transmission networks. Organically coordinating heterogeneous domains in Wide Area Networks (WANs) is essential for establishing end-to-end customized service delivery. However, the peer-to-peer centralized communication mode between Au-tonomous Domains (ADs) hinders their connectivity and makes it challenging to support diverse intra-domain routing protocols for the desired Quality of Service (QoS). To achieve customized service scheduling for predictable QoS in WANs, we propose a contract theory-based incentive mechanism. In specific, we first select trusted ADs with high service qualities by computing their reputations through a subjective logic model. The Service Provider (SP) decomposes the overall QoS requirements by domains logically from a global perspective. These decomposed QoS metrics will splice differentiated service capabilities from ADs to obtain the expected end-to-end connection. To address information asymmetry between the SP and ADs, we formulate contribution-reward contract items and devise an optimization problem of maximizing the whole system utility. The optimal contract problem is solved through constraints of individual rationality and incentive compatibility. Simulation results indi-cate the feasibility and effectiveness of our scheme on service customization and economic benefits. Tao Huang 0005, Sha Tan, Qinqin Tang, Renchao Xie, F. Richard Yu |
ICC | 1 |
| 2024 | Accelerating Mega-Scale Satellite Network Simulation in NS-3 via MPI-based ParallelizationabstractDue to the high costs of low Earth orbit (LEO) satellite manufacturing and launch, as well as the complexity of in-orbit network protocol debugging, simulating and verifying satellite network protocols on the ground before satellite launch holds significant importance. Compared to the expensive emulation with one-to-one replication, simulation (e.g., using ns-3) can achieve discrete event processing at a relatively lower cost by extending the wall clock time. However, very few studies have used ns-3 for LEO satellite network simulation, facing challenges such as faithfully simulating the on/off state switching of inter-satellite links (ISLs) and achieving simulation performance scalability for high-density satellite constellations. In this work, we propose a system to accelerate mega-scale LEO satellite network simulation in ns-3 via MPI-based parallelization. Specifically, we simulate ISLs based on ns-3's P2P channels/P2P remote channels and achieve runtime link connection/disconnection by implementing stateful traffic dropping inside the network interface. Then, we conduct concurrent simulation with ns-3's parallel and distributed simulation capability and partition the satellite constellation into multiple simulation processes through a hierarchical clustering algorithm and automated scripts, considering satellite locality and inter-process workload balance. Our evaluation shows significant speed improvements via parallelization, e.g., a 373% speedup with 12 processes for LEO-192, and a 156% speedup with 3 processes for LEO-3072. Haibin Song, Tian Pan 0001, Guohao Ruan, Ying Wan 0001, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 7 |
| 2024 | A VLAN-based Network Testbed for Lightweight Satellite Constellation EmulationabstractConsidering the high costs of satellite manufacturing and launch, as well as the complexity of in-orbit debugging, pre-launch emulation on the ground will significantly reduce the development costs of low Earth orbit (LEO) satellite networks. LEO satellite network emulation faces challenges in emulating mega-scale constellations in a lightweight and scalable manner, as well as efficiently handling the frequent link on/off switching for both inter-satellite networks and terrestrial access networks. Existing simulation/emulation tools, such as NS-3, Mininet, QualNet, fall short in addressing these issues effectively. In this work, we propose a lightweight satellite emulation testbed based on Docker containers and the VLAN protocol. In the data plane, our testbed uses Docker containers to emulate satellites/terminals, and uses VETH-pairs and bridges to emulate inter-satellite networks and terrestrial access networks. Furthermore, these virtual network elements can horizontally scale across multiple servers for mega-scale constellation emulation. In the control plane, the real-time constellation topology changes are efficiently emulated through the configuration of VLAN segmentation according to satellite movement patterns. Evaluation shows the testbed's low resource occupancy and high efficiency, with 100 nodes consuming only 1000MB memory and 100 links switching in less than 2s. Tian Pan 0001, Yan Zhang 0063, Jiang Liu 0010, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 5 |
| 2024 | Gaia: Ground Station-Centric Mobility Management for LEO Satellite NetworksabstractFor mega-scale low Earth orbit (LEO) satellite constellations, the relative position changes between satellites and ground terminals pose challenges for end-to-end TCP session maintenance, which serves as the substrate for many Internet services. Mobile IP resolves the session maintenance issue by in-troducing a binding mechanism between the care-of address and home address; however, this also leads to inefficient triangular routing. Our recently proposed LISP-LEO, through partition-satellite mapping, routes traffic to the service satellite above the destination terminal's partition, addressing the triangular routing problem. However, due to the corner case of partition-satellite mapping, LISP-LEO introduces the issue of the last-hop route selection, as well as the associated per-terminal registration states on the satellite, making the solution non-scalable. In this work, we propose Gaia, a ground station-centric mobility management scheme for LEO satellite networks. Gaia maintains a precise mapping of each ground terminal's IP and its geographical location. When receiving traffic from a source terminal, the access satellite can, based on the geographical location of the destination terminal carried by the traffic, directly locate the satellite above the destination terminal and tunnel the traffic to it. In addition, to reduce the satellite's burden, we add a DNS-like querying mechanism by offloading the mapping of terminal IPs and geographical locations to the ground station. The evaluation shows that Gaia outperforms Mobile IP and LISP-LEO in both end-to-end latency and on-board resource consumption. Xiaxin Zhou, Tian Pan 0001, Zhaokun Yang, Sirui Su, Guohao Ruan, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 8 |
| 2024 | BiCC: Bilateral Congestion Control in Cross-datacenter RDMA NetworksabstractWith the development of network-intensive applications like machine learning and cloud storage, there are two growing trends: (i) RDMA has been widely deployed to enhance underlying high-speed networks; (ii) applications are deployed on geographically distributed datacenters to meet customer demands (e.g., low access latency to services or regular data backups). To fully utilize the benefits of RDMA, we desire to support long-haul RDMA transport for cross-datacenter applications. Different from common intra-datacenter communications, the hybrid of long-haul and intra-datacenter traffic complicates the congestion state, and the considerably long control loop makes it more severe. We revisit existing congestion control methods and find they are insufficient to address the hybrid traffic congestion.Note that regional datacenters are connected by dedicated long-haul optical fiber and datacenter interconnection (DCI) switches directly. In this paper, we propose Bilateral Congestion Control (BiCC), a novel solution relying on two-side DCI-switches to bilaterally alleviate the hybrid traffic congestion in the sender-side and receiver-side datacenter while serving as a building block for existing host-driven methods. BiCC can shorten the control loop to a single datacenter scale and aggregate congestion information across the whole datacenter. We implement BiCC on commodity P4-based switches and conduct evaluations using both testbed experiments and NS3 simulations. The extensive evaluation results show that BiCC ensures fast congestion avoidance. Thus, BiCC reduces the average FCT for intra-datacenter and inter-datacenter traffic by up to 53% and 51%, respectively, in large-scale simulations. Zirui Wan, Jiao Zhang 0002, Mingxuan Yu, Xinghua Zhao, Tao Huang 0005 |
INFOCOM | 7 |
| 2024 | SARO: Intelligent Data-Driven Routing Optimization for LEO Satellite NetworksabstractThe rapid development of satellite networks has precipitated an increasing demand for high bandwidth. However, compared to terrestrial networks, bandwidth resources in satellite networks are often more constrained. Hence, reasonable traffic scheduling is crucial. Because satellite networks present complex network structures, uneven traffic patterns, and strong dynamics, traditional traffic scheduling algorithms often fail in accurate analysis and modeling. Numerous intelligent routing schemes based on learning have been proposed and validated. However, due to the limitations of network modeling and generalization, it is difficult for them to quickly adapt to changing network conditions. In this paper, we propose SARO, an intelligent routing algorithm that integrates Deep Reinforcement Learning (DRL), Graph Neural Networks (GNN), and In-band Network Telemetry (INT). Compared to traditional deep learning algorithms, SARO not only overcomes the challenge of acquiring network features but also addresses the issue of limited generalization performance. Experiments show that regardless of changes in topology, SARO’s maximum link utilization is reduced by 7.6% to 15.6% compared to baseline algorithms, demonstrating SARO performs excellently in terms of load balancing and generalization. Jiao Zhang 0002, Tian Pan 0001, Tao Huang 0005 |
ISCC | 4 |
| 2024 | FTA-detector: Troubleshooting Gray Link Failures Based on Fault Tree AnalysisabstractDetecting link failures is critical to ensuring the operation of data center networks (DCNs). However, some gray link failures may go undetected by switches, leading to silent packet drops. In this paper, we propose FTA-detector, a gray link failure detection and localization approach leveraging Fault Tree Analysis (FTA), a technique previously applied in the field of reliability engineering. On the data plane, we collect fine-grained hop-by-hop information through In-band Network Telemetry (INT), detect the bidirectional connectivity of end-to-end paths through a novel aging mechanism, and implement fast reroute in response to gray link failures. On the control plane, we introduce a faulty link localization algorithm based on FTA to recommend the most likely faulty links. Specifically, we use Top K and progressive failure repair to discover and repair link faults as early as possible during failure inference, significantly reducing the overall computation complexity of sequential root cause analysis. For large-scale network topology, we propose a divide and conquer optimization scheme for scalability. To verify the efficiency of our system, we build a virtual network test platform with P4 switch software and Redis database. The test results show that FTA-detector can troubleshoot multi-point failures in DCNs in a very short time with high accuracy. Yan Zou, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Qingqiang Yi, Ying Wan 0001, Jiao Zhang 0002, Tao Huang 0005 |
NOMS | 8 |
| 2024 | LuoShen: A Hyper-Converged Programmable Gateway for Multi-Tenant Multi-Service Edge Clouds
Tian Pan 0001, Xionglie Wei, Yisong Qiao, Tiesheng Cheng, Wenqiang Su, Yuke Hong, Zhengzhong Wang, Chongjing Dai, Peiqiao Wang, Xuetao Jia, Jianyuan Lu, Enge Song, Biao Lyu, Ennan Zhai, Jiao Zhang 0002, Tao Huang 0005, Dennis Cai, Shunmin Zhu |
NSDI | 23 |
| 2024 | R-Pingmesh: A Service-Aware RoCE Network Monitoring and Diagnostic SystemabstractRoCE services are sensitive to network failures and performance bottlenecks, which become more common as the RoCE network scales. In addition, some non-network problems behave like network problems and can waste troubleshooting time. However, existing mechanisms cannot quickly detect and locate network problems or determine whether the service problem is network-related. Kefei Liu 0004, Zhuo Jiang, Jiao Zhang 0002, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang 0003, Haohan Xu, Dongyang Song, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
SIGCOMM | 19 |
| 2024 | Online Resource Allocation for Large-scale Deterministic Networks with Historical DataabstractDeterministic IP (DIP) networking is a time division duplex technique providing delay-bounded transmissions in large-scale networks. Dedicated resources are allocated to each time-sensitive (TS) flow, reducing queuing delay uncertainties. However, most research relies on offline scheduling, unfit for dynamic networks. Some online heuristics prioritize early flows, risking later traffic access issues. To tackle this, we propose an online algorithm using historical data for better foresight. We develop and analyze a competitive algorithm considering historical data, demonstrating its effectiveness through simulations. Weiqian Tan, Binwei Wu, Shuo Wang 0006, Tao Huang 0005 |
VTC Spring | 4 |
| 2024 | A Secure and Efficient State Channels-Based Network Service Business Settlement SchemeabstractWith the continuous innovation of network technology, emerging network service models have begun to be proposed, which also puts forward higher requirements for business settlement. Business settlement, as the foundation of network services, involves the interests of users and service providers. The traditional centralized settlement schemes no longer respond to the needs of both parties in terms of transparency, security, and fairness. Therefore, some blockchain-based settlement solutions have been proposed to address these issues. However, due to the additional overhead brought by blockchain, it is difficult for these solutions to ensure both security and efficiency. In this paper, we propose a state channels-based business settlement scheme (SCBS) to complete network service settlement securely and efficiently. In SCBS, the settlement process can be effectively carried out off-chain. When non-cooperative behavior occurs, both parties can create, resolve, and refute the dispute through the state channel to assure safety. Finally, the test results from Hyperledger Fabric platform demonstrate the feasibility and effectiveness of our SCBS. Wei Chen 0131, Ru Huo, Shuo Wang 0006, Tao Huang 0005 |
WCNC | 4 |
| 2024 | Delay-Prioritized Task Scheduling with Load Balancing in Computing Power NetworksabstractIn the era of data-driven intelligent Internet, efficient utilization of computing power is paramount. Yet, current cloud-edge collaboration architectures face challenges with computing power isolation, adversely affecting efficiency and user experience. Computing Power Networks (CPNs) leverage networks with cloud-native applications to connect and manage resources, offering a blueprint for a collaborative computing ecosystem. In the CPN, scheduling stands as a pivotal function. However, conventional scheduling often neglects the interplay between computing and networks. To rectify this, we present a collaborative task scheduling system in CPNs that simultaneously contemplates the selections of computing nodes and network links. Aiming to maintain a balanced load for both computing and network resources, we formulate the scheduling challenge as a Constrained Markov Decision Process (CMDP). This approach focuses on optimizing both execution delay and success rate of computing tasks with load balancing constraints in CPNs. To facilitate the resolution of the CMDP, we introduce a Lyapunov-optimized Deep Reinforcement Learning (DRL) algorithm, which reconfigures the long-term constraint into immediate optimization. We provide numerical results to demonstrate the effectiveness of our suggested policy and algorithm. Renchao Xie, Qinqin Tang, Tao Huang 0005 |
WCNC | 4 |
| 2024 | Multi-path CQF for Low-Jitter and High-Reliable Packet Delivery in Time-Sensitive NetworksabstractTime-Sensitive networking (TSN) has put forward a series of standards, such as cyclic queuing and forwarding (CQF) and frame replication and elimination for reliability (FRER), to achieve deterministic latency and high reliability. However, most work studies these two mechanisms separately, while directly combining CQF and FRER (DCCF) will inevitably introduce distinct multiple-path delays, seriously impair scheduling capa-bilities and result in a large jitter. In this paper, we propose a Multi-path CQF (MCQF) mecha-nism. Firstly, MCQF enables flexible end-to-end delay calculation by extending the ping-pong queues of CQF to multi-queues. Then, we formulate a joint routing and scheduling mathematical model to maximize the number of schedulable flows and satisfy diverse latency and reliability requirements. Moreover, a hop-by-hop offset scheduling (HOS) algorithm is designed to achieve low jitter by aligning the packet delays on multiple disjoint paths. Evaluation results show that MCQF performs better than CQF on reliability. Compared to DCCF, MCQF greatly reduces the jitter and improves the schedulable flow number by about 31.9 %. Yudong Huang, Shuo Wang 0006, Guizhen Li, Xinyuan Zhang 0011, Dongran Xu, Tao Huang 0005 |
WCNC | 6 |
| 2024 | Blaze: Delay-Aware Cloud-Edge Collaborative Service Function Chain Deployment with Network CalculusabstractWith the rapid development of Internet of the Things (IoT) technology, IoT services have higher and higher requirements for latency. In the IoT environment, virtual network functions (VNFs) are deployed on general-purpose hardware and are sequentially connected to form service function chain (SFC) to provide network services for IoT devices. However, the high latency of the link between the cloud center and the edge nodes and the resource capacity limitation of the edge nodes pose challenges to the deployment of SFCs in IoT devices. In this paper, we study the cloud-edge collaborative SFC deployment problem. We applied the network calculus theory to the cloud-edge collaborative SFC deployment for the first time, aiming to provide the end-to-end delay guarantee for the deployed SFC. We model the SFC deployment problem as Mixed Integer Nonlinear Programming (MINLP). Then we propose a heuristic algorithm (Blaze) to solve this problem. Blaze is proven to complete the deployment of SFCs in polynomial time. Finally, the algorithm is evaluated by experimental simulation. The experimental results show that compared with the existing state-of-the-art corresponding algorithms, the proposed algorithm achieves better performance in terms of the number of VNFs deployed in the cloud, resource consumption of edge nodes, and SFC request acceptance rate. Huimin Luo, Jiao Zhang 0002, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
WCNC | 5 |
| 2024 | Joint Transmission and Transcoding in Computing Power Networks for Livecast: A Quantum-inspired Optimization ApproachabstractWith the continuous development of network technology and hardware devices, panoramic live cast is promising and widely used in various industries. To meet massive heterogeneous viewer demands, live cast video streams must be transcoded into multiple versions and then transmitted to viewers. By offloading transcoding workloads to network nodes closer to broadcasters and viewers, computing power networks (CPN) have been considered an effective means to provide viewers with a higher quality of experience (QoE). In this paper, we study the joint video transmission and transcoding resource allocation in the CPN-based panoramic livecast system. Considering the versatility and simplicity, we offer a multi-layer network model to capture the stochastic characteristics of transmission and transcoding processes and transform the joint resource allocation problem into a broader shortest path problem (SPP). As the scale of networks expands, the SPP may become intractable on classical computers. This paper explores the viability of solving the SPP on quantum computers by utilizing quantum resources such as superposition and entanglement, then proposes a shortest path algorithm based on the quantum approximate optimization algorithm (QAOA-SPP) for jointly optimizing the latency and overhead of transmission and transcoding. Simulation results indicate that our algorithm can achieve better performance than the benchmark schemes under reasonable parameter settings. Renchao Xie, Qinqin Tang, Tao Huang 0005 |
WCNC | 4 |
| 2024 | DAmpADF: A framework for DNS amplification attack defense based on Bloom filters and NAmpKeeper
Yunwei Dai, Tao Huang 0005, Shuo Wang 0006 |
Comput. Secur. | 2 |
| 2024 | Joint Task Dispatching and Bandwidth Allocation with Hard Deadlines in Distributed Serverless Edge Computing Systems
Tao Huang 0005 |
J. Grid Comput. | 3 |
| 2024 | Hirail: Core-Agnostic Deterministic Networks for Long-Distance Time-Sensitive IIoT ApplicationsabstractWith the emergence of time-sensitive IIoT applications, such as remote operation and industrial control, a long-distance deterministic forwarding service is highly desirable. However, most of the existing research is limited to local area networks, or requires costly replacement of core network devices. Enabling incremental deterministic networks based on off-the-shelf technologies is a significant challenge. This paper designs a core-agnostic and cost-effective solution named Hirail to achieve the smooth evolution of long-distance deterministic networks. Firstly, we investigate that a time-discrete shaper (TDS) can be deployed at the ingress node to enable millisecond-level bounded delay. TDS functions similarly to the concept of buying time-stamped tickets for each flow prior to getting on a high-speed rail, thus avoiding the expensive modification of core devices. Then, to alleviate the flow aggregation problem under long-distance links, we utilize the inband network telemetry to construct the delay-aware network map and conduct adaptive source routing based on the map. Finally, an adjustable buffer at the last hop is devised for jitter reduction. Evaluation results show that Hirail can meet the bounded delay and jitter demands, and outperforms other solutions in terms of performance and overhead. Tao Huang 0005, Yudong Huang, Xinyuan Zhang 0011, Shuo Wang 0006, Hongyang Du 0001, Dusit Niyato, F. Richard Yu, Yunjie Liu 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Network Coding-Based Multipath Transmission for LEO Satellite Networks With Domain ClusterabstractIn the large-scale dynamic Low Earth Orbit (LEO) satellite networks, the conventional TCP-based single-path transmission encounters challenges such as prolonged propagation delay, frequent connection failures, and suboptimal resource utilization. In this paper, we propose an Integrated Multi-Path Network Coding (IMPNC) transmission scheme. This scheme leverages multiple paths for end-to-end transmission to achieve bandwidth aggregation and redundant backup. The multi-path transmission is facilitated by Multi-Path Quick UDP Internet Connection (MPQUIC) protocol to adapt to the limited satellite bandwidth and caching resources. The proposed approach involves encoding packets at nodes along the paths, addressing the significant out-of-order problem arising from variable delays on different paths. Additionally, we present a Software Defined Networking (SDN)-based domain clustering architecture, which offers a more streamlined control approach, reducing overall complexity. Furthermore, we formulate the domain clustering problems as mixed-integer nonlinear programming and the coding-based routing problem as a Steiner tree problem. Evaluation results demonstrate that the proposed scheme effectively reduces the latency over 25.1%, enhances bandwidth utilization by 19.6%, and ensures reliable data transmission by reducing retransmission probability by 4.1%. Man Ouyang, Ran Zhang 0004, Jiang Liu 0010, Tao Huang 0005, Jincheng Tong, Ning Xin, F. Richard Yu |
IEEE Internet Things J. | 5 |
| 2024 | AIIN: An APN-Integrated Approach Toward Reactive Telemetry Notification for IFITabstractIn situ flow information telemetry (IFIT) is a state-of-the-art in-band telemetry framework for operator networks and can serve as the information foundation for network intelligence in the emerging sixth-generation regime. However, the performance advantage of IFIT comes at the cost of excessive telemetry data notification overhead, which makes it challenging to promote IFIT extensively. Therefore, we propose an approach named APN-integrated IFIT information notification (AIIN) to provide data notification overhead adaptability to IFIT. AIIN introduces a requirement-aware capability and reactive differentiated treatment into IFIT. In AIIN, we first enhance application-aware networking (APN) and integrate it into IFIT notification to support the explicit expression of data notification requirements. Then, oriented toward the different timeliness (telemetry data lag time) and accuracy (telemetry data retention rate) requirements expressed in APN, we design different behavioral treatment models to define reactive functions and procedures to make network devices explicitly process these requirements without decisions. The AIIN prototype is implemented on P4 switches. We also deploy the prototype on the China Environment for Network Innovation (CENI) network. Emulation results show that AIIN can achieve nanosecond line speed performance with differentiated and reactive data notification overhead reduction and, in the best case, can reduce bandwidth occupation by approximately 84%. Weihong Wu, Jiang Liu 0010, Jianwei Mao, Shuping Peng 0001, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Internet Things J. | 7 |
| 2024 | FastTS: Enabling Fault-Tolerant and Time-Sensitive Scheduling in Space-Terrestrial Integrated NetworksabstractThe emerging space-terrestrial integrated network (STIN) assumes a pivotal role within the 6G vision, promising to deliver seamless global coverage and connectivity. Achieving advanced, high-reliability, and time-sensitive (TS) services in a resource-constrained and failure-prone space environment is critical, but also presents challenges. Existing space-terrestrial communication approaches either suffer from temporary link failures with unstable reliability, or intolerable service latency due to the extensive coverage and uneven traffic distribution. This paper presents FastTS, a heuristic resilient and performant scheduling strategy to achieve fault-tolerant and time-sensitive scheduling in futuristic STINs. First, we model the high-dynamic and failure-prone topology in space, and formulate the scheduling problem as a mixed non-linear problem with the objective of minimizing the average task completion time. To approach the optimal solution, joint time-variant routing and frame replication and elimination for reliability (FRER) redundancy under resource constraints are formally considered in our design. During the path-stable duration, FastTS prioritizes the multipath selection with higher redundancy scores, all while ensuring a bounded low latency for TS services based on time-sensitive networking (TSN) techniques. Specifically, our FastTS is divided into three phases: time-sensitive multipath generation (TMG), series-parallel redundancy scoring (SPRS), and SPRS-based time-variant routing (STR). Finally, simulation results show that FastTS exhibits outstanding performance improvements in terms of packet delay, scheduling success ratio, task completion time and packet loss rate, when compared to other state-of-the-art methods. Guoyu Peng, Shuo Wang 0006, Tao Huang 0005, Fengtao Li, Kangzhe Zhao, Yudong Huang, Zehui Xiong |
IEEE J. Sel. Areas Commun. | 3 |
| 2024 | Joint Service Deployment and Task Scheduling for Satellite Edge Computing: A Two-Timescale Hierarchical ApproachabstractIn this paper, we establish a two-timescale framework for the joint service deployment and task scheduling problem in satellite edge computing networks.We aim to optimize the computing performance of networks with diverse quality-of-service (QoS) guarantees for computing tasks. Specifically, to capture the small-timescale network dynamics and task randomness, we formulate the task scheduling problem as a constrained Markov decision process (CMDP) to minimize the energy consumption, load imbalance and packet loss of networks while ensuring the long-term delay. The Lyapunov technique is employed to deal with the delay constraints. A soft actor-critic (SAC)-based deep reinforcement learning (DRL) framework is designed to learn the stationary scheduling policy. We further explore the significant impact of deploying diverse services on the performance of task scheduling in satellite edge computing. Considering that frequent deployment of services will incur huge deployment overhead, we optimize the service deployment on a larger timescale. The optimization problem is modeled as an integer programming problem to improve the service capability of networks and reduce service deployment costs. A heuristic-based atomic orbital search (AOS) approach is proposed to obtain the superior policy with low complexity. Due to the correlation between the problems of two timescales, a hierarchical solution is constructed to iteratively find the excellent solution. Finally, extensive simulations are conducted to validate the effectiveness and superiority of the proposed scheme. Qinqin Tang, Renchao Xie, Zeru Fang, Tao Huang 0005, Tianjiao Chen, Ran Zhang 0004, F. Richard Yu |
IEEE J. Sel. Areas Commun. | 4 |
| 2024 | Connected and Autonomous Vehicles in Web3: An Intelligence-Based Reinforcement Learning Approachabstract“Read-write-own” based Web3 has been proposed as a promising user-centric Internet to open the new generation of the World Wide Web, where Web3 users can independently manage data and derive value from creating content without relying on intermediaries. Connected and autonomous vehicles (CAVs) in Web3 can trade models in a self-controlled and decentralized credible way, which is a fundamentally and principally innovation based on novel architecture. Effectively implementing such paradigms involves proper model trading strategies. However, reinforcement learning (RL)-based strategies face challenges of poor generalization ability, low feasibility, and the exploration-exploitation dilemma. It is also difficult to define an explicit and appropriate reward function. Therefore, in this paper, we propose an intelligence-based reinforcement learning (IRL) approach for CAVs in Web3. We present a framework to enable model transactions between CAVs. Also, we provide a decentralized identifier (DID)-based identity management system for resource description and data verification to access Web3, followed by the mechanism and supporting smart contracts. Furthermore, we formulate the model trading issue as an active inference to form higher-level cognition about the environment without rewards. Then we use IRL to solve it. And we use “intelligence”, a high-level indicator, to quantify the efficiency of such cognition. It can evaluate the difference between the predicted state and the real state in policy exploration. The proposed scheme shows good generalization and can auto-balance exploration and exploitation, simultaneously achieving outperforming performance on the model trading issue with no rewards. In simulations, the performance of the proposed scheme is compared with existing methods. Yuzheng Ren, Renchao Xie, F. Richard Yu, Ran Zhang 0004, Yuhang Wang 0019, Ying He 0006, Tao Huang 0005 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Secure incentive mechanism for energy trading in computing force networks enabled internet of vehicles: a contract theory approach
Wen Wen 0011, Lu Lu 0016, Renchao Xie, Qinqin Tang, Yuexia Fu, Tao Huang 0005 |
J. Supercomput. | 6 |
| 2024 | Efficient and Non-Repudiable Data Trading Scheme Based on State Channels and Stackelberg GameabstractAs the Internet of Things gathers pace and popularity, more and more data is collected at the edge. To unleash the value of data and make it tradable, data markets have been proposed. However, existing data markets generally depend on broker or blockchain, which inevitably raises concerns about one or more aspects of fairness, security, or efficiency. In addition, to promote data trading in the data market, a data trading incentive mechanism is also essential. In this paper, we propose a novel data trading scheme based on state channels and Stackelberg game. First, we propose aStateChannels-basedDataTrading (SCDT) framework to support non-repudiable and efficient data trading. The framework can arbitrate disputes arising from off-chain data trading through state channels, enabling traders to conduct efficient transactions off-chain without worrying about security issues. Second, we propose an optimal incentive mechanism to solve the pricing and purchasing problems. The tripartite interactions among the data seller, resource seller, and user service platform are formulated as a Stackelberg game to maximize the profits of all participants. Finally, we implement the data trading framework and analyze the incentive mechanism, which reveals the feasibility of the framework and the rationality of the incentive mechanism. Wei Chen 0131, Ru Huo, Chuang Sun 0002, Shuo Wang 0006, Tao Huang 0005 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Dual-Timescales Optimization of Task Scheduling and Resource Slicing in Satellite-Terrestrial Edge Computing NetworksabstractIn this paper, we optimize network computational performance and ensure diverse quality of service (QoS) for tasks by developing a dual-timescale joint optimization framework for satellite-terrestrial integrated edge computing networks (STECN). In our architecture, STECN can handle intelligent tasks for the Internet of remote things (IoRT) devices based on multiple configured applications deployed. Specifically, we formulate task scheduling as a Markov decision process (MDP) to minimize network energy consumption and task processing delay at small timescales. A deep reinforcement learning (DRL) framework is designed for policy learning. Recognizing the impact of resource slicing on task scheduling in STECN and the deployment overhead from frequent changes, we further optimize resource slicing at larger timescales. To enhance network service capability under dynamic demand, we establish a resource slice gap index, characterizing the difference between actual resources and service demand. By a heuristic-based artificial electric field (AEF) approach, we obtain an optimal strategy with low complexity. Considering the correlation between two timescales, the optimal solution is found by iteratively constructing a hierarchical solution. In addition, to guarantee the global load balancing of the network, we introduce a self-attention mechanism, which allows the knowledge of other satellites to be taken into account when slicing the satellite resources. Finally, extensive simulations confirm the effectiveness and superiority of the proposed scheme. Tao Huang 0005, Zeru Fang, Qinqin Tang, Renchao Xie, Tianjiao Chen, F. Richard Yu |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Coordinating Services and Networks With NaaS Tickets Towards Service Customization in Distributed CloudsabstractDistributed clouds decentralize cloud resources, moving from a single high-level point in the network to multiple low-level points, allowing for the dynamic distribution of services across the “cloud-edge-end”. Nevertheless, the “best-effort” traditional networks suffer from unpredictable service quality and limited collaboration between services and networks. To address these shortcomings, we present a novel solution named “Network-as-a-Service (NaaS) Tickets,” inspired by traffic tickets in transportation systems to empower distributed clouds with customized service capabilities. Specifically, we first propose NaaS Tickets-enabled service-customized distributed clouds (NT-SCDC) to realize on-demand and service-oriented interconnection in a wide area. To establish a solid connection between services and networks, we introduce an auction-driven matching mechanism for NaaS Tickets. Then, the matching problem is formulated via an online framework MatOnline, which translates the long-term market problem into a series of one-shot auctions for NaaS Tickets. Based on the Vickrey-Clarke-Groves (VCG) mechanism, we develop MatVCG algorithm to handle one-shot matching problems, guaranteeing truthfulness, individual rationality, and social welfare. Moreover, we improve the performance of MatOnline to find the minimum feasible scale-down ratio with reduced budget expenditure. Experimental results demonstrate our algorithm achieves a stable competitive ratio on social welfare, effectively meeting customized demands Tao Huang 0005, Sha Tan, Qinqin Tang, Renchao Xie, F. Richard Yu |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | CPPer-FL: Clustered Parallel Training for Efficient Personalized Federated LearningabstractIn this paper, a clustered parallel training algorithm is designed for personalized federated learning (Per-FL), called CPPer-FL. CPPer-FL improves the communication and training efficiency of Per-FL from two perspectives, namely, less burden for the central server and lower interaction idling delay. CPPer-FL adopts a client-edge-center learning architecture, which offloads the central server's model aggregation and communication burden to distributed edge servers. Also, CPPer-FL redesigns the cascading model synchronization and updating procedure in conventional Per-FL and changes it to a parallel manner, thus improving the interaction efficiency in the training process. Further, for the proposed hierarchical architecture, two approaches are proposed to cater to Per-FL: similarity-based clustering for client-edge association and personalized model aggregation for parallel model updating, such that clients' personal features can be preserved in the training process. The convergence of CPPer-FL has been formally analyzed and proved. Evaluation results validate the communication efficiency, model convergence, and model accuracy improvement. Ran Zhang 0004, Fangqi Liu 0002, Jiang Liu 0010, Mingzhe Chen, Qinqin Tang, Tao Huang 0005, F. Richard Yu |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Breaking the Inertial Thinking: Non-Blocking Multipath Congestion Control Based on the Single-Subflow Reinforcement Learning ModelabstractThe Multipath TCP (MPTCP) protocol has received more attention due to the increasing number of terminals with multiple network interfaces. To meet the higher network performance demand of terminal services, many researches leverage reinforcement learning (RL) for MPTCP congestion control (CC) algorithms to improve the performance of MPTCP. However, we observe two limitations of existing RL-based mechanisms that make them impractical: 1) Fail to break the restriction of the input and output dimensions of RL, making the mechanisms unadaptable to the varying number of subflows. 2) Frequent model decisions block packet transmission, leading to under-utilization of bandwidth. This paper breaks the inertial thinking By “inertial thinking” here, we are referring to the initial reaction of others when dealing with CC in MPTCP. Given the interdependence between MPTCP subflows, scholars have traditionally opted for coupled CC. However, we have challenged this conventional thinking by independently handling the CC of different subflows in a single MPTCP flow and ensuring fairness. to overcome the above limitations and proposes Maggey, a non-blocking CC mechanism that applies the single-subflow model to multipath transmission. To this end, Maggey employs loosely coupled design principles and a unique reward function to ensure the fairness of the algorithm. Additionally, Maggey introduces iterative training to ensure the accuracy of training of the single-subflow model. Furthermore, a mode transition framework is artfully designed to avoid blocking, preserving the flexibility of RL-based CCs. These two features enhance the practicability of Maggey and the paper analyze the stability of Maggey. We implement Maggey in the Linux kernel and evaluate the performance of Maggey through extensive emulation and live experiments. The evaluation results show that Maggey boosts 26% throughput over DRL-CC at high bandwidth and improves 2%-60% throughput over traditional algorithms under different network conditions. Besides, Maggey maintains fairness in different scenarios. Dehui Wei, Jiao Zhang 0002, Yuanjie Liu, Tian Pan 0001, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2024 | Diagnosing End-Host Network Bottlenecks in RDMA ServersabstractIn RDMA (Remote Direct Memory Access) networks, end-host networks, including intra-host networks and RNICs (RDMA NIC), were considered robust and have received little attention. However, as the RNIC line rate rapidly increases to multi-hundred gigabits, the intra-host network becomes a potential performance bottleneck for network applications. Intra-host network bottlenecks can result in degraded intra-host bandwidth and increased intra-host latency. In addition, RNIC network problems can result in connection failures and packet drops. Host network problems can severely degrade network performance. However, when host network problems occur, they can hardly be noticed due to the lack of a monitoring system. Furthermore, existing diagnostic mechanisms cannot efficiently diagnose host network problems. In this paper, we analyze the symptom of host network problems based on our long-term troubleshooting experience and propose Hostping, the first monitoring and diagnostic system dedicated to host networks. The core idea of Hostping is to conduct 1) loopback tests between RNICs and endpoints within the host to measure intra-host latency and bandwidth, and 2) mutual probing between RNICs on a host to measure RNIC connectivity. We have deployed Hostping on thousands of servers in our distributed machine learning system. Not only can Hostping detect and diagnose host network problems we already knew in minutes, but it also reveals eight problems we did not notice before. Kefei Liu 0004, Jiao Zhang 0002, Zhuo Jiang, Xiaolong Zhong, Lizhuang Tan, Tian Pan 0001, Tao Huang 0005 |
IEEE/ACM Trans. Netw. | 8 |
| 2024 | PACC: A Proactive CNP Generation Scheme for Datacenter NetworksabstractThe rapid upgrade of link speed and the prosperity of new applications in data center networks (DCNs) lead to a rigorous demand for ultra-low latency and high throughput. To mitigate the overhead of traditional software-based packet processing at end-hosts, RDMA (Remote Direct Memory Access) has been widely adopted in DCNs. Particularly, congestion control (CC) mechanisms designed for RDMA have attracted much attention to avoid performance deterioration when packets lose. However, through comprehensive analysis, we found that existing RDMA CC schemes have limitations of a sluggish response to congestion and unawareness of tiny microbursts due to the long end-to-end control loop. In this paper, we propose PACC, a proactive and accurate switch-driven RDMA CC algorithm with easy deployability. PACC is driven by PI controller-based computation, threshold-based flow discrimination and weight-based allocation at the switch. It leverages real-time queue length to generate accurate congestion feedback proactively and piggybacks it to the corresponding source without modification to end-hosts. We theoretically analyze the stability, convergence and key parameter settings of PACC. Then, we implement PACC in a testbed consisting of DPDK-based end-hosts and Tofino P4 switches. In our evaluation, PACC achieves better fairness, fast reaction, high throughput, and 6$\sim$69% lower FCT (Flow Completion Time) than DCQCN, TIMELY, HPCC and RoCC. Jiao Zhang 0002, Xiaolong Zhong, Mingxuan Yu, Haoyu Pan, Zixuan Guan, Biyao Che, Zirui Wan, Tian Pan 0001, Tao Huang 0005 |
IEEE/ACM Trans. Netw. | 11 |
| 2024 | INT-Label: Lightweight In-Band Network-Wide Telemetry via Distributed LabelingabstractIn-band Network Telemetry (INT) enables hop-by-hop device-internal state exposure for maintaining and troubleshooting data center networks. To achievenetwork-widetelemetry coverage, orchestration on top of the INT primitive is required. A straightforward solution would flood the network with INT probe packets for maximum measurement coverage, which leads to a huge bandwidth overhead. A refined solution leverages the SDN controller to collect the network topology information and carry out centralized probing path planning, which, however, is inefficient in reacting to topology changes. To tackle the above problems, we proposeINT-label, a lightweight In-band Network-Wide Telemetry architecture via the distributed labeling approach. INT-label periodically labels the sampled packets with device-internal states. It is cost-effective with a minor bandwidth overhead and able to seamlessly adapt to topology changes. In order to reduce the number of labeled packets, we introduce a times-based probabilistic labeling algorithm, which allows fewer packets to carry more INT information than the interval-based algorithm. In addition, to counteract the degradation of telemetry resolution due to loss of labeled packets, we design a feedback mechanism which can adaptively change the instant labeling frequency. We provide theoretical proof that INT-label can achieve network-wide telemetry. We analyze the impact of transmission delay on coverage rate and labeling times distribution under the INT-label architecture. Evaluation on software P4 switches suggests that INT-label can achieve 99.72% measurement coverage under the labeling frequency of 20 times per second. With the adaptive labeling enabled, even if 60% of the packets are lost, the coverage can still reach 92%. Enge Song, Tian Pan 0001, Haoyu Song 0001, Qiang Fu 0011, Yingjiang Liu, Chenhao Jia, Chuanying Yuan, Minglan Gao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Trans. Parallel Distributed Syst. | 10 |
| 2024 | Delay-Prioritized and Reliable Task Scheduling With Long-Term Load Balancing in Computing Power NetworksabstractIn the era driven by big data and algorithms, the efficient collaboration of pervasive computing power is crucial for rapidly meeting computing demands and enhancing resource utilization. However, current mainstream end-edge-cloud collaboration faces challenges of computing isolation, adversely affecting resource efficiency and user experience. The Computing Power Network (CPN) is a novel architecture designed to sense and collaborate ubiquitous computing resources through networks. Nevertheless, the expansion of its scope and the integration of networks complicate task scheduling. To address this, we design a collaborative scheduling system that considers the joint selection of computing nodes and network links, aiming to reduce delay, enhance reliability, and ensure long-term load balance. First, we propose a delay-prioritized reliable scheduling policy based on a dual-priority mechanism for forwarding and computing. Second, we define the scheduling problem as a Constrained Markov Decision Process (CMDP) and introduce Lyapunov optimization to transform constraints into instantaneous optimizations, achieving a long-term balanced load of computing and network resources. Lastly, we employ an enhanced Deep Reinforcement Learning (DRL) approach to solve the problem. Performance evaluation demonstrates that compared to standard DRL, the proposed algorithm effectively reduces delay and improves reliability while maintaining long-term load balance, resulting in an overall performance improvement of 54.7%. Renchao Xie, Qinqin Tang, Tao Huang 0005, Zehui Xiong, Tianjiao Chen, Ran Zhang 0004 |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | Adaptive joint placement of edge intelligence services in mobile edge computing
Ru Huo, Chuang Sun 0002, Shuo Wang 0006, Tao Huang 0005 |
Wirel. Networks | 5 |
| 2023 | Amphis: Rearchitecturing Congestion Control for Capturing Internet Application VarietyabstractTCP was designed to provide stream-oriented communication service for bulk data transfer applications (e.g., FTP and Email). With four-decade development, Internet applications have undergone significant changes, which now involve highly dynamic traffic pattern and message-oriented communication paradigm. However, the impact of this substantial evolution on congestion control (CC) has not been fully studied. Most of the network transports today still make the long-held assumption about application traffic, i.e., a byte stream with an unlimited data arrival rate. Tian Pan 0001, Shuihai Hu, Guangyu An, Xincai Fei, Fanzhao Wang, Yueke Chi, Minglan Gao, Hao Wu 0023, Jiao Zhang 0002, Tao Huang 0005, Jingbin Zhou |
APNet | 10 |
| 2023 | uTAS: Ultra-Reliable Time-Aware Shaper for Time-Sensitive NetworksabstractRecent studies leverage time-aware shaper (TAS) and frame replication and elimination for reliability (FRER) techniques to achieve deterministic latency and high reliability for time-triggered (TT) flows. FRER requires transmitting TT flows on$k$disjoint paths to tolerate transient and permanent failures. However, directly allocating timeslot resources for all$k$TT flows will dramatically increase the computational complexity of gate control lists (GCLs), seriously impair scheduling capabilities, and result in a wastage of bandwidth. In this paper, we propose an ultra-reliable time-aware shaper (uTAS). uTAS only allocates timeslots for one TT flow to ensure deterministic transmission. The$k - 1$replica TT (RT) flows are delivered using a best-effort strategy. On this basis, we propose an adaptive window scheduling (AWS) algorithm based on network calculus, which aims to guarantee that RT flows reach their destinations within the deadline. Evaluation results show that uTAS can meet the flow reliability and deadline requirements. Compared to directly combining TAS and FRER (DCTF), uTAS reduces the total number of GCLs by approximately 72.9%. Guizhen Li, Shuo Wang 0006, Yudong Huang, Xingyu Zhong, Guiyu Zhang, Luying Bai, Tao Huang 0005 |
GLOBECOM | 7 |
| 2023 | Fast In-Network Functionality Embedding in Software-Defined Service-Centric NetworkingabstractJoint resource allocation in integrated networking, computing, and caching frameworks has attracted plenty of attention. In software-defined service-centric networking (SDSCN), we investigate an energy cost economical functionality embedding (FE) problem. The FE problem is a non-convex quadratically constrained quadratic programming (QCQP) problem which is difficult to obtain its global optimal solutions. In this paper, we propose a low-complexity high-performance algorithm for energy-economical FE design in large-scale SDSCN systems by leveraging the alternating direction method of multipliers (ADMM) together with successive convex approximation (SCA). In specific, the FE problem is first approximated as a sequence of convex subproblems via SCA. Each convex subproblem is then reformulated as a novel ADMM form to enable parallel computations and closed-form solutions. Numerical results show that our fast algorithm reduces the complexity by orders of magnitude and obtains favorable performances compared with state-of-the-art algorithms. Renchao Xie, Tao Huang 0005, Yunjie Liu 0001 |
GLOBECOM | 3 |
| 2023 | A Contract-Based Incentive Mechanism for Resources Trading in Computing Force NetworksabstractRecently, Computing Force Networks (CFN) is emerging to deeply integrate and flexibly schedule multi-layer, multi-domain, distributed, and heterogeneous computing force resources among the cloud, edge network, and end devices. In CFN, the market monopoly of large resource providers results in a lack of bargaining power for small-sized providers. Existing resources pricing strategies ignore dynamic market factors affecting prices. Moreover, there is information asymmetry between resource consumers and providers. These problems destroy the fairness of the trading market and damage the benefits of consumers and providers, making them have a low degree of willingness to participate in resources trading. Therefore, this paper proposes a contract theory-based incentive mechanism to solve the above problems and motivate resource consumers and providers to join CFN. The proposed scheme classifies resource commodities into different types and enables the trading platform to provide appropriate rewards based on commodities' types and contributions. More specifically, we formulate an optimization problem aiming at maximizing the trading platform's utility and obtain an optimal contract scheme based on the individual rationality and incentive compatible constraints. Simulation results verify the feasibility and effectiveness of our scheme. Wen Wen 0011, Lu Lu 0016, Yuexia Fu, Qinqin Tang, Renchao Xie, Tao Huang 0005 |
GLOBECOM | 7 |
| 2023 | Joint Task Scheduling and Intelligence Optimization in CPN-Enabled Connected Intelligence SystemsabstractAs Artificial Intelligence (AI) has flourished in various industries in recent years, the evolutionary trend of endogenous network intelligence continues to accelerate. Connected intelligence, which aims to achieve a widely distributed and collaborative evolution of intelligence, has received much attention. Meanwhile, the emerging Computing Power Network (CPN) provides more robust computation and communication capabilities for intelligence training and intelligent application processing. In this context, the integration of CPN and connected intelligence becomes a potential solution to drive the digital and intelligent transformation of networks. In this paper, we propose a scheme to jointly consider task scheduling, routing, and intelligence capability improvement during the processing of smart applications represented by Digital Twin (DT) in the CPN-enabled connected intelligence systems. We formulate the problem of jointly optimizing the processing time consumption and training accuracy improvement. We solve the problem using a modified NSGA-II algorithm and numerical results show that our approach is effective in optimizing the overall average time consumption and improving the accuracy of the intelligence models distributed in the system during task processing. Gaochang Xie, Renchao Xie, Qinqin Tang, Zongping Li, Tao Huang 0005 |
GLOBECOM | 7 |
| 2023 | A Novel Scheduling Scheme for Earth Observation in LEO Satellite SystemsabstractEarth observation applications, such as emergency surveillance and disaster relief, are thriving due to the availability of earth observation satellites that provide timely and objective observation data at different spatial and temporal scales. Such observation data is processed at the satellite edge with orbital edge computing, leading to potentially reduced bandwidth cost and transmission delay. However, most existing studies primarily focus on optimizing computation offloading but ignore the consideration of object observation and observation data transmission. To fill this gap, this paper proposes a novel scheduling approach that jointly considers observation satellites, relay satellites, and computing satellites in LEO satellite systems, aiming to maximize the number of completed observation tasks while taking into account various requirements of observation, transmission, and computation resources. Specifically, we first formulate the problem of jointly scheduling observation, relay, and computing satellites to maximize the number of accomplished observation tasks. Then, we decompose the formulated problem into two sub-problems and design a resource-aware algorithm called ORCA to determine the optimal scheduling of observation, relay, and computing satellites. Simulation results demonstrate that ORCA outperforms existing algorithms in terms of completing a number of observation tasks. Ran Zhang 0004, Changqing Luo, Jiang Liu 0010, Geyong Min, Tao Huang 0005 |
GLOBECOM | 6 |
| 2023 | Performance Modeling and Analysis of Distributed Deep Neural Network Training with Parameter ServerabstractWith the growth of dataset size and the development of hardware accelerators, the application of deep neural networks (DNN) in various fields has made great breakthroughs. In order to improve the training speed of DNN, distributed training has been widely used. However, the imbalance between computation and communication makes distributed training difficult to achieve maximum efficiency. Therefore there is a need to detect the bottleneck state and verify the effect of some optimization schemes. Testing on a physical cluster incurs additional time and cost overhead. This paper builds a DNN-specific performance model that is used for bottleneck detection and tuning at a low cost. We build this model through detailed analysis and reasonable assumptions. We also focus on fine-grained modeling of scalability and network components, which are key factors affecting performance. Then we verify the performance model with an average error of 5% on testbed and emulator. Finally, we provide use cases of the performance model. Jiao Zhang 0002, Dehui Wei, Tian Pan 0001, Tao Huang 0005 |
GLOBECOM | 5 |
| 2023 | INT-Balance: In-Band Network-Wide Telemetry with Balanced Monitoring Path PlanningabstractIn-band Network Telemetry (INT) empowers high-resolution network monitoring by collecting hop-by-hop device-internal states through the data plane without frequently disturbing the control plane. To achieve network-wide monitoring, a high-level orchestration is made to provision multiple monitoring paths to cover the entire network. The path number and path overlapping are kept minimum to maximally reduce the telemetry overhead. However, in production deployment, except for the telemetry overhead, the telemetry timeliness is equally important for fine-grained monitoring, which creates new requirements of balanced monitoring path planning. Given the INT probes from multiple paths are collected to the central controller for analysis, the late arrival of even one probe will delay the analysis process and affect the monitoring timeliness. To address the problem, we propose INT-balance, a novel path planning algorithm for balanced INT path generation. In INT-balance, we first break the original network graph into multiple path segments at the odd vertices. Then, we iteratively splice the two shortest path segments with the joint endpoints to form a longer path segment until the path segment number reaches half the number of the odd vertices. INT-balance generates the minimum number of INT paths with well-balanced path lengths, covering every edge of the network graph without any path overlapping. Evaluation on a network of 100 switches shows that the path length variance of INT-balance is 67% less than that of INT-path, while the algorithm execution time is increased only by 0.012s. Yan Zhang 0063, Tian Pan 0001, Enge Song, Jiang Liu 0010, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 6 |
| 2023 | P4RSS: Load-Aware Intra-Server Load Balancing with Programmable Switching ASICsabstractOff-the-shelf x86 servers are widely deployed as middleboxes in edge and public clouds, such as cloud gateways and load balancers. They follow the “run-to-completion” model and achieve parallel traffic processing by distributing packet flows across multiple CPU cores using the RSS (receive side scaling) capability of NICs. However, RSS can cause inter-core load imbalance as it conducts stateless hashing without considering the CPU core utilization. As a result, multiple heavy-hitter flows can potentially overload a single CPU core when they are hashed onto that core. In this research, we propose P4RSS, a load-aware intra-server load balancing solution that leverages the P4 data plane. Specifically, a P4 ASIC is placed in front of the CPU to perform stateful traffic load balancing among multiple CPU cores based on real-time monitoring of core utilization. In addition, flow affinity maintenance and heavy hitter throttling are also offloaded to the P4 ASIC to free up valuable CPU computing resources. P4RSS can be implemented in the form of either hyper-converged server switches or P4-based SmartNICs. Evaluation results demonstrate that P4RSS reduces the standard deviation of CPU core utilization by 22%~53% compared to RSS. This not only improves the stability of middleboxes but also allows for higher CPU utilization without overprovisioning. Yan Zou, Tian Pan 0001, Lu Lu 0016, Kehan Yao, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 6 |
| 2023 | Hostping: Diagnosing Intra-host Network Bottlenecks in RDMA Servers
Kefei Liu 0004, Zhuo Jiang, Jiao Zhang 0002, Xiaolong Zhong, Lizhuang Tan, Tian Pan 0001, Tao Huang 0005 |
NSDI | 8 |
| 2023 | Poster: Programmable Cycle-Specified Queue for Deterministic NetworkingabstractThe emerging time-critical applications pose intense demands for enabling large-scale deterministic networks. In this paper, we propose a new Programmable Cycle-Specified Queue (PCSQ) for wide-area deterministic packet scheduling. We implement the first end-to-end high-precision rotation dequeuing, which enables microsecond-level time slot resource reservation (noted as T) and especially jitter control of up to 2T. We prototype the PCSQ scheduler on an FPGA. The PCSQ-enabled switches can guarantee bounded delay and jitter transmission on a realistic testbed. Yudong Huang, Shuo Wang 0006, Shiyin Zhu, Guoyu Peng, Xinyuan Zhang 0011, Tian Pan 0001, Tao Huang 0005, Zuopin Cheng, Daorong Guo, Lianqing Zhang, Juyan Lei, Liangzhang Xu, Wei Wang 0494, Xinmin Liu, Xuejun You, Yunjie Liu 0001 |
SIGCOMM | 7 |
| 2023 | Maximizing Optical Inter-DC Emergency Backup Reliability in Unpredictable DisastersabstractThe emergency backup problem of optical inter-datacenter (inter-DC) networks is widely studied to avoid data loss caused by disasters. However, when unpredictable disasters such as earthquakes and large-scale power outages attack the optical inter-DC network, it is impossible to predict the network damage under a discrete time domain through the early warning system (EWS). Designing a reliable emergency backup scheme is becoming a key challenge in the optical inter-DC network in unpredictable disasters. Previous work only optimized emergency backup in the optical inter-DC networks in predictable disasters, which would lead to the unreliability of emergency backup of the optical inter-DC networks in unpredictable disasters. Therefore, this work proposes a reliable emergency backup method for the optical inter-DC networks in unpredictable disasters. Firstly, the disaster probability propagation model of the optical inter-DC network in unpredictable disasters is established based on the Markov process to quantify the network damage under a discrete time domain and install the emergency backup reliability problem. Next, constructing a time-sensitive variable time extension network (TS-VTEN) transforms the dynamic emergency backup reliability problem into a static flow problem. Finally, the problem is expressed as integer linear programming (ILP). The proposed ILP method outperforms state-of-the-art emergency backup methods by simultaneously achieving high reliability and time efficiency. Mingwei Cui, Weihong Wu, Tao Huang 0005 |
VTC2023-Spring | 5 |
| 2023 | FIBFT: An Improved Byzantine Consensus Mechanism for Edge ComputingabstractBlockchain has been widely used to solve data privacy and security issues in edge computing scenarios. However, the blockchain based on edge computing still has some performance problems, such as insufficient scalability, difficulty in balancing security and edge device power consumption, and inability to simultaneously meet low latency, high throughput, high security and privacy issues, etc. In order to solve these problems, this paper proposes a generally improved Byzantine consensus mechanism based on the K-medoids clustering algorithm - FIBFT. Considering the different performance characteristics of each node in the network, the node’s state is first abstracted into a multi-dimensional state space containing eigenvalues, and then the nodes are divided into subnets by the efficient K-medoids clustering algorithm. Each subnet uses a Byzantine consensus mechanism based on arbitration for consensus and data interaction, and the consensus data could be exchanged between the subnets without interfering with the consensus process. The research results show that FIBFT has better scalability and throughput while ensuring high security compared with the traditional Byzantine consensus algorithm. Ningjie Gao, Ru Huo, Shuo Wang 0006, Tao Huang 0005 |
WCNC | 4 |
| 2023 | ED-VNE: A profit-oriented VNE optimization scheme of energy and delay in 5G SlaaS
Ying Wang 0141, Jiang Liu 0010, Mingwei Cui, Weihong Wu, Tao Huang 0005 |
Comput. Networks | 5 |
| 2023 | A Task-Oriented Hybrid Routing Approach based on Deep Deterministic Policy Gradient
Zongxuan Sha, Ru Huo, Chuang Sun 0002, Shuo Wang 0006, Tao Huang 0005 |
Comput. Commun. | 5 |
| 2023 | SCRT: A Secure and Efficient State-Channel-Based Resource Trading Scheme for Internet of ThingsabstractWith the development of edge computing technology, the resource-limited Internet of Things (IoT) devices can offload computation-intensive artificial intelligence tasks, such as model training and inference to edge servers through resource trading. However, due to the increase in the number of intelligent applications and the rise of peer-to-peer (P2P) resource trading, the existing resource trading schemes based on the blockchain can no longer meet the needs of efficiency and security at the same time. In this article, a new state channels-based resource trading scheme is proposed for IoT, which can improve scalability without sacrificing security and fairness. In our scheme, most of the trading process could be completed off-chain, and the blockchain is used as an adjudicator to determine malicious behavior according to the users’ actions. Moreover, a method without being reliant on support from third parties is presented to defend against execution forks that must be faced when using the state channels. Finally, the feasibility and efficiency of our scheme are experimentally verified in the realistic testbed. Wei Chen 0131, Ru Huo, Chuang Sun 0002, Shiqin Zeng, Shuo Wang 0006, Tao Huang 0005 |
IEEE Internet Things J. | 6 |
| 2023 | Workflow Scheduling in Serverless Edge Computing for the Industrial Internet of Things: A Learning ApproachabstractServerless edge computing is seen as a promising enabler to execute differentiated Industrial Internet of Things (IIoT) applications without managing the underlying servers and clusters. In IIoT serverless edge computing, IIoT workflow scheduling for cloud-edge collaborative processing is closely related to the service quality of users. However, serverless functions decomposed by IIoT applications are limited in their deployment at the edge due to the resource-constrained nature of edge infrastructures. In addition, the scheduling of complex IIoT applications supported by serverless computing is more challenging. Therefore, considering the limited function deployment and the complex dependencies of serverless workflows, we model the workflow application as directed acyclic graph and formulate the scheduling problem as a multiobjective optimization problem. A dueling double deep Q-network-based solution is proposed to make scheduling decisions under dynamically changing systems. Extensive simulation experiments are conducted to validate the superiority of the proposed scheme. Renchao Xie, Dier Gu, Qinqin Tang, Tao Huang 0005, F. Richard Yu |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Collective Deep Reinforcement Learning for Intelligence Sharing in the Internet of Intelligence-Empowered Edge ComputingabstractEdge intelligence is emerging as a new interdiscipline to push learning intelligence from remote centers to the edge of the network. However, with its widespread deployment, new challenges arise in terms of training efficiency and service of quality (QoS). Massive repetitive model training is ubiquitous due to the inevitable needs of users for the same types of data and training results. Additionally, a smaller volume of data samples will cause the over-fitting of models. To address these issues, driven by the Internet of intelligence, this paper proposes a distributed edge intelligence sharing scheme, which allows distributed edge nodes to quickly and economically improve learning performance by sharing their learned intelligence. Considering the time-varying edge network states including data collection states, computing and communication states, and node reputation states, the distributed intelligence sharing is formulated as a multi-agent Markov decision process (MDP). Then, a novel collective deep reinforcement learning (CDRL) algorithm is designed to obtain the optimal intelligence sharing policy, which consists of local soft actor-critic (SAC) learning at each edge node and collective learning between different edge nodes. Simulation results indicate our proposal outperforms the benchmark schemes in terms of learning efficiency and intelligence sharing efficiency. Qinqin Tang, Renchao Xie, F. Richard Yu, Tianjiao Chen, Ran Zhang 0004, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | DRL-Based Service Function Chain Edge-to-Edge and Edge-to-Cloud Joint Offloading in Edge-Cloud NetworkabstractIn this paper, we study service function chain (SFC) offloading in the edge-cloud network. Two offloading options are available for fully-loaded edge nodes: edge to edge (E2E) offloading and edge to cloud (E2C) offloading. Both E2E offloading and E2C offloading have been optimized in existing research, and Deep Reinforcement Learning (DRL) methods were adopted to achieve excellent performances. However, DRL-based SFC E2E and E2C joint offloading is still a research gap. In this paper, we propose a DRL-based SFC E2E and E2C joint offloading optimization algorithm for the edge-cloud network to maximize the utilization efficiency of edge resources. Twin Delayed Deep Deterministic policy gradient (TD3) algorithm is applied to the optimization problem. The simulation results indicate that the proposed algorithm has excellent convergence performance and improves the utilization efficiency of edge resources in edge-cloud network scenarios of diverse scales. Wentao Fan 0002, Fan Yang 0046, Peilong Wang, Mao Miao, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2023 | Flexible Cyclic Queuing and Forwarding for Time-Sensitive Software-Defined NetworksabstractTime-Sensitive Networking (TSN) is emerging to support critical real-time applications in Industry 4.0. Recent proposals leverage Cyclic Queuing and Forwarding (CQF) to achieve bounded-delay transmission for cyclic flows in TSN. However, the CQF is not flexible enough in two aspects. First, it cannot achieve zero jitter. The Ping-Pong queue-based model in CQF will introduce the jitter of two cycles, which is inapplicable to industrial automation scenarios where isochronous flows require zero jitter. Second, it may require setting the maximum queue length to a fixed value in advance and scheduling the flows offline, which is challenging for dynamic traffic scheduling. In this paper, we firstly present a time-aware cyclic-queuing (TACQ) mechanism to enable zero jitter for CQF. TACQ consists of a novel no-wait shaper (NWS) and a cyclic-queuing shaper (CQS). The NWS handles isochronous flows by strictly limiting the transmission time of flows that do not overlap on each output port and each period. The CQS is extended from CQF to schedule cyclic flows. Then, we propose a variable time slot mechanism and a novel incremental routing and scheduling (IRAS) algorithm based on software-defined networking (SDN) to online schedule dynamic flows. Simulation results show that TACQ significantly reduces the delay of isochronous flows and achieves zero jitter compared with CQF. And the IRAS algorithm approaches 96.1% of the optimal solution in scheduling 2000 flows with a feasible per-flow computational time. Yudong Huang, Shuo Wang 0006, Xinyuan Zhang 0011, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2023 | SLIT: Achieving Fast Bandwidth Isolation Across Virtual MachinesabstractNetwork performance guarantee in the cloud is one of the hottest research topics recently. However, current approaches either lack scalability or fail to achieve high performance. To avoid the hard tradeoff between the scalability and performance, we adopt a hybrid approach to get the best of both worlds. In this paper, we follow the stateless design principle and propose SLIT, which provides fast VM-level network performance isolation and maintains core-stateless characteristics. Specifically, at the end host, we develop a novel network-adaptive labeling mechanism and it computes the correct scheduling priority for each hop. At the switch, packets are scheduled based on their labels and we develop a novel straggler detection algorithm to find the new incoming flow. Our evaluation results show that SLIT can approximate the optimal performance of Weighted Fair Queuing (WFQ) effectively. Chengyuan Huang, Jiao Zhang 0002, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2023 | RCC: Enabling Receiver-Driven RDMA Congestion Control With Congestion Divide-and-Conquer in Datacenter NetworksabstractThe development of datacenter applications leads to the need for end-to-end communication with microsecond latency. As a result, RDMA is becoming prevalent in datacenter networks to mitigate the latency caused by the slow processing speed of the traditional software network stack. However, existing RDMA congestion control mechanisms are either far from optimal in simultaneously achieving high throughput and low latency or in need of additional in-network function support. In this paper, by leveraging the observation that most congestion occurs at the last hop in datacenter networks, we propose RCC, a receiver-driven rapid congestion control mechanism for RDMA networks that combines explicit assignment and iterative window adjustment. Firstly, we propose a network congestion distinguish method to classify congestions into two types, last-hop congestion and in-network congestion. Then, an Explicit Window Assignment mechanism is proposed to solve the last-hop congestion, which enables senders to converge to a proper sending rate in one-RTT. For in-network congestion, a PID-based iterative delay-based window adjustment scheme is proposed to achieve fast convergence and near-zero queuing latency. RCC does not need additional in-network support and is friendly to hardware implementation. In our evaluation, the overall average FCT (Flow Completion Time) of RCC is$4{\sim }79\%$better than Homa, ExpressPass, DCQCN, TIMELY, and HPCC. Jiao Zhang 0002, Xiaolong Zhong, Zirui Wan, Tian Pan 0001, Tao Huang 0005 |
IEEE/ACM Trans. Netw. | 6 |
| 2022 | Power-Aware Traffic Engineering for Data Center Networks via Deep Reinforcement LearningabstractThe issue of high energy consumption and low energy utilization in data center networks (DCNs) has always been the focus of attention of both academia and industry. One general solution is to select a subset of network devices that can meet the traffic transmission requirements, thereby turning off the remaining redundant devices. However, modeling the problem as integer linear programming introduces significant time overhead, while heuristic approaches often suffer from poor generalizability. In this paper, we propose GreenDCN.ai, a closed-loop control system, which utilizes In-band Network Telemetry to collect the network-wide device-internal state, and leverages a Deep Reinforcement Learning-based energy-saving algorithm to make rapid decisions to turn on or off network device ports in response to the real-time network state. The trained GreenDCN.ai can adaptively adjust its energy-saving strategy without human intervention when the DCN topology changes. Besides, based on the regularity of the DCN topology, we design two training complexity reduction methods to address the non-convergence issue under large-scale DCN topologies. Specifically, we split the large-scale DCN topology into sub-topologies for parallel training on each sub-topology without breaking the DCN topology connectivity. Evaluation on software P4 switches suggests that GreenDCN.ai can achieve stable convergence within 590 episodes, generate effective action decisions within$\boldsymbol{79}\upmu\mathrm{s}$, and save about 34% to 39% of the network energy consumption. Minglan Gao, Tian Pan 0001, Enge Song, Mengqi Yang, Tao Huang 0005, Yunjie Liu 0001 |
GLOBECOM | 5 |
| 2022 | High-Performance and Low-Cost VPP Gateway for Virtual Cloud NetworksabstractThe virtual cloud network has been the first choice for most enterprises to expand local networks due to its convenience, flexibility, and elasticity. Cloud gateways are the key to steering traffic among these virtual networks, which require high throughput and low delay, usually in Tbps and microseconds (us), to improve the performance of cloud regions. However, as cloud traffic growth far exceeds Moore's law, previous software cloud gateways are facing the performance bottleneck that they are vulnerable to the attack of heavy-hitter flows. In this paper, we propose a high-performance and low-cost software cloud gateway for accelerating virtual cloud networks. By revisiting the Vector Packet Processing (VPP) framework, we design a custom control plane to enable various network functions of the cloud gateway, which enhances the routing and forwarding actions in the data plane. And we also provide a common interface for users to flexibly configure and manage the gateway. This pure software architecture has a low deployment cost. Compared with other software gateways with roughly the same price, the processing performance is close to the NIC's line speed, which is suitable for high concurrency cloud scenarios. Shuo Wang 0006, Yudong Huang, Tao Huang 0005, Yunjie Liu 0001 |
GLOBECOM | 4 |
| 2022 | LISP-LEO: Location/Identity Separation-based Mobility Management for LEO Satellite NetworksabstractIn space-terrestrial integrated networks, the relative motion between LEO satellites and ground terminals is inevitable, which will trigger the reassignment of the terminal IP addresses and disrupt the ongoing TCP connections. Traditional Mobile IP protocol can solve the problem by using the home agent and the tunneling mechanism. However, for space-terrestrial integrated networks, Mobile IP is inefficient as it introduces (1) increased latency when registering with the remote home agent, (2) high packet loss due to large registration latency, (3) triangular routing to the remote home agent. To address the above issues, we propose LISP-LEO, a location/identity separation-based mobility management protocol for LEO satellite networks. Specifically, (1) we divide the Earth's surface into partitions and maintain a partition-satellite mapping table in real-time according to the regularity of satellite motion, (2) we always route traffic to the satellite above the destined terminal by querying the partition-satellite mapping table, which eliminates triangular routing and the related performance overheads, (3) we handle the corner case that multiple satellites occur above the destined terminal by proposing last-hop relay. The evaluation convinces that, for the LEO-48 constellation, LISP-LEO produces a 55.0% reduction in the RTT and a 45.8% reduction in the number of forwarding hops in the worst routing case compared with Mobile IP. Tian Pan 0001, Xuebei Zhang, Tao Huang 0005, Yunjie Liu 0001 |
GLOBECOM | 5 |
| 2022 | Joint Resource Allocation for Software-Defined Serverless Service-Centric NetworkingabstractRecently, there are significant advances in networking, computing, and caching (NCC). Nevertheless, few attempts have been made to explore the potential of the promising server-less computing paradigm in NCC integrated frameworks. In this paper, we consider a software-defined serverless service-centric networking (SD-SSCN) framework that not only dynamically orchestrates NCC by combining software-defined networking and service-centric networking technologies but also strongly focuses on the context of serverless computing. To achieve a cost-efficient green SD-SSCN system, we first formulate the joint resource allocation problem to minimize an overall average cost model. We derive this problem as a nonlinear integer programming problem, then we relax it as a quadratic programming problem and develop a primal-dual interior-point algorithm to find joint resource allocation solutions. Simulation results show that our proposed SD-SSCN framework significantly outperforms the traditional networks in terms of the average cost in the context of serverless computing. Renchao Xie, Tao Huang 0005, Yunjie Liu 0001 |
GLOBECOM | 3 |
| 2022 | Cost-Effective and Deployment-Friendly L4 Load Balancers Based on Programmable SwitchesabstractRecent proposals leverage emerging programmable switches to implement high-throughput and low-latency load balancers in the datacenter. However, most of them store per-connection states in programmable switches to ensure consistent load balancing decisions, which is costly due to the limited on-chip memory. Other proposals avoid storing per-connection states but have difficulty in large-scale deployment due to the modification of running applications. We present CDLB, a cost-effective and deployment-friendly Layer-4 load balancer based on programmable switches in the datacenter. To be cost-effective, CDLB leverages an improved hash-based algorithm that can maintain per-connection consistency in a static environment. To be easier in deployment, we introduce a small state table and design a protocol between load balancers and the controller to maintain per-connection consistency in a dynamic environment without modifying running applications. The state table stores small amounts of connection states temporarily to ensure consistent load balancing decisions under device variations. We implemented CDLB with bmv2 in the mininet environment. The evaluation results show that CDLB greatly reduces overhead and achieves better performance compared with stateful load balancers which store per-connection states in programmable switches, and CDLB maintains per-connection consistency well in a dynamic environment. Shuo Wang 0006, Tao Huang 0005 |
GLOBECOM | 3 |
| 2022 | Lightweight Route Flooding via Flooding Topology Pruning for LEO Satellite NetworksabstractWith the low latency and high coverage, the low earth orbit (LEO) satellite systems are attracting more and more venture capitals as well as research attentions. Due to their highly dynamic constellation topologies, routing protocols on the ground have to be tailored to efficiently adapt to the regular topology changes. However, for irregular topology changes caused by exceptional link failure/recovery, network-wide route flooding is still necessary for route convergence. But, this will cause significant traffic flooding redundancy due to the high density of constellation topologies. For larger-scale constellations, the redundancy issue will be exacerbated. To lessen the redundancy, this work proposes a lightweight route flooding mechanism by generating a sparse flooding topology that prunes the original full-mesh topology, and only flooding the route information on the sparse topology. By considering the maximum flooding hop as well as the robustness of the flooding topology, we design an algorithm to calculate the optimal topology instead of just applying the minimum spanning tree. The evaluation shows that, for the LEO-96 constellation, the new flooding topology has a 28.1% reduction in the inter-satellite links (ISLs) compared with the original topology, and the new flooding mechanism has a 37.52% reduction in the traffic flooded and a 10.03% reduction in the route convergence time compared with OSPF. Such improvements will be amplified on larger-scale constellations. Guohao Ruan, Tian Pan 0001, Chengcheng Lu, Zhengjie Luo, Houtian Wang, Jiao Zhang 0002, Yushi Shen, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 8 |
| 2022 | Large-scale Deterministic Transmission among IEEE 802.1Qbv Time-Sensitive NetworksabstractIEEE 802.1Qbv (TAS) is the most widely used technique in Time-Sensitive Networking (TSN) which aims to provide bounded transmission delays and ultra-low jitters in industrial local area networks. With the development of emerging technologies (e.g., cloud computing), many wide-range time-sensitive network services emerge, such as factory automation, connected vehicles, and smart grids. Nevertheless, TAS is a Layer 2 technique for local networks, and cannot provide large-scale deterministic transmission. To tackle this problem, this paper proposes a hierarchical network containing access networks and a core network. Access networks perform TAS to aggregate time-sensitive traffic. In the core network, we exploit DIP (a well-known deterministic networking mechanism for backbone networks) to achieve long-distance deterministic transmission. Due to the differences between TAS and DIP, we design cross-domain transmission mechanisms at the edge of access networks and the core network to achieve seamless deterministic transmission. We also formulate the end-to-end scheduling to maximize the amount of accepted time-sensitive traffic. Experimental simulations show that the proposed network can achieve end-to-end deterministic transmission even in high-loaded scenarios. Weiqian Tan, Binwei Wu, Shuo Wang 0006, Tao Huang 0005 |
ICC | 4 |
| 2022 | Delay-Aware Cooperative Caching for On-Chain Authentication in LEO Satellite Communication SystemsabstractUser authentication on the blockchain has been considered a promising solution to secure communications in LEO satellite communication systems. Due to resource-limited LEO satellites, the blockchain needs to be deployed in the terrestrial network component of LEO satellite communication systems, consequently resulting in high authentication delays. To fill the gap, we propose to cache the blockchain at LEO satellites and update the blockchain periodically and design a delay-aware cooperative caching scheme for on-chain authentication by considering the query delay and the synchronization delay. Specifically, we first propose to divide LEO satellites into multiple clusters which have the same copy of all the blocks belonging to the blockchain. Then, we model the clustering problem as a coalition formation game. Afterward, we design a distributed delay-aware coalition formation algorithm, which is called DAC, to find an optimal coalition partition. Extensive simulation results show the efficacy of the proposed scheme. Jiang Liu 0010, Ran Zhang 0004, Xinyuan Zhang 0011, Changqing Luo, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 6 |
| 2022 | MIMIC: SmartNIC-aided Flow Backpressure for CPU Overloading Protection in Multi-Tenant CloudsabstractIn multi-tenant clouds, off-the-shelf x86 boxes are widely deployed as middleboxes. With the rapid growth of cloud traffic and the migration to NFV deployment in recent years, CPU overloading at middleboxes becomes more of an issue. From our data centers, we observed that the CPU overloading was caused by heavy hitters. To address this issue, we propose MIMIC, a cloud-scale flow backpressure system, implemented onto our existing SmartNIC with FPGA acceleration. MIMIC rate-limits the selected heavy hitters through a new per-flow backpressure protocol and a new heavy-hitter detection system, to protect the other tenants. The detection system is based on hierarchical memory design, leveraging on-chip SRAM and off-chip DRAM, which can handle highly concurrent cloud traffic without the losses of flow information. We extend the design by adding a pre-filtering procedure for rapid detection. To avoid CPU being flooded by FPGA through frequent heavy-hitter reporting, due to their performance disparity, the CPU queries the FPGA on demand. The backpressure protocol is non-invasive to protect tenant privacy and allows controllable rate-limiting through the novel use of ECN and meter tables. The SmartNIC acts as a man in the middle to facilitate heavy-hitter detection and per-flow backpressuring. In a production setting, we observe that MIMIC can react quickly and bring down CPU load to the normal level within 10ms without packet losses. Enge Song, Nianbing Yu, Tian Pan 0001, Qiang Fu 0011, Xionglie Wei, Yisong Qiao, Jianyuan Lu, Yijian Dong, Mingxu Xie, Jinkui Mao, Zhengjie Luo, Chenhao Jia, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Shunmin Zhu |
ICNP | 16 |
| 2022 | Behavior-decoupled Labeling Mechanism in Generalized SRv6abstractIn this paper, we study the problem of compression efficiency of SRH (Segment Routing Header) in G-SRv6 (Generalized SRv6). We propose the Function-decoupled Segment Routing Mechanism (FDSRM) to optimize the SID list in SRH while ensuring that the routing policy decision would not be affected. FDSRM logically decouples the functions of SID/G-SID according to the function in routing decision and instruction indication. Based on FDSRM, we propose a mathematical optimization framework that leverages the LSTM neural network to optimize the SID allocation to adapt to future traffic. Simulation results indicate that FDSRM can improve the probability of SRH compression by 50.86%, and compress more than 24.50 bytes when the hop count is greater than 20. Weihong Wu, Anbang Pei, Tao Huang 0005 |
ICNP | 4 |
| 2022 | WebQMon.ai: Gateway-Based Web QoE Assessment Using Lightweight Neural Networks
Enge Song, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
ICSOC | 6 |
| 2022 | Workflow Scheduling Using Hybrid PSO-GA Algorithm in Serverless Edge Computing for the Internet of ThingsabstractIn this paper, we design a task scheduling scheme for Internet of Things (IoT) workflow applications in serverless edge computing. Notice the fact that complex applications in traditional serverless computing are decomposed into several stateless, dependent functions, whose execution environments are pre-deployed at the resource-finite edge domain, we model the workflow application as Directed Acyclic Graph (DAG) by considering the distribution of edge resources and the deployment of serverless functions. We further formulate the scheduling problem as a multi-objective optimization problem to reduce the time consumption, energy consumption, and cost simultaneously. Then, considering the diversity of solution space and the fast convergence to optimal solutions, an improved hybrid algorithm that combines Particle Swarm Optimization and Genetic Algorithm (PSO–GA) is introduced and utilized to make the scheduling decision. Finally, extensive simulation experiments are conducted to validate the superiority of the proposed scheme. Renchao Xie, Dier Gu, Qinqin Tang, Tao Huang 0005, F. Richard Yu |
VTC Spring | 4 |
| 2022 | A blockchain-based and privacy-preserved authentication scheme for inter-constellation collaboration in Space-Ground Integrated Networks
Ran Zhang 0004, Jiang Liu 0010, Tao Huang 0005, Yunjie Liu 0001, F. Richard Yu |
Comput. Networks | 4 |
| 2022 | Learning-Based Computation Offloading for IoRT Through Ka/Q-Band Satellite-Terrestrial Integrated NetworksabstractIn this article, we propose a multilayer Ka/Q-band satellite–terrestrial integrated network for the Internet of Remote Things (IoRT) to achieve a high transmission rate with communication robustness in dynamic network environments. Under this architecture, we investigate how to jointly manage the offloading path selection and resource allocation to offload computation-intensive and delay-sensitive tasks in the IoRT. Considering continuous low earth orbit (LEO) satellite movements and Markovian rainfall changes, the computation offloading problem is described as a Markov decision process (MDP) formulation with the objective of maximizing the number of offloaded tasks with satisfied delay requirements and minimizing the power consumption of the LEO satellites. A deep reinforcement learning (DRL) approach is leveraged to make optimal decisions by taking account of dynamic queues of IoRT devices, channel conditions that vary with rainfall intensities and satellite positions, and computing capabilities of ground stations. Extensive simulations are conducted to validate the effectiveness and superiority of our proposed scheme. Tianjiao Chen, Jiang Liu 0010, Qiang Ye 0002, Weihua Zhuang, Weiting Zhang, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Internet Things J. | 6 |
| 2022 | Sharding-Hashgraph: A High-Performance Blockchain-Based Framework for Industrial Internet of Things With Hashgraph MechanismabstractIn recent years, with the development and widespread use of blockchain, many projects have introduced blockchain technology to solve the increasingly serious security problems of the Industrial Internet of Things (IIoT). However, due to the conflict between the operational performance and security of the blockchain system, the conflict between transparency and privacy, and the compatibility issues with a large number of IIoT devices running together, the mainstream blockchain system cannot be applied to IIoT scenarios. In order to solve these problems, in this article, we propose an IIoT distributed data system based on blockchain technology. We provide a novel system architecture for different IIoT devices to deploy high-performance blockchain systems in many scenarios, such as smart factory networks. To improve the performance of the blockchain network, we adopt the sharding hashgraph consensus mechanism and introduce a node evaluation mechanism based on the state of the node, which is applied to divide a large number of nodes into many shards dynamically. We abstract the node sharding problem as a joint optimization problem and use deep reinforcement learning to solve it. Finally, we compared with asynchronous Byzantine consensus algorithms, such as HoneybadgerBFT and BEAT, which validated the performance of this system architecture. Ningjie Gao, Ru Huo, Shuo Wang 0006, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Internet Things J. | 4 |
| 2022 | Reliable and Low-Overhead Clustering in LEO Small Satellite NetworksabstractLow earth orbit (LEO) small satellites have attracted great interests in civilian and military applications due to their low cost and high service performance. However, the enormous scale and high dynamism of small satellites pose challenges to network flexibility and scalability. Therefore, the hierarchical satellite network structure is introduced as an effective approach to enhance the satellite network capabilities further. In this regard, small satellites’ clustering is of fundamental importance for designing such a hierarchical structure. Satellite clusters are always prone to instability due to unpredictable link failures and frequent topology changes. In this article, we study the small satellite clustering problem of jointly optimizing the cluster reliability and the network management overhead. A coalition game-theoretic framework is introduced to obtain low computational complexity by adopting the clustering-decision-making process in an automated and fully distributed fashion. A distributed coalition formation algorithm based on the optimization of reliability and management overhead is developed for the clustering problem. Finally, extensive simulations have been conducted, and the results show that our proposed clustering scheme is able to produce better results than the baseline schemes. Jiang Liu 0010, Xinyuan Zhang 0011, Ran Zhang 0004, Tao Huang 0005, F. Richard Yu |
IEEE Internet Things J. | 4 |
| 2022 | Distributed Task Scheduling in Serverless Edge Computing Networks for the Internet of Things: A Learning ApproachabstractBy delegating the infrastructure management, such as provisioning or scaling to third-party providers, serverless edge computing has recently been widely adopted in several applications, especially Internet of Things (IoT) applications. Task scheduling is a critical issue in serverless edge computing as it significantly impacts the quality of user experience. In contrast to the centralized scheduling in the cloud center, serverless edge task scheduling is more challenging due to the heterogeneous and resource-constrained nature of edge resources. This article aims to study the distributed task scheduling for the IoT in serverless edge computing networks, in which heterogeneous serverless edge computing nodes are rational individuals with interests to optimize their own scheduling utility while the nodes only have access to local observations. The task scheduling competition process is formulated as a partially observable stochastic game (POSG) to enable serverless edge computing nodes to noncooperatively schedule tasks and allocate computing resources depending on their locally observed system state, which takes into account the associated task generation state, data queue state, communication channel state, and previous computing resource allocation state. To solve the proposed POSG and deal with the partial observability, a multiagent task scheduling algorithm based on the dueling double deep recurrent$Q$-network (D3RQN) method is developed to approximate the optimal task scheduling and resource allocation solution. Finally, extensive simulation experiments are conducted to validate the effectiveness and superiority of the proposed scheme. Qinqin Tang, Renchao Xie, F. Richard Yu, Tianjiao Chen, Ran Zhang 0004, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Internet Things J. | 6 |
| 2022 | Buffer-Aware Virtual Reality Video Streaming With Personalized and Private Viewport PredictionabstractViewport prediction and prefetch have an important influence on VR video streaming performance. This work proposes a novel federated learning-based viewport prediction model training algorithm, ComPer-FedAvg. The proposed algorithm leverages a VR video’s common viewing pattern and users’ personal viewing patterns to train the prediction model in a distributed and privacy-preserving manner. Further, considering the VR video viewport prediction accuracy, a stochastic game is formulated to solve the VR streaming network’s communication resource allocation problem, where limited communication resource blocks are auctioned to users to achieve the optimal overall VR viewing experience. For each user, the auction is decomposed into two disjoint subproblems, namely, the optimal number of data rate requesting and true value claiming (bidding). The optimal true value claiming has been analytically proved to be equal to the VR viewing reward with given data rate. Due to the lack of global information when users request data rate, we reformulate users’ data rate requesting problem as a POMDP problem. A novel deep reinforcement learning algorithm is adopted to solve the problem. Evaluation and simulation results show the proposed viewport prediction and VR streaming schemes outperform conventional solutions in terms of prediction accuracy and VR viewing experience. Ran Zhang 0004, Jiang Liu 0010, Fangqi Liu 0002, Tao Huang 0005, Qinqin Tang, Shangguang Wang, F. Richard Yu |
IEEE J. Sel. Areas Commun. | 4 |
| 2021 | HierCC: Hierarchical RDMA Congestion ControlabstractRDMA has been increasingly deployed in data centers to decrease latency and CPU utilization. However, existing RDMA congestion control schemes fail to address instantaneous large queue build-up or bandwidth under-utilization associated with frequent traffic bursty. In this paper, we argue that traffic uncertainty is the essential reason that constrains data center congestion control from simultaneously achieving high throughput and deterministic latency. Since aggregated flows within the same rack are relatively long-lived, we propose HierCC, which aggregates flows destined to the same IP in a rack and hierarchically controls the rate of flows. The rate of aggregate flows between racks is controlled by a credit-based congestion control mechanism. Then the bandwidth obtained by an aggregate flow in a rack is allocated to the corresponding individual flows from that rack promptly and accurately. We evaluate HierCC using SystemC and large-scale NS3 simulations. Results indicate that HierCC can significantly mitigate buffer usage and reduce the 99th percentile FCT by up to 20% and 40% compared with HPCC and DCQCN under a realistic workload, respectively. Jiao Zhang 0002, Zixuan Guan, Zirui Wan, Yinben Xia, Tian Pan 0001, Tao Huang 0005, Dezhi Tang |
APNet | 7 |
| 2021 | Enabling In-band Network Telemetry in Software-based Virtual SwitchesabstractSoftware-based virtual switches are indispensable in multi-tenant cloud networks. They either work as bridges between virtual machines and the underlying networks, or act as the virtual network function carriers for flexible service chaining and orchestration. Therefore, high-accuracy monitoring of virtual switches is significant for ease of data center network management. The recently proposed In-band Network Telemetry, which relies on the protocol-independent switch architecture (PISA), can achieve the monitoring requirements. However, not all the virtual switches with production quality are P4-based or built under the PISA. In this work, we provide the design and implementation of label-based INT and probe-based INT on top of OVS and VPP, the two mainstream software-based virtual switches with non-PISA architecture. Extensive evaluation shows that our implementation has low performance overhead in terms of forwarding latency, packet loss ratio and CPU consumption. Under 100Mbps traffic pressure, the CPU overhead of INT on OVS and VPP are less than 0.1% and 0.5%, and the switch latency are added by less than$4 \mu\mathrm{s}$and$3\mu\mathrm{s}$, respectively. Tian Pan 0001, Xingchen Lin, Yan Zhang 0063, Houtian Wang, Tao Huang 0005, Yunjie Liu 0001 |
GLOBECOM | 6 |
| 2021 | Multi-path Transmission Scheme Based on Segment Control in Low-Earth-Orbit Satellite NetworkabstractBecause of the challenges brought by the high dynamic topology of satellite networks to the transport layer, this paper is mainly devoted to a multipath transmission control protocol (MPTCP) path selection scheme based on segment control technology and software-defined networking (SDN). We describe the signaling interaction mode of MPTCP and the process of segment control technology applied in the satellite network. According to the requirements of real-time and accuracy of data transmission, we consider the link delay, stability, and packet loss rate, and construct the scheme as a maximum-flow minimum-cost problem. The experimental results show that the proposed scheme can meet the low delay requirements of delay-sensitive traffic flow, improve bandwidth utilization, and ensure more efficient and reliable data transmission. Man Ouyang, Xuefei Duan, Jiang Liu 0010, Ran Zhang 0004, Tao Huang 0005, Lu Hua |
HPSR | 5 |
| 2021 | TACQ: Enabling Zero-jitter for Cyclic-Queuing and Forwarding in Time-Sensitive NetworksabstractRecent proposals leverage Cyclic-Queuing and Forwarding (CQF) to achieve bounded-delay transmission for cyclic flows in Time-Sensitive Networking (TSN). However, the Ping-Pong queue-based model in CQF will introduce the jitter of two cycles, which is inapplicable to industrial automation scenarios where isochronous flows such as synchronized frames and motor control loops require zero jitter.We present TACQ, a first time-aware cyclic-queuing mechanism to support the co-transmission of cyclic flows and isochronous flows. Our key insight is that isochronous flows should enable zero-jitter while minimally impacting the bounded-delay of cyclic flows. To achieve this goal, we design a novel no-wait shaper (NWS) that handles isochronous flows with as little time-slot as possible. For cyclic flows, we extend the CQF to schedule flows with a double-closed state on the Tx-gate and compute the open time on the Rx-gate. Simulation results show that TACQ significantly reduces the delay of isochronous flows by 81.3% and achieves zero jitter compared with CQF. And the NWS effectively schedules 87.2% flows at 500 flows level in feasible execution time. Yudong Huang, Shuo Wang 0006, Binwei Wu, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 4 |
| 2021 | Online Routing and Scheduling for Time-Sensitive NetworksabstractRecent proposals leverage Time-Aware Shaper (TAS) to achieve precise transmission in Time-Sensitive Networking (TSN). However, most of the proposals require the information of all time-triggered flows to be known in advance and synthesize the gate control list of each switch offline, making the mechanisms they designed inapplicable to industrial automation scenarios where the devices are changed dynamically and the flows should be scheduled online. In this paper, we propose an online routing and scheduling mechanism of TAS for time-sensitive networks. In order to maximize the number of schedulable flows and reduce bandwidth waste, we devise the variable time slot mechanism and minimize the sending start time of each flow. Based on these mechanisms, a novel incremental routing and scheduling (IRAS) algorithm is designed to achieve per-flow deployment, with a pre-routing algorithm to reduce synthesis time. The evaluations show that the IRAS algorithm approaches 96.5 % of the optimal solution in scheduling 2000 flows, and has a feasible per-flow computational time from sub-seconds to less than ten seconds. Yudong Huang, Shuo Wang 0006, Tao Huang 0005, Binwei Wu, Yunxiang Wu, Yunjie Liu 0001 |
ICDCS | 3 |
| 2021 | A Refined Dijkstra's Algorithm with Stable Route Generation for Topology-Varying Satellite NetworksabstractSpaceX plans ambitiously to launch approximately 12,000 satellites from 2019 to 2024, expected to be a complement or even competitor to ground networks. However, the mega-scale satellite network is topology-varying and the frequency of inter-satellite link (ISL) handovers increases rapidly as the topology expands, which will further arouse a massive number of route updates with considerable packet travel delay or even packet loss during the route convergence. The classic Dijkstra's algorithm is adopted for space route calculation, however, it always selects the default shortest path from multiple equal-cost shortest paths between two satellite nodes. To reduce the route change as much as possible during the periodical topology change, in this work, we refined the original Dijkstra and propose StableRoute to select the most appropriate route from the equal-cost candidates with the least route updates compared with the routing table last round. In this way, the end-to-end paths can be maintained as far as possible without time-to-time oscillation. Evaluation shows that it reduces 41% of the route updates in a 36 × 36 topology compared with Dijkstra, and the reduction rate will rise persistently with the growth of the satellite constellation. Zhengjie Luo, Tian Pan 0001, Enge Song, Houtian Wang, Wenhao Xue, Tao Huang 0005, Yunjie Liu 0001 |
ICDCS | 6 |
| 2021 | INT-probe: Lightweight In-band Network-Wide Telemetry with Stationary ProbesabstractVisibility is essential for operating and troubleshooting intricate networks. In-band Network Telemetry (INT) has been embedded in the latest merchant silicons to offer high-precision device and traffic state visibility. INT is actually an underlying technique and each INT instance covers only one monitoring path. The network-wide measurement coverage therefore requires a high-level orchestration to provision multiple INT paths. An optimal path planning is expected to produce a minimum number of paths with a minimum number of overlapping links. Eulerian trail has been used to solve the general problem. However, in production networks, the vantage points where one can deploy probes to start and terminate INT paths are constrained. In this work, we propose an optimal path planning algorithm, INT-probe, which achieves the network-wide telemetry coverage under the constraint of stationary probes. INT-probe formulates the constrained path planning into an extended multi-depot k-Chinese postman problem (MDCPP-set) and then reduces it to a solvable minimum weight perfect matching problem. We analyze algorithm's theoretical bound and the complexity. Extensive evaluation on both wide area networks and data center networks with different scales and topologies are conducted. We show INT-probe is efficient, high-performance, and practical for real-world deployment. For a large-scale data center networks with 1125 switches, INT-probe can generate 112 monitoring paths (reduced by 50.4 %) by allowing only 1.79% increase of the total path length, promptly resolving link failures within 744.71ms. Tian Pan 0001, Xingchen Lin, Haoyu Song 0001, Enge Song, Zizheng Bian, Hao Li 0011, Jiao Zhang 0002, Fuliang Li, Tao Huang 0005, Chenhao Jia, Bin Liu 0001 |
ICDCS | 9 |
| 2021 | Loom: Switch-based Cloud Load Balancer with Compressed StatesabstractLayer-4 load balancers play a critical role in large-scale data centers. Recently, load balancers implemented on programmable switches have attracted much attention since they overcome the inflexibility of dedicated load balancers and high latency of software load balancers. However, keeping per-connection state easily leads to storage exhaustion, especially under resource exhaustion attacks. Although several stateless load balancers are proposed to address this issue, the state management burden is offloaded to backend servers, causing high deployment and running costs. In this paper, a load balancer called Loom with compressed states is proposed for large-scale data centers. Firstly, we propose a novel classifier-based load balancer idea to avoid directly maintaining per-connection state. Then, a circulating Bloom filter structure is proposed that can efficiently classify connections as well as be implemented on existing programmable switches. Theoretical analysis shows that Loom can maintain 11 ~ 30x more concurrent connections than those directly storing the 5-tuple of connections. Loom is implemented in hardware P4 switches and experimental results indicate that 11 ~ 29x more concurrent connections can be maintained in Loom, which is close to the theoretical results. Besides, Loom is resistant to resource exhaustion attacks and reduces the percentage of broken connections by up to 57% with an SYN flood. Jiao Zhang 0002, Shubo Wen, Tian Pan 0001, Tao Huang 0005 |
ICNP | 5 |
| 2021 | Receiver-Driven RDMA Congestion Control by Differentiating Congestion Types in Datacenter NetworksabstractThe development of datacenter applications leads to the need for end-to-end communication with microsecond latency. As a result, RDMA is becoming prevalent in datacenter networks to mitigate the latency caused by the slow processing speed of the traditional software network stack. However, existing RDMA congestion control mechanisms are either far from optimal in simultaneously achieving high throughput and low latency or in need of additional in-network function support. In this paper, by leveraging the observation that most congestion occurs at the last hop in datacenter networks, we propose RCC, a receiver-driven rapid congestion control mechanism for RDMA networks that combines explicit assignment and iterative window adjustment. Firstly, we propose a network congestion distinguish method to classify congestions into two types, last-hop congestion and innetwork congestion. Then, an Explicit Window Assignment mechanism is proposed to solve the last-hop congestion, which enables senders to converge to a proper sending rate in one-RTT. For in-network congestion, a PID-based iterative delay-based window adjustment scheme is proposed to achieve fast convergence and near-zero queuing latency. RCC does not need additional innetwork support and is friendly to hardware implementation. In our evaluation, the overall average FCT (Flow Completion Time) of RCC is 4~79% better than Homa, ExpressPass, DCQCN, TIMELY, and HPCC. Jiao Zhang 0002, Jiaming Shi, Xiaolong Zhong, Zirui Wan, Tian Pan 0001, Tao Huang 0005 |
ICNP | 7 |
| 2021 | INT-label: Lightweight In-band Network-Wide Telemetry via Interval-based Distributed LabellingabstractThe In-band Network Telemetry (INT) enables hop-by-hop device-internal state exposure for reliably maintaining and troubleshooting data center networks. For achieving network-wide telemetry, orchestration on top of the INT primitive is further required. One straightforward solution is to flood the INT probe packets into the network topology for maximum measurement coverage, which, however, leads to huge bandwidth overhead. A refined solution is to leverage the SDN controller to collect the topology and carry out centralized probing path planning, which, however, cannot seamlessly adapt to occasional topology changes. To tackle the above problems, in this work, we propose INT-label, a lightweight In-band Network-Wide Telemetry architecture via interval-based distributed labelling. INT-label periodically labels device-internal states onto sampled packets, which is cost-effective with minor bandwidth overhead and able to seamlessly adapt to topology changes. Furthermore, to avoid telemetry resolution degradation due to loss of labelled packets, we also design a feedback mechanism to adaptively change the instant label frequency. Evaluation on software P4 switches suggests that INT-label can achieve 99.72% measurement coverage under a label frequency of 20 times per second. With adaptive labelling enabled, the coverage can still reach 92% even if 60% of the packets are lost in the data plane. Enge Song, Tian Pan 0001, Chenhao Jia, Wendi Cao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
INFOCOM | 6 |
| 2021 | GreenTE.ai: Power-Aware Traffic Engineering via Deep Reinforcement LearningabstractPower-aware traffic engineering via coordinated sleeping is usually formulated into Integer Programming problems, which are generally NP-hard with unbounded computation time for large-scale networks. This results in delayed control decision making in dynamic network environments. Motivated by advances in deep Reinforcement Learning, we consider building intelligent systems that learn to adaptively change router/switch’s power state according to changing network conditions. Neural network’s forward propagation can greatly speed up power on/off decision making. Generally, conducting RL requires a learning agent to iteratively explore and perform the "good" actions based on the feedback from the environment. By coupling Software-Defined Networking for performing centrally calculated actions to the environment and In-band Network Telemetry for collecting feedback from the environment, we develop GreenTE.ai, a closed-loop control/training system to automate power-aware traffic engineering. Furthermore, we propose novel techniques to enhance the learning ability and reduce the learning complexity. With both energy efficiency and traffic load balancing considered, GreenTE.ai can generate reasonable power saving actions within 276ms under a network testbed of 11 software P4 switches. Tian Pan 0001, Xiaoyu Peng, Zizheng Bian, Xingchen Lin, Enge Song, Fuliang Li, Yang Xu 0010, Tao Huang 0005 |
IWQoS | 9 |
| 2021 | Sailfish: accelerating cloud-scale multi-tenant multi-service gateways with programmable switchesabstractThe cloud gateway is essential in the public cloud as the central hub of cloud traffic. We show that horizontal scaling of software gateways, once sustainable for years, is no longer future-proof facing the massive scale and rapid growth of today's cloud. The root cause is the stagnant performance of the CPU core, which is prone to be overloaded by heavy hitters as traffic growth goes far beyond Moore's law. To address this, we propose \emph{Sailfish}, a cloud-scale multi-tenant multi-service gateway accelerated by programmable switches. The new challenge is that large forwarding tables due to multi-tenancy cannot be fit into the limited on-chip memories. To this end, we devise a multi-pronged approach with (1) hardware/software co-design for table sharing, (2) horizontal table splitting among gateway clusters, (3) pipeline-aware table compression for a single node. Compared with the x86 gateway of a similar price, Sailfish reduces latency by 95% (2μs), improves throughput by more than 20x in bps (3.2Tbps) and 71x in pps (1.8Gpps) with packet length < 256B. Sailfish has been deployed in Alibaba Cloud for more than two years. It is the first P4-based cloud gateway in the industry, of which a single cluster carries dozens of Tbps traffic, withstanding peak-hour traffic in large online shopping festivals. Tian Pan 0001, Nianbing Yu, Chenhao Jia, Jianwen Pi, Yisong Qiao, Jianyuan Lu, Enge Song, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu |
SIGCOMM | 13 |
| 2021 | A novel identity resolution system design based on Dual-Chord algorithm for industrial Internet of Things
Renchao Xie, F. Richard Yu, Tao Huang 0005, Yunjie Liu 0001 |
Sci. China Inf. Sci. | 4 |
| 2021 | Dynamic Computation Offloading in IoT Fog Systems With Imperfect Channel-State Information: A POMDP ApproachabstractDriven by the growing popularity of mobile applications, such as the Internet of Things (IoT), fog computing has been envisioned as a promising approach to enhance the computation capability of mobile devices and reduce the energy consumption. In this article, we aim to investigate the dynamic computation offloading problem in the IoT fog system under the fast time-varying wireless channel conditions. Our work differs from the existing work, which is based on the assumption that the channel-state information can be perfectly obtained by the offloading agent (e.g., the IoT device). In reality, due to hardware limitation, short sensing time, and network connectivity issues in IoT fog systems, it is difficult for the IoT device to have the perfect knowledge of a dynamic channel environment. Therefore, in this article, we propose a partially observable offloading scheme to enable the IoT device to make the optimal offloading decision with imperfect channel-state information. The optimization problem is formulated as a partially observable Markov decision process (POMDP) formulation, with the objective of minimizing the IoT device's energy consumption while meeting its requirement on task processing delay. To find the optimal offloading solution, an offline algorithm based on the deep recurrent $Q$ -network (DRQN) is developed. Finally, extensive simulation experiments are performed to evaluate the effectiveness of the proposed offloading scheme. Renchao Xie, Qinqin Tang, Chenghao Liang, F. Richard Yu, Tao Huang 0005 |
IEEE Internet Things J. | 5 |
| 2021 | NB-Cache: Non-Blocking In-Network Caching for High-Performance Content RoutersabstractInformation-Centric Networking (ICN) provides scalable and efficient content distribution at the Internet scale due to in-network caching and native multicast. To support these features, a content router needs high performance at its data plane, which consists of three forwarding steps: checking the Content Store (CS), then the Pending Interest Table (PIT), and finally the Forwarding Information Base (FIB). In this work, we build an analytical model of the router and identify that CS is the actual bottleneck. Then, we propose a novel mechanism called “NB-Cache” to address CS’s performance issue from a network-wide point of view. In NB-Cache, when packets arrive at a router whose CS is fully loaded, instead of being blocked and waiting for the CS, these packets are forwarded to the next-hop router, whose CS may not be fully loaded. This approach essentially utilizes Content Stores of all the routers along the forwarding path in parallel rather than checking each CS sequentially. NB-Cache follows a design pattern of on-demand load balancing and can be formulated into a non-trivial N-queue bypass model. We use the Markov chain to establish its theoretical base and find an algorithm for automated transition rate matrix generation. Experiments show significant improvement of data plane performance: 70% reduction in round-trip time (RTT) and 130% increase in throughput. NB-Cache decouples the fast packet forwarding from the slower content retrieval thus substantially reducing CS’s heavy dependency on fast but expensive memory. Tian Pan 0001, Xingchen Lin, Enge Song, Jiao Zhang 0002, Hao Li 0011, Jianhui Lv, Tao Huang 0005, Bin Liu 0001, Beichuan Zhang 0001 |
IEEE/ACM Trans. Netw. | 8 |
| 2020 | PLB: Adaptive Partial Congestion-aware Load Balancing for Datacenter NetworksabstractIn order to accommodate ever-increasing new tenants and applications, datacenter networks (DCNs) require an efficient load balancing scheme to fully utilize their bisection bandwidth. Equal-cost MultiPath routing (ECMP) is a widely used load-balancing mechanism in the DCN. However, ECMP blindly hashes traffic to parallel paths and results in imbalance and collisions. Motivated by ECMP's shortcomings, some recent schemes provide more visibility into networks via active probing. They could be broadly classified as probing all the paths or a fixed number of paths (e.g., 3 paths) each probe interval. However, they all suffer from some limitations. Probing all paths introduces high probing overhead while probing a fixed number of paths is suboptimal when the network topology and traffic load change. To our best knowledge, none of the existing schemes adapt the number of paths being probed to the network conditions. Enlightened by the defects of previous work, we introduce PLB, an adaptive partial congestion-aware load-balancing mechanism. At its heart, PLB randomly probes partial paths each probe interval and the number of them changes according to the network topology and the traffic load. Besides, PLB splits flow into flowlets and makes careful routing/rerouting decisions for them. Through analysis, we formulate the correlations between the number of paths being probed and the network conditions. Furthermore, simulations with realistic workloads validate our conclusions and show that PLB reduces overall flow completion times compared to the state-of-the-art load balancing schemes both in symmetric and asymmetric topologies. Kefei Liu 0004, Jiao Zhang 0002, Dehui Wei, Tao Huang 0005 |
GLOBECOM | 5 |
| 2020 | INT-filter: Mitigating Data Collection Overhead for High-Resolution In-band Network TelemetryabstractIn-band Network Telemetry (INT) enables fine-grained network monitoring to ease the management of large-scale networks, which, however, relies on the real-time collection of a huge amount of telemetry data through the southbound interface. For example, the INT telemetry data upload rate of a 28-pod FatTree topology reaches 3Tbps under a probe frequency of 100 times/s, which is rather unacceptable since the controller-switch link bandwidth is limited. To mitigate the telemetry data collection overhead, in this work, we propose INT-filter, a novel measurement architecture that deploys the same prediction algorithm on both the data plane and the control plane to predict the traffic state in the near future instead of uploading all the telemetry data. Such prediction-based approach leverages the observation that there is considerable redundancy in the telemetry data sequence. In addition, we design an integration mechanism that conducts predictions using multiple methods simultaneously and uploads the predicted result from the least-error method to further decrease the upload volume. Extensive evaluation suggests that INT-filter can achieve at least 33.6% data collection decrease under a 10ms probe interval. With prediction integration, the upload reduction can further reach 58.5%. Enge Song, Tian Pan 0001, Chenhao Jia, Wendi Cao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
GLOBECOM | 6 |
| 2020 | Multi-Constraint Virtual Network Embedding Algorithm For Satellite NetworksabstractSatellite network constellation is promising in providing efficient global Internet access. While the constellation scale, the user population, and service variety in satellite networks are too large, requiring efficient resource allocation and network management. Network virtualization is an efficient solution to achieve preceding objectives, but conventional schemes on terrestrial networks are not well adapted to satellite networks. Therefore, in this work, we establish a network virtualization model considering topology dynamics, quality of service requirement, and resource constraint. Then we formulate Virtual Network Embedding (VNE) into optimization problems, and we propose a multi-constraint virtual network embedding algorithm to solve the problem. Finally, we evaluate the proposed scheme and prove its adaptability to satellite networks. Jiang Liu 0010, Ran Zhang 0004, Tao Huang 0005 |
GLOBECOM | 4 |
| 2020 | Optimal Proactive Caching Placement for Named Data Networking with Interest AggregationabstractOn-path caching is a building block in Named Data Networking that helps eliminate redundant traffic. The performance of redundancy elimination depends on both Content Store (CS) and Pending Interest Table (PIT), i.e., CS caches content for future reuse, and PIT aggregates repetitive requests in a short period. However, contemporary proactive caching strategies only take account of CS while neglecting PIT. In this work, we integrate both PIT and CS into the proactive caching model, derive how to calculate aggregated request rate, and propose an algorithm to calculate the aggregated request rate across the tree topology. Then we formulate caching placement into optimization problems and solve them with a decomposition-based evolutionary algorithm. The simulation results show that the proposed scheme outperforms conventional solutions. Ran Zhang 0004, Jiang Liu 0010, Tao Huang 0005, Renchao Xie, F. Richard Yu, Yunjie Liu 0001 |
GLOBECOM | 3 |
| 2020 | DRA-IG: The Balance of Performance Isolation and Resource Utilization Efficiency in Network SlicingabstractNetwork slicing (NS) is a promising technology of 5G that provides customized end-to-end network service to multi-tenant. How to improve resource utilization efficiency with guarantee of performance isolation in a shared infrastructure is one of the main challenges in resource allocation problem of NS. To address this challenge, we characterize the degree of performance isolation based on the relationship between the requested resource amount, the allocated resource amount, and the time-varying network loads. We propose a dynamic resource allocation problem with probabilistic isolation guarantee (DRAIG), which is formulated as a chance constrained program. As the true probability distribution of network loads is usually unknown, we use the Conditional Valuate-at-Risk (CVaR) measure to provide a distributionally robust formulation that approximate the basic chance constraints of DRA-IG in a data-driven manner. We estimate the second-order moment of network loads by the periodic history information. Then, we further reformulate the distributionally robust optimization problem as a tractable semidefinite programming (SDP). Finally, numerical evaluation verifies the effectiveness of the proposed method. Jiang Liu 0010, Tao Huang 0005, Yunjie Liu 0001 |
ICC | 3 |
| 2020 | Data-driven Routing Optimization based on Programmable Data PlaneabstractTo meet the growing demand for high bandwidth of Multimedia network, IP Network Providers spend millions of dollars overprovisioning bandwidth of their network. However, due to the lack of reasonable traffic scheduling, the over-provisioning network still has a severe issue of utilization imbalance. Traffic Engineering (TE) is proposed to solve this problem. Network measurement and routing optimization strategies are two key components of TE. Effective real-time network measurement provides the basis for the generation of route optimization strategies, which makes the network congestion-aware. Existing out-band network telemetry that transmits extra probes to measure network status has the problem of inaccurate measurement information in the network. Besides, the relationship between complex network status and routing optimization strategy is difficult to describe with an exact mathematical model. Therefore, we propose a novel TE approach, which is called DPRO. It combines In-band Network Telemetry based on programmable language P4 with Reinforcement Learning to minimize network max-link-utilization. Extensive experiments show that our approach significantly outperforms several widely-used baseline methods in terms of max-link-utilization. Qian Li 0006, Jiao Zhang 0002, Tian Pan 0001, Tao Huang 0005, Yunjie Liu 0001 |
ICCCN | 4 |
| 2020 | Hieff: Enabling Efficient VNF Clusters by Coordinating VNF Scaling and Flow SchedulingabstractA cluster of Virtual Network Functions (VNF) can serve massive fluctuating traffic by managing VNF instances and distributing flows. However, how to schedule the flows and manage VNF scaling efficiently in a VNF cluster is still an open question. Existing solutions such as hash based schemes encounter imbalance and passive flow remapping obstacles while flow table-based scheme suffers from high processing latency and flow entries overflow challenges. In this paper, we design and present Hieff, an efficient NFV system that coordinates VNF scaling and flow scheduling within a VNF cluster. The key idea of Hieff is to precisely manage the heavy flows with a flow table while simply allowing light flows be distributed by hash. Though this idea has been explored in previous work, we are the first to apply it with the VNF scaling process. We mathematically model the Hieff system and propose a heuristic algorithm to determine the optimized VNF scaling and flow scheduling strategies. We implement Hieff based on BESS and Click and use real-world tracing to evaluate the system. Results show that Hieff can handle co-existing massive flows efficiently with low latency while balancing the load of VNF instances at low cost. Jiao Zhang 0002, Tao Huang 0005 |
IPCCC | 4 |
| 2020 | Rapid Detection and Localization of Gray Failures in Data Centers via In-band Network TelemetryabstractNetwork reliability becomes increasingly important in modern data center networks (DCNs). The DCNs are expected to work sustainably under internal failures and assist network operators in troubleshooting them rapidly. However, some network failures will happen silently with packets discarded without producing any explicit notification before causing tremendous damage to the network. To troubleshoot these "gray failures", in this work, we present a rapid gray failure detection and localization mechanism based on the recently proposed In-band Network Telemetry (INT). Specifically, we leverage simplified INT probe packets to conduct network-wide telemetry to help the servers under ToR switches obtain all the feasible paths between sources and destinations. Once a network failure occurs, the affected thus unavailable paths will immediately be detected and flushed out of the path information table at each server by a timeout mechanism. Hence, servers can proactively perform source routing-based fast traffic reroute to avoid massive packet loss and retain uninterrupted quality of experience. At the meantime, all the aged path entries will be uploaded to a remote controller for centralized failure localization by identifying common path elements. To verify the feasibility of our design, we build a virtual network testbed with software P4 switches and a Redis database. Evaluation shows that our system can successfully detect network gray failures and reroute the affected traffic in no time while complete failure localization within only a few seconds. Chenhao Jia, Tian Pan 0001, Zizheng Bian, Xingchen Lin, Enge Song, Tao Huang 0005, Yunjie Liu 0001 |
NOMS | 7 |
| 2020 | Service-aware optimal caching placement for named data networking
Ran Zhang 0004, Jiang Liu 0010, Renchao Xie, Tao Huang 0005, F. Richard Yu, Yunjie Liu 0001 |
Comput. Networks | 4 |
| 2020 | Threshold-oblivious on-line web QoE assessment using neural network-based regression modelabstractThe evaluation of the web‐browsing quality of experience (QoE) is difficult to complete through traditional methods (e.g. deducing formulas or setting thresholds) due to the diversity of websites and their contents. To evaluate web‐browsing QoE through a general way, the authors propose a web QoE evaluation architecture based on machine learning, consisting of two parts: traffic classification sub‐system and QoE prediction sub‐system. When evaluating user experience, traffic classification sub‐system first classifies the packets generated by visiting a website into a flowthrough some fields in the packet header, to model each website separately. The traffic classification accuracy of packets over six websites reaches 96.63%. Then, in the network layer, the traffic metric cumulative traffic volume is generated from the size and arrival time of packets. When a user visits a web page, their regression model predicts the above‐the‐fold time (ATF) and thus QoE. The output of the regression model is an exact ATF value that is mapped to user experience. In addition, reversing input variables further improves the model, which is evaluated on two popular websites. The QoE prediction results of the improved method for 5400 visits are obtained within 0.0975 s, reaching 0.9 . Enge Song, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Wendi Cao, Tao Huang 0005 |
IET Commun. | 7 |
| 2020 | Decentralized Computation Offloading in IoT Fog Computing System With Energy Harvesting: A Dec-POMDP ApproachabstractRecently, fog computing has emerged as a prospective technique to provide pervasive and agile computation services for Internet-of-Things (IoT) devices and support advanced applications. Introducing the energy harvesting (EH) technique into the fog computing system can extend the battery lifetime and provide a higher quality of experiences (QoE) for IoT devices. In the EH-enabled IoT fog system, computation offloading is an important issue and has attracted much attention. In most existing works, it is assumed that the IoT device is fully aware of the system state. However, in practical offloading problems, the IoT device may not be able to obtain accurate system state information, and only have a partial observation of the environment. Therefore, in this article, we investigate the decentralized partially observable offloading problem in the EH-enabled IoT fog system, in which multiple IoT devices cooperate to maximize the network performance while meeting their QoE requirements. We formulate the optimization problem as a decentralized partially observable Markov decision process (Dec-POMDP) in which each IoT device makes the task offloading decisions according to its local observation of the environment. The Lagrangian approach and the policy gradient method are adopted to find the optimal solution for the proposed problem. Due to the high complexity of solving the Dec-POMDP, a learning-based decentralized offloading algorithm with low complexity is presented to find the approximate optimal solution. Finally, extensive experimental evaluation and comparison are carried out to show the effectiveness of the proposed scheme. Qinqin Tang, Renchao Xie, F. Richard Yu, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Internet Things J. | 4 |
| 2020 | The source-multicast: A sender-initiated multicast member management mechanism in SRv6 networks
Weihong Wu, Jiang Liu 0010, Tao Huang 0005 |
J. Netw. Comput. Appl. | 3 |
| 2020 | Fast Switch-Based Load Balancer Considering Application Server StatesabstractLarge-scale services are generally hosted on multiple application servers to scale out in today's data centers. Load balancers distribute users' requests across these servers. Software load balancer and switch-based load balancer are two typical classes of load balancers. However, most of the existing mechanisms either exhibit high processing latency at load balancers or likely lead to unbalanced requests distribution without considering the disparity of the application servers. In this paper, we study how the disparity of application servers significantly impacts the response time of requests. A fast switch-based Load Balancer considering Application Server states (LBAS) then is proposed to minimize the processing latency at both load balancers and application servers. The data plane of LBAS is well designed to store millions of connections in limited storage capacity without violating per-connection consistency. Besides, a partial dynamic weighting algorithm based on the Ridge Regression theory is designed and implemented to decrease the processing latency at application servers. We implement LBAS using the P4 programming language and conduct a series of extensive experiments to evaluate the performance. The results demonstrate that the proposed LBAS mechanism significantly reduces the response time of requests compared with Uniform random, Static weight, and Spotlight in various scenarios. Jiao Zhang 0002, Shubo Wen, Jinsheng Zhang, Tian Pan 0001, Tao Huang 0005, Linquan Zhang, Yunjie Liu 0001, F. Richard Yu |
IEEE/ACM Trans. Netw. | 6 |
| 2020 | Deep Reinforcement Learning (DRL)-Based Device-to-Device (D2D) Caching With Blockchain and Mobile Edge ComputingabstractDevice-to-Device (D2D) caching assists Mobile Edge Computing (MEC) based caching in offloading inter-domain traffic by sharing cached items with nearby users, while its performance relies heavily on caching nodes' sharing willingness. In this paper, a Blockchain-based Cache and Delivery Market (CDM) is proposed as an incentive mechanism for the distributed caching system. Under given incentive mechanisms, both D2D and MEC caching nodes' willingness is guaranteed by satisfying their expected reward for cache sharing. Besides, for the distributed CDM, content delivery related transactions are executed by smart contracts. To achieve consensus on transactions and prevent frauds, a consensus protocol among the smart contract execution nodes (SCENE) is necessary. To minimize the latency of reaching consensus while guaranteeing its confidence level, we propose partial Practical Byzantine Fault Tolerance (pPBFT) protocol. Further, the model of cache sharing and transaction execution consensus is proposed, and we further formulate caching placement and SCENE selection as Markov Decision Process problems. Due to the complexity and dynamics of the problems, a deep reinforcement learning approach is adopted to solve the problem. The simulation results show that the proposed schemes outperform conventional solutions in terms of traffic offloading, content retrieval latency, and consensus latency. Ran Zhang 0004, F. Richard Yu, Jiang Liu 0010, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2019 | Service-Aware Optimal Caching Placement for Named Data NetworkingabstractBuilt-in caching in Named Data Networking (NDN) promises to provide efficient content delivery, where the dedicated on-path caching scheme is deployed to serve users' requests on the forwarding path. In this work, to utilize limited caching resources to achieve optimal performance, the caching placement decision is made by jointly considering the content popularity, underlying network topology, forwarding strategy and caching service mechanism in NDN. More specifically, we propose a service-aware caching model. In the model, we first define the Cache Service Matrix (CSM), which describes the position where each user's request is served for each piece of content. In order to make CSM comply with the caching placement, underlying topology, forwarding strategy, and on-path caching service mechanism, we propose an algorithm to calculate CSM under the preceding constraints. With CSM, the utility of caching placement could be derived correctly, and we formulate the optimal caching placement into optimization problems. Moreover, the differential grouping co-evolutionary (DG2-E) algorithm is adopted to decompose and solve the NP-hard optimization problems. Simulation results show the proposed scheme outperforms state of the art solutions in terms of inter-domain traffic reducing and request-response accelerating under arbitrary topologies. Ran Zhang 0004, Jiang Liu 0010, Renchao Xie, Tao Huang 0005, F. Richard Yu |
GLOBECOM | 4 |
| 2019 | OPSPF: Orbit Prediction Shortest Path First Routing for Resilient LEO Satellite NetworksabstractWith global coverage as well as ultra-low latency, the Low-Earth-Orbit (LEO) satellite constellation is regarded as an ideal complement to the terrestrial network infrastructure. One technical issue in LEO satellite networks is efficient and resilient routing. Considering the periodic topology changes, straightforwardly leveraging terrestrial routing protocols, such as OSPF, will incur endless route convergence, consuming expensive inter-satellite link bandwidth. Prior work proposes several snapshot-based routing approaches, which either require to store a sequence of routing table snapshots in limited satellite memory, or have to maintain frequent interaction with the ground stations. In this work, we propose OPSPF, a novel routing protocol dedicated to LEO satellite networks. OPSPF takes advantage of the regularity of the constellation and conducts periodic route calculation for instantaneous routing table generation, which well handles the regular topology changes. Moreover, OPSPF proposes an on-demand dynamic routing mechanism, dedicated to the irregular topology changes caused by link failure/recovery. Evaluation shows, compared with OSPF, OPSPF has zero route convergence overhead during regular topology changes and 57% reduction of the communication overhead and 82% reduction of the route convergence time during irregular topology changes. Tian Pan 0001, Tao Huang 0005, Wenhao Xue, Yunjie Liu 0001 |
ICC | 2 |
| 2019 | INT-path: Towards Optimal Path Planning for In-band Network-Wide TelemetryabstractWith the ever-increasing complexity of networks, fine-grained network monitoring enables better network reliability and timely feedback control. The In-band Network Telemetry (INT) allows cost-effective network monitoring by encapsulating device-internal states into probe packets. However, INT only specifies an underlying device-level primitive while how to achieve network-wide traffic monitoring remains undefined. In this work, we propose INT-path, a network-wide telemetry framework, by decoupling the system into a routing mechanism and a routing path generation policy. Specifically, we embed source routing into INT probes to allow specifying the route the probe packet takes through the network. Above the mechanism, we develop an Euler trail-based path planning policy to generate non-overlapped INT paths that cover the entire network with a minimum path number. Besides, an exhaustive analysis of algorithm's run-time complexity is also provided. INT-path can “encode” the network-wide traffic status into a series of “bitmap images”, transforming network troubleshooting into pattern recognition problems. INT-path is very suitable for deployment in data center networks thanks to their symmetric network topologies. Tian Pan 0001, Enge Song, Zizheng Bian, Xingchen Lin, Xiaoyu Peng, Jiao Zhang 0002, Tao Huang 0005, Bin Liu 0001, Yunjie Liu 0001 |
INFOCOM | 7 |
| 2019 | RABA: Resource-Aware Backup Allocation For A Chain of Virtual Network FunctionsabstractNetwork Function Virtualization (NFV) turns a sequence of network functions on hardwares into a service chain of virtual network functions (VNFs) provisioned on virtual machines or containers. However, the chain of VNFs may suffer from interruption as long as one VNF fails due to software faults or hardware malfunctions. A common approach to ensuring high availability is to provide backup nodes for primary VNFs. However, existing work on allocating backup nodes have not considered the heterogeneous resource demands of different VNFs. In this paper, we formalize the resource-aware backup allocation problem, which aims to minimize the backup resource consumption while meeting the overall availability demand. To this end, we prove the NP-hardness of this problem and propose the RABA-CDDE algorithm based on differential evolution to solve it. Besides, to reduce the computation overhead of RABA-CDDE, a greedy algorithm is proposed. Our extensive evaluation shows that the proposed algorithms can reduce the resource consumption by about 15% and 35% respectively compared to the state-of-art solutions in dedicated and shared protection scenarios. Jiao Zhang 0002, Chunyi Peng 0001, Linquan Zhang, Tao Huang 0005, Yunjie Liu 0001 |
INFOCOM | 5 |
| 2019 | NB-cache: non-blocking in-network caching for high-speed content routersabstractInformation-Centric Networking (ICN) provides scalable and efficient content distribution at the Internet scale due to its in-network caching and native multicast capabilities. To support these features, a content router needs high performance at its data plane, which consists of three forwarding steps: checking the Content Store (CS), then the Pending Interest Table (PIT), and finally the Forwarding Information Base (FIB). While prior works focus on performance optimization of a single step, we build an analytical model of content router's entire data plane and identify that CS is the actual bottleneck in the pipeline. Compared with PIT and FIB, CS is more challenging because it has more data to read/write, may have more entries in its table to store and lookup, and needs to organize content objects to sustain frequent cache replacement. Then, we propose a novel mechanism called "NB-Cache" to address CS's performance issue from a network-wide point of view rather than a single router's. In NB-Cache, when packets arrive at a router whose CS is fully loaded, instead of being blocked and waiting for the CS, these packets are forwarded to the next-hop router, whose CS may not be fully loaded. This approach essentially utilizes Content Stores of all the routers along the forwarding path in parallel rather than checking each CS sequentially. Our experiments show significant improvement of data plane performance: 70% reduction in round-trip time (RTT) and 130% increase in throughput. Tian Pan 0001, Xingchen Lin, Jiao Zhang 0002, Hao Li 0011, Jianhui Lv, Tao Huang 0005, Bin Liu 0001, Beichuan Zhang 0001 |
IWQoS | 6 |
| 2019 | Virtual Time Machine for Reproducible Network EmulationabstractReproducing network emulation experiments on diverse physical platforms with varying computation and communication resources is non-trivial. Many state-of-the-art network emulation testbeds do not guarantee timing fidelity. Consequently, results obtained from these testbeds can be misleading, especially when insufficient physical resources are provided to run the experiments. Reproducibility is far from being the norm. In this paper, we present a novel approach that can guarantee reproducible results for network emulation. Our system, called the Virtual Time Machine (VTM), takes advantage of both time dilation and carefully controlled scheduling of the virtual machines. Time dilation allows sufficiently scaled resources to run the experiments in virtual time, and controlled VM scheduling prescribes the precise timing of message passing for distributed applications---independent of the resource provisioning of the underlying physical testbed. Preliminary experiments show that VTM can guarantee reproducible results with varying time dilation, resource subscription, and VM scheduling scenarios. Jiang Liu 0010, Tao Huang 0005, Jason Liu 0001 |
SIGSIM-PADS | 3 |
| 2019 | Energy-efficient computation offloading in 5G cellular networks with edge computing and D2D communicationsabstractComputation offloading has been considered as one of the key research issues in edge computing fields. In order to reduce the energy consumption of the mobile terminal, the energy efficiency issue of computation offloading has attracted a lot of attention from academia and industry. In this study, the authors propose an energy‐efficient computation offloading scheme in 5G cellular networks with edge computing and device‐to‐device (D2D) communications. They consider the computation offloading to fog computing devices via D2D communications and mobile edge computing (MEC) servers via cellular networks. And thus the computation task execution model can be composed of local execution, fog computing device execution and MEC server execution. Then, they formulate the computation offloading issue as stochastic optimisation problem, and use the Lyapunov optimisation technology framework to solve this problem. Finally, extensive simulation results are presented to illustrate the effectiveness of the proposed scheme. Qingmin Jia, Renchao Xie, Qinqin Tang, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
IET Commun. | 5 |
| 2019 | Energy-efficient hierarchical cooperative caching optimisation for 5G networksabstractThe caching in fifth generation (5G) networks has been considered as a promising technique to reduce the duplicate traffic transmission and improve the users' quality of experience (QoE). Although many works have been done for caching in 5G networks, most of them focus on the content caching policy design to optimise the users' QoE, the issue to realise the energy efficiency of the whole network is not fully considered. Therefore, in this study, by introducing the hierarchical cooperative caching property, i.e. the core gateway and the base stations can cooperative cache the content, the authors study the problem of the hierarchical cooperative caching policy to realise the energy efficiency. They then formulate the hierarchical cooperative content placement problem as an integer programming problem to minimise the total energy consumption. Also, to reduce the computation complexity, the optimisation algorithm based on the idea of quantum‐inspired evolutionary is proposed, which has the fast convergence and approximate to the optimal solution. Finally, extensive simulation results are illustrated to demonstrate the performance of the proposed scheme. Renchao Xie, Qinqin Tang, Tao Huang 0005 |
IET Commun. | 3 |
| 2019 | Future Internet: trends and challengesabstractTraditional networks face many challenges due to the diversity of applications, such as cloud computing, Internet of Things, and the industrial Internet. Future Internet needs to address these challenges to improve network scalability, security, mobility, and quality of service. In this work, we survey the recently proposed architectures and the emerging technologies that meet these new demands. Some cases for these architectures and technologies are also presented. We propose an integrated framework called the service customized network which combines the strength of current architectures, and discuss some of the open challenges and opportunities for future Internet. We hope that this work can help readers quickly understand the problems and challenges in the current research and serves as a guide and motivation for future network research. Jiao Zhang 0002, Tao Huang 0005, Shuo Wang 0006, Yunjie Liu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2019 | Service-differentiated QoS routing based on ant colony optimisation for named data networking
Rui Hou 0003, Lang Zhang, Yuzhou Chang, Tao Huang 0005, Jiangtao Luo |
Peer-to-Peer Netw. Appl. | 6 |
| 2019 | Service Function Chain Composition, Placement, and Assignment in Data CentersabstractWith the development of network function virtualization (NFV), service function chains (SFCs) are deployed via virtual network functions (VNFs). In general, the SFCs are served via composition and then deployed into data center infrastructures. However, most of the existing works neglect SFC composition. Furthermore, they consider that VNF instances are independently deployed for each SFC, which may underutilize the computational power of servers. We consider, for each required VNF in the chain, the operator can either place it on a new instance or assign it to an established instance if the residual resource of that instance is sufficient. Such a deployment scheme can leverage resources more efficiently and we define it as SFC placement and assignment. In this paper, we first combine SFC composition, placement and assignment together to enhance resource allocation. We present the system model and formulate the problem as 0-1 integer programming. We aim to improve the VNF instance utilization as well as reduce the link consumption. A heuristic approach called Jcap is developed to solve the problem in two stages. The simulations show that Jcap achieves competitive performance with the optimal results obtained from mathematical model. Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2018 | Hierarchical collaborative caching in 5G networksabstractCaching in mobile networks can reduce the redundant data transmission and cope with the challenge of the explosive growth of mobile data traffic. It has been considered as a promising technology in 5G networks and has been attracting a lot of attention in recent years. Although many existing works have addressed the content placement problem or the cache optimisation problem, most of them do not consider the issue of hierarchical collaborative caching. Collaborative caching can further alleviate the traffic pressure and reduce the user‐perceived latency by reducing duplicate content transmission. Therefore, in this study, the authors consider a hierarchical collaborative caching framework with the cache deployment at the distributed gateway and mobile edge computing servers, and then design a novel caching strategy based on this framework. They formulate the hierarchical collaborative content placement problem as an optimisation problem to maximise the latency saving under the constraint of limited cache capacity. Since finding the optimal solution is an NP‐hard problem, they propose a genetic placement algorithm to find the near‐optimal solution to reduce the computation complexity. Numerical experiment results show that the proposed algorithms can significantly improve the performance compared with the reference algorithms. Qinqin Tang, Renchao Xie, Tao Huang 0005, Yunjie Liu 0001 |
IET Commun. | 3 |
| 2018 | A novel forwarding and routing mechanism design in SDN-based NDN architectureabstractCombining named data networking (NDN) and software-defined networking (SDN) has been considered as an important trend and attracted a lot of attention in recent years. Although much work has been carried out on the integration of NDN and SDN, the forwarding mechanism to solve the inherent problems caused by the flooding scheme and discard of interest packets in traditional NDN is not well considered. To fill this gap, by taking advantage of SDN, we design a novel forwarding mechanism in NDN architecture with distributed controllers, where routing decisions are made globally. Then we show how the forwarding mechanism is operated for interest and data packets. In addition, we propose a novel routing algorithm considering quality of service (QoS) applied in the proposed forwarding mechanism and carried out in controllers. We take both resource consumption and network load balancing into consideration and introduce a genetic algorithm (GA) to solve the QoS constrained routing problem using global network information. Simulation results are presented to demonstrate the performance of the proposed routing scheme. Jia Li 0006, Renchao Xie, Tao Huang 0005 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2018 | Joint Resource Allocation for Software-Defined Networking, Caching, and ComputingabstractAlthough some excellent works have been done on networking, caching, and computing, these three important areas have traditionally been addressed separately in the literature. In this paper, we describe the recent advances in jointing networking, caching, and computing and present a novel integrated framework: software-defined networking, caching, and computing (SD-NCC). SD-NCC enables dynamic orchestration of networking, caching, and computing resources to efficiently meet the requirements of different applications and improve the end-to-end system performance. Energy consumption is considered as an important factor when performing resource placement in this paper. Specifically, we study the joint caching, computing, and bandwidth resource allocation for SD-NCC and formulate it as an optimization problem. In addition, to reduce computational complexity and signaling overhead, we propose a distributed algorithm to solve the formulated problem, based on recent advances in alternating direction method of multipliers (ADMM), in which different network nodes only need to solve their own problems without exchange of caching/computing decisions with fast convergence rate. Simulation results show the effectiveness of our proposed framework and ADMM-based algorithm with different system parameters. Qingxia Chen, F. Richard Yu, Tao Huang 0005, Renchao Xie, Jiang Liu 0010, Yunjie Liu 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2018 | Multi-Attributes-Based Coflow Scheduling Without Prior Knowledge
Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2017 | OpenSched: Programmable Packet Queuing and Scheduling for Centralized QoS ControlabstractIn this work, we propose OpenSched, a layered architecture that glues the QoS apps, the controller and the switches together to maximally unleash the power of centralized QoS control. Specifically, our design consists of (1) a flexible northbound interface via the ``builder pattern'', (2) one-to-many controller-switch interactions via device abstraction, thread pooling and Java NIO's selector mechanism, and (3) efficient southbound protocol handling as well as QoS policy execution via a producer-consumer model at the switch side. We build a prototype based on ONOS and OVS with 2340 lines of Java code and 1097 lines of c code. OpenSched is expected to facilitate flexible network resource provisioning. Tian Pan 0001, Tao Huang 0005, Jianwei Mao, Yunjie Liu 0001 |
ANCS | 2 |
| 2017 | Software Defined Networking, Caching and Computing Resource Allocation with Imperfect NSIabstractWe propose a novel framework called Software Defined Networking, Caching and Computing (SD-NCC) which integrates networking, caching and computing in a systematic way to improve the end-to-end system performance. In SDNCC, the more in-network resources it utilizes, the less network usage it costs under the same service demands. However only minimizing the total network usage leads to bottlenecks in the network, making the network fragile to traffic bursts. In this paper, we study the joint networking, caching and computing resource allocation issue and formulate it as an optimization problem to make a trade off between minimizing network usage and balancing servers' load. In addition, taking into consideration the inaccurate measurement of network state information (NSI), we reformulate this problem under imperfect NSI. Because the joint allocation problems with imperfect NSI are large-scale combinational optimization problems, we propose a discrete stochastic approximation(DSA) algorithm to deal with it. Finally, simulations are conducted to demonstrate the effectiveness of proposed framework and algorithms. Simulation results show that SD-NCC can significantly improve the end-to-end performance by sharing the physical infrastructure and information resources. Besides, DSA algorithms can achieve near-optimal performance. Qingxia Chen, Renchao Xie, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
GLOBECOM | 3 |
| 2017 | Joint Forwarding Strategy and Resource Allocation in Information-Centric HWNsabstractNamed Data Networking (NDN) is a prominent fully- fledged Information-Centric Networking (ICN) architecture. NDN can help users to take advantage of multiple access networks in Heterogeneous Wireless Networks (HWNs) more efficiently than IP. In HWNs with NDN, which we call information-centric HWNs, jointly designing forwarding strategy and resource allocation has great potential to improve network performance, which is ignored in the literatures. To fill in this blank, we propose a jointly designed forwarding strategy and resource allocation algorithm called Dynamic Forwarding and Resource Allocation (DFRA) that can adapt variable wireless environment. We also establish the fundamental throughput limitations of information-centric HWNs and prove that DFRA is throughput-optimal. By the cooperation between forwarding strategy and resource allocation, DFRA enables users to utilize wireless communication resource in information-centric HWNs more efficiently. From simulation results, DFRA can provide larger network throughput, faster download speed and better fairness than forwarding strategy that doesn't explicitly cooperate with resource allocation. Renchao Xie, Tao Huang 0005, Ru Huo, Jiang Liu 0010, Yunjie Liu 0001 |
GLOBECOM | 3 |
| 2017 | Modeling CCN Packet Forwarding EngineabstractWith the in-network caching capability embedded, packet forwarding in CCN (content-centric networking) becomes rather sophisticated. To enable wire-speed forwarding, previous works have reported exciting component-level performance achievements. However, a proper system-level model for exact bottleneck identification and accordingly performance tuning is still absent. In this paper, we build two such models dedicated to the two common implementation variants of a CCN router, i.e., the pipeline model and the run- to-completion model, respectively. By carefully investigating and analyzing the interactions between FIB (forwarding information base), PIT (pending interest table) and CS (content store) in the two models, we quantitatively identify that CS is the exact performance bottleneck of the entire CCN packet forwarding engine. This conclusion is very timely (if not too late) because in the past, researchers invest a lot of time and effort in optimizing FIB and PIT while very limited for CS. According to the mathematics, we suggest that the research community should shift more attention to CS performance tuning or totally rethink the entire packet forwarding architecture. Tian Pan 0001, Tao Huang 0005 |
GLOBECOM | 2 |
| 2017 | Energy-Efficient Content Placement for Layered Video Content Delivery over Cellular NetworksabstractWith the ever-increasing demand for high quality video, mobile video transmission optimization over a limited wireless network capacity has attracted extensive attention. Scalable Video Coding (SVC) is a main solution to provide better Quality of Experience (QoE) by encoding each video into one mandatory base layer and several optional enhancement layers. Deployment of caching in wireless networks has been considered as another effective method to mitigate redundant data transmission over backhaul links and to reduce the end-to-end video transmission delay. Although some works have been done for layered video content over cellular networks with caching, most of them focus on video quality selection or video caching to optimize the users' QoE. The problem of energy- efficient content placement is largely ignored. To fill this gap, we focus on the problem of energy- efficient content placement for layered video content delivery over cellular networks in this paper. Our design objective is to maximize the energy cost savings. We formulate the energy- efficient content placement problem as a convex optimization problem. Then, by solving the optimization problem, we can obtain the optimal set of content placement parameters for the Mobile Network Operator (MNO) to design an optimal caching policy for layered video contents. Finally, simulation results are presented to show the performance of the proposed content placement scheme. Junfeng Xie 0002, Renchao Xie, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
GLOBECOM | 3 |
| 2017 | Leveraging multiple coflow attributes for information-agnostic coflow schedulingabstractRecently, designing information-agnostic coflow scheduling mechanisms attracts much attention since by leveraging priority queues, they could reduce coflow completion time in data-parallel clusters without a priori knowledge, such as flow size, coflow size. However, existing information-agnostic mechanisms generally schedule coflows only according to the sent data size of different coflows and ignore other useful coflow-level attributes like width, length and communication patterns. In this paper, we investigate that the coflow completion time could be further decreased by jointly leveraging multiple coflow-level attributes. Based on this investigation, we present a Multiple-attributes-based Coflow Scheduling (MCS) mechanism to reduce the coflow completion time. In MCS, a Shortest and Narrowest Coflow First (SNCF) algorithm is designed to separate coflows based on their widths and estimated lengths at the start of a coflow. During the transmission of coflows, one type of demotion thresholds employed in previous coflow scheduling mechanisms is too crude for various coflows. Therefore, we proposed a double-threshold scheme to adjust the priorities of narrow (small coflow width) and wide (large coflow width) coflows according to different thresholds. Trace-driven simulations with production workloads show that MCS outperforms the previous information-agnostic scheduler Aalo, and reduces the coflow completion time of small coflows. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
ICC | 3 |
| 2017 | Energy-efficient cache resource allocation and QoE optimization for HTTP adaptive bit rate streaming over cellular networksabstractWith the ever-increasing demand for high quality video, mobile video transmission optimization over limited wireless network capacity has been attracted extensive attention. HTTP Adaptive Bit Rate (ABR) streaming is a main solution to provide better Quality of Experience (QoE) by adapting multimedia content over wireless channels real-timely. Deployment of caching in wireless network has been considered as another effective method to mitigate redundant data transmission over backhaul links and to reduce the end-to-end video transmission delay. Although some works have been done for HTTP ABR streaming caching, they only consider the users' QoE. The problem of energy-efficient cache resource allocation is largely ignored. In this paper, we focus on the problem of optimal cache resource allocation for HTTP ABR streaming in cellular networks. Our design objective is to maximize both the users' QoE and energy cost saving. We formulate the content cache management problem as two sub-optimization problems. Then, by solving the two sub-optimization problems, we can obtain the optimal set of playback rates selected by users and the MNO's caching policy for each individual content. Finally, simulation results are presented to show the performance of the proposed cache resource allocation scheme. Junfeng Xie 0002, Renchao Xie, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
ICC | 3 |
| 2017 | Adaptively adjusting ECN marking thresholds for datacenter networksabstractECN thresholds have limited operational range and very strict scope. Lower thresholds exacerbate the queue underflow while higher thresholds increase the queueing delays. In this paper, an Adaptive ECN (A-ECN) marking scheme is proposed to enhance the performance of ECN. A-ECN can adaptively adjust ECN marking thresholds in different scenarios to achieve good generality. Therefore, network operators can directly deploy A-ECN in various environments regardless of underlying queue types and bandwidth. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
ICNP | 3 |
| 2017 | A clustering-based approach for Virtual Network Function Mapping and AssigningabstractNetwork Function Virtualization has attracted attention from both academia and industry as it can help the service provider to obtain agility and flexibility in network service deployment. In general, the enterprises require their flows to pass through a specific sequence of virtual network function (VNF) that varies from service to service. In addition, for each VNF required in the coming service demands, the operator can either launch a new instance for it or assign it to an established instance. This makes the network service deployment tasks even more complicated. In this paper, we first propose a method based on min-K-cut to cluster the VNFs. With clustering results as guidance, we determine whether to launch or reuse the instance to improve utilization rate of the VNF instance. Furthermore, for purpose of decreasing link bandwidth occupation, we aggregate the instances that are deployed with VNFs from the same cluster into the same server or rack. We evaluate our approach considering the average link bandwidth occupied by every accepted demand, the instance utilization rate and the total number of served demands. The simulation shows that our approach reduces link occupation effectively, and, meanwhile, guarantees the VNF instance utilization rate advantageously. Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001 |
IWQoS | 3 |
| 2017 | Skipping congestion-links for coflow schedulingabstractData transfer duration accounts for a great proportion of job completion time in big-data systems. To reduce the time spent on data transfer, some traffic scheduling mechanisms at coflow-level are proposed recently. Most of them abstract datacenter networks as an ideal non-blocking big-switch, and the bottleneck is located at egress or ingress ports of end-hosts instead of in networks. Thus, they mainly focus on how to allocate port capacities of end-hosts to jobs without considering innetwork congestion. However, link congestion frequently occurs in datacenter networks due to network oversubscription and load imbalance. When link congestion occurs, bottleneck locations will move from the ports of end-hosts to network links. In this paper, we design and implement SkipL, a congestionaware coflow scheduler which could detect congestion and schedules coflows at end-hosts to effectively reduce coflow completion time. In addition, to be easily deployed in cloud environments, SkipL does not require to control flow routes. SkipL prototype system is implemented in Linux. The results of experiments conducted in a real small testbed and simulations conducted in the flow-level simulator show that SkipL reduces the average Coflow Completion Time(CCT) compared to the per-flow fair sharing scheduling method and Varys. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
IWQoS | 3 |
| 2017 | Senz: A Context Awareness Middleware System Used in Mobile DevicesabstractWith the continuing penetration of sensor devices and the development of wireless communication techniques, increasing number of applications involving context awareness in ubiquitous computing have been used in daily life. How to collect data from mobile devices at a low energy cost and to mine contextual habits of users remains a key challenge for ubiquitous computing. We proposed an efficient context- awareness computing middleware system, Senz. By leveraging high-efficiency mobile data transmission method, using high-performance context recognition algorithms, and combining mobile data with online third-party data, this middleware system can recognize various user behavior patterns reliably, accurately and efficiently. Experiments show that the Senz recognition accuracy of context activity is above 83% on average and the energy cost is relatively low. By integrating Senz SDK and cloud computing engine, developers can build rich user experience apps with better understanding of users' behavior data and providing various personalized contextual services. Hengyang Zhang, Tao Huang 0005, Yunjie Liu 0001, Shixiang Zhu, Yuanying Chi |
VTC Spring | 2 |
| 2017 | Flow distribution-aware load balancing for the datacenter
Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
Comput. Commun. | 3 |
| 2017 | Efficient caching resource allocation for network slicing in 5G core networkabstractNetwork slicing has been considered as one of the key technologies in the next generation mobile network (fifth generation – 5G), which can create virtual network and provide customised services on demand. Most of the current work on network slicing mainly focuses on virtualisation technology, especially in virtual resource allocation. However, caching as a significant approach to improve the content delivery and quality of experience for end‐users has not been well considered in network slicing. In this study, the authors consider in‐network caching combining with network slicing, and propose an efficient caching resource allocation scheme for network slicing in 5G core network. They first formulate the caching resource allocation issue as an integer linear programming model, and then propose a caching resource allocation scheme based on chemical reaction optimisation (CRO) algorithm, which can significantly improve the caching resource utilisation. The CRO algorithm is a population‐based optimisation metaheuristic, which has advantages in searching optimal solution and computation complexity. Finally, extensive simulation results are presented to illustrate the performance of the proposed scheme. Qingmin Jia, Renchao Xie, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
IET Commun. | 3 |
| 2017 | Jointly optimized congestion control, forwarding strategy, and link scheduling in a named-data multihop wireless networkabstractAs a promising future network architecture, named data networking (NDN) has been widely considered as a very appropriate network protocol for the multihop wireless network (MWN). In named-data MWNs, congestion control is a critical issue. Independent optimization for congestion control may cause severe performance degradation if it can not cooperate well with protocols in other layers. Cross-layer congestion control is a potential method to enhance performance. There have been many cross-layer congestion control mechanisms for MWN with Internet Protocol (IP). However, these cross-layer mechanisms for MWNs with IP are not applicable to named-data MWNs because the communication characteristics of NDN are different from those of IP. In this paper, we study the joint congestion control, forwarding strategy, and link scheduling problem for named-data MWNs. The problem is modeled as a network utility maximization (NUM) problem. Based on the approximate subgradient algorithm, we propose an algorithm called ‘jointly optimized congestion control, forwarding strategy, and link scheduling (JOCFS)’ to solve the NUM problem distributively and iteratively. To the best of our knowledge, our proposal is the first cross-layer congestion control mechanism for named-dataMWNs. By comparison with the existing congestion control mechanism, JOCFS can achieve a better performance in terms of network throughput, fairness, and the pending interest table (PIT) size. Renchao Xie, Tao Huang 0005, Yunjie Liu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2017 | FlowTrace: measuring round-trip time and tracing path in software-defined networking with low communication overheadabstractIn today’s networks, load balancing and priority queues in switches are used to support various quality-of-service (QoS) features and provide preferential treatment to certain types of traffic. Traditionally, network operators use ‘traceroute’ and ‘ping’ to troubleshoot load balancing and QoS problems. However, these tools are not supported by the common OpenFlow-based switches in software-defined networking (SDN). In addition, traceroute and ping have potential problems. Because load balancing mechanisms balance flows to different paths, it is impossible for these tools to send a single type of probe packet to find the forwarding paths of flows and measure latencies. Therefore, tracing flows’ real forwarding paths is needed before measuring their latencies, and path tracing and latency measurement should be jointly considered. To this end, FlowTrace is proposed to find arbitrary flow paths and measure flow latencies in OpenFlow networks. FlowTrace collects all flow entries and calculates flow paths according to the collected flow entries. However, polling flow entries from switches will induce high overhead in the control plane of SDN. Therefore, a passive flow table collecting method with zero control plane overhead is proposed to address this problem. After finding flows’ real forwarding paths, FlowTrace uses a new measurement method to measure the latencies of different flows. Results of experiments conducted in Mininet indicate that FlowTrace can correctly find flow paths and accurately measure the latencies of flows in different priority classes. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001, F. Richard Yu |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2016 | Joint Resource Allocation for Software Defined Networking, Caching and ComputingabstractRecently, there are significant advances in the areas of networking, caching and computing. Nevertheless, these three important areas have traditionally been addressed separately in the existing research. In this paper, we present a novel framework that integrates networking, caching and computing in a systematic way and enables dynamic orchestration of these three resources to improve the end-to-end system performance and meet the requirements of different applications. Then, we consider the bandwidth, caching and computing resource allocation issue and formulate it as a joint caching/computing strategy and servers selection problem to minimize the combination cost of network usage and energy consumption in the framework. To minimize the combination cost of network usage and energy consumption in the framework, we formulate it as a joint caching/computing strategy and servers selection problem. In addition, we solve the joint caching/computing strategy and servers selection problem using an exhaustive-search algorithm. Simulation results show that our proposed framework significantly outperforms the traditional network without in-network caching/computing in terms of network usage and energy consumption. Qingxia Chen, F. Richard Yu, Tao Huang 0005, Renchao Xie, Jiang Liu 0010, Yunjie Liu 0001 |
GLOBECOM | 3 |
| 2016 | Joint user association and rate allocation for HTTP adaptive streaming in heterogeneous cellular networksabstractHypertext transfer protocol based (HTTP) adaptive streaming (HAS) of video over wireless networks has brings huge challenge for the mobile networks. Although some works have been done for video streaming delivery in heterogeneous cellular networks, most of them are focus on the video streaming scheduling or the caching strategy design. The problem of joint user association and rate allocation to maximize the system utility while satisfying the requirement of the quality of experience of users is largely ignored. In this paper, the problem of joint user association and rate allocation for HTTP adaptive streaming in heterogeneous cellular networks is studied, we model the optimization problem as a mixed integer programming problem. To reduce the computational complexity, an optimal rate allocation using the Lagrangian dual method under the assumption of knowing user association for BSs is first solved. Then we use the many-to-one matching model to analyze the user association problem, and the joint user association and rate allocation based on the distributed greedy matching algorithm is proposed. Finally, extensive simulation results are illustrated to demonstrate the performance of the proposed scheme. Renchao Xie, F. Richard Yu, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
ICC | 3 |
| 2016 | FDALB: Flow distribution aware load balancing for datacenter networksabstractWe present FDALB, a flow distribution aware load balancing mechanism aimed at reducing flow collisions and achieving high scalability. FDALB, like the most of centralized methods, uses a centralized controller to get the view of networks and congestion information. However, FDALB classifies flows into short flows and long flows. The paths of short flows and long flows are controlled by distributed switches and the centralized controller respectively. Thus, the controller handles only a small part of flows to achieve high scalability. To further reduce the controller's overhead, FDALB leverages end-hosts to tag long flows, thus switches can easily determine long flows by inspecting the tag. Besides, FDALB can adaptively adjust the threshold at each end-host to keep up with the flow distribution dynamics. Shuo Wang 0006, Jiao Zhang 0002, Tao Huang 0005, Tian Pan 0001, Jiang Liu 0010, Yunjie Liu 0001 |
IWQoS | 3 |
| 2016 | Special issue on future network: software-defined networkingabstractComputer networks have to support an everincreasing array of applications, ranging from cloud computing in datacenters to Internet access for users.In order to meet the various demands, a large number of network devices running different protocols are designed and deployed in networks.As a result, network management and the deployment of new protocols and applications are quite challenging.On one hand, network operators have to manage so many different network devices and manually configure these devices using different tools.On the other hand, vendors use different physical infrastructures as well as software interfaces to manufacture devices, which makes it difficult for researchers to implement new functions in devices.Therefore, network infrastructure and architecture design face great challenges.Software-defined networking (SDN) has been proposed as a new way to facilitate network evolution.SDN decouples the data and control planes, and removes the control plane from network hardware.In SDN, all the devices are controlled by a centralized controller through open protocols, such as OpenFlow, BGP, and NETCONF.Then, control functions are implemented in the centralized controller to realize operational efficiency and reduce costs.Thus, it dramatically simplifies the network, and brings many potential benefits in terms of network management, network virtualization, trouble shooting, and other Tao Huang 0005, F. Richard Yu, Yunjie Liu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2016 | Caching resource sharing in radio access networks: a game theoretic approachabstractDeployment of caching in wireless networks has been considered an effective method to cope with the challenge brought on by the explosive wireless traffic. Although some research has been conducted on caching in cellular networks, most of the previous works have focused on performance optimization for content caching. To the best of our knowledge, the problem of caching resource sharing for multiple service provider servers (SPSs) has been largely ignored. In this paper, by assuming that the caching capability is deployed in the base station of a radio access network, we consider the problem of caching resource sharing for multiple SPSs competing for the caching space. We formulate this problem as an oligopoly market model and use a dynamic non-cooperative game to obtain the optimal amount of caching space needed by the SPSs. In the dynamic game, the SPSs gradually and iteratively adjust their strategies based on their previous strategies and the information given by the base station. Then through rigorous mathematical analysis, the Nash equilibrium and stability condition of the dynamic game are proven. Finally, simulation results are presented to show the performance of the proposed dynamic caching resource allocation scheme. Junfeng Xie 0002, Renchao Xie, Tao Huang 0005, Jiang Liu 0010, F. Richard Yu, Yunjie Liu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2016 | Guaranteeing Delay of Live Virtual Machine Migration by Determining and Provisioning Appropriate BandwidthabstractThe proliferation of cloud services makes virtualization technology more important. One important feature of virtualization is live Virtual Machine (VM) migration. Two main metrics of evaluating a live VM migration mechanism are total migration time and downtime. Most existing literature on live VM migration focus on designing migration mechanisms to shorten the two metrics or making a tradeoff between them. Few of them can be applied to applications with delay requirements, such as a VM backup process that needs to be done in a specific time. This will negatively impact the user experiences and reduce the profit of cloud service providers. Besides, the frequently varied bandwidth required by the widely used pre-copy mechanism is difficult to be provided by current network technologies. In this work, we theoretically analyze how much bandwidth is required to guarantee the total migration time and downtime of a live VM migration, and then propose a novel transport control mechanism to guarantee the computed bandwidth. The experimental results demonstrate that the bandwidth obtained from the proposed reciprocal-based model guarantees the expected total migration time and downtime, and the proposed transport control mechanism ensures that the live VM migration flow obtains the expected bandwidth even if there are background flows. Jiao Zhang 0002, Fengyuan Ren, Ran Shu 0001, Tao Huang 0005, Yunjie Liu 0001 |
IEEE Trans. Computers | 4 |
| 2015 | A distributed energy-efficient algorithm in green Content-Centric NetworksabstractIn Content-Centric Networking (CCN), most existing works do not consider energy savings by turning off network devices in CCN. In this paper, we systematically analyze the energy efficiency problem in CCN by turning off the content routers and network links. We formulate the energy consumption issue as a Mixed Integer Linear Programming (MILP) model, and propose a centralized solution via spanning tree heuristic and a fully distributed consensus optimization algorithm via the alternating direction method of multipliers (ADMM) to solve the problem for CCN. By duplicating flow variables, the energy consumption problem decomposes into node specific subproblems with local variables. These variables are iteratively driven into consensus via the ADMM. Simulation results reveal that the proposed distributed algorithm is amenable to energy-efficient implementation, due to smaller amount of local information exchange at each iteration. Moreover, the proposed algorithm can converge to final status in a significantly smaller number of iterations compared to the method based on dual decomposition. In addition, our algorithm scales better to large networks and it does not require intensive finetuning of the step size. Chao Fang 0001, F. Richard Yu, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
ICC | 3 |
| 2015 | Modeling of miss-probability in content-centric networking
Tao Huang 0005, Chao Fang 0001, F. Richard Yu, Yunjie Liu 0001 |
Sci. China Inf. Sci. | 1 |
| 2015 | An energy-efficient distributed in-network caching scheme for green content-centric networks
Chao Fang 0001, F. Richard Yu, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
Comput. Networks | 3 |
| 2015 | Congestion-aware adaptive forwarding in datacenter networks
Jiao Zhang 0002, Fengyuan Ren, Tao Huang 0005, Yunjie Liu 0001 |
Comput. Commun. | 3 |
| 2015 | Virtual network embedding based on real-time topological attributesabstractAs a great challenge of network virtualization, virtual network embedding/mapping is increasingly important. It aims to successfully and efficiently assign the nodes and links of a virtual network (VN) onto a shared substrate network. The problem has been proved to be NP-hard and some heuristic algorithms have been proposed. However, most of the algorithms use only the local information of a node, such as CPU capacity and bandwidth, to determine how to map a VN, without considering the topological attributes which may pose significant impact on the performance of the embedding. In this paper, a new embedding algorithm is proposed based on real-time topological attributes. The concept of betweenness centrality in graph theory is borrowed to sort the nodes of VNs, and the nodes of the substrate network are sorted according to the correlation properties between the former selected and unselected nodes. In this way, node mapping and link mapping can be well coupled. A simulator is built to evaluate the performance of the proposed virtual network embedding (VNE) algorithm. The results show that the new algorithm significantly increases the revenue/cost (R/C) ratio and acceptance ratio as well as reduces the runtime. Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2015 | Capacity analysis for cognitive heterogeneous networks with ideal/non-ideal sensingabstractDue to irregular deployment of small base stations (SBSs), the interference in cognitive heterogeneous networks (CHNs) becomes even more complex; in particular, the uncertainty of spectrum mobility aggravates the interference context. In this case, how to analyze system capacity to obtain a closed-form expression becomes a crucial problem. In this paper we employ stochastic methods to formulate the capacity of CHNs and achieve a closed-form expression. By using discrete-time Markov chains (DTMCs), the spectrum mobility with respect to the arrival and departure of macro base station (MBS) users is modeled. Then an integral method is proposed to derive the interference based on stochastic geometry (SG). Also, the effect of sensing accuracy on network capacity is discussed by concerning false-alarm and miss-detection events. Simulation results are illustrated to show that the proposed capacity analysis method for CHNs can approximate the conventional sum methods without rigorous requirement for channel station information (CSI). Therefore, it turns out to be a feasible and efficient way to capture the network capacity in CHNs. Tao Huang 0005, Yinglei Teng, Mengting Liu 0006, Jiang Liu 0010 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2014 | A distributed energy consumption optimization algorithm for content-centric networks via dual decompositionabstractDue to the in-network caching capability, Content-Centric Networking (CCN) has emerged as one of the most promising architectures for the diffusion of contents over the Internet. Most existing works on CCN focus on network resource utilization, and the energy efficiency aspect is largely ignored. In this paper, we formulate the energy consumption issue as a Mixed Integer Linear Programming (MILP) problem, and propose a centralized solution via spanning tree heuristic and a fully distributed energy consumption optimization algorithm via dual decomposition (DD) to solve the problem for CCN. The dual decomposition method transforms the centralized energy consumption optimization problem into the router status, link status, and link flow subproblems. Simulation results reveal that the proposed scheme exhibits a fast convergence speed, and achieves superior energy efficiency compared to other widely used schemes in CCN. Chao Fang 0001, F. Richard Yu, Tao Huang 0005, Jiang Liu 0010, Yunjie Liu 0001 |
GLOBECOM | 3 |
| 2013 | A virtual network mapping algorithm based on integer programmingabstractThe virtual network (VN) embedding/mapping problem is recognized as an essential question of network virtualization. The VN embedding problem is a major challenge in this field. Its target is to efficiently map the virtual nodes and virtual links onto the substrate network resources. Previous research focused on designing heuristic-based algorithms or attempting two-stage solutions by solving node mapping in the first stage and link mapping in the second stage. In this study, we propose a new VN embedding algorithm based on integer programming. We build a model of an augmented substrate graph, and formulate the VN embedding problem as an integer program with an objective function and some constraints. A factor of topology-awareness is added to the objective function. The VN embedding problem is solved in one stage. Simulation results show that our algorithm greatly enhances the acceptance ratio, and increases the revenue/cost ( R/C ) ratio and the revenue while decreasing the cost of the VN embedding problem. Jianya Chen, Tao Huang 0005, Yunjie Liu 0001 |
J. Zhejiang Univ. Sci. C | 4 |
| 2011 | A new algorithm based on the proximity principle for the virtual network embedding problemabstractThe virtual network embedding/mapping problem is a core issue of network virtualization. It is concerned mainly with how to map virtual network requests to the substrate network efficiently. There are two steps in this problem: node mapping and link mapping. Current studies mainly focus on developing heuristic algorithms, since both steps are computationally intractable. In this paper, we propose a new algorithm based on the proximity principle, which considers the distance factor besides the capacity factor in the node mapping step. Thus, the two steps of the embedding problem can be better integrated and the substrate network resource can be used more efficiently. Simulation results show that the new algorithm greatly enhances the performance of the revenue/cost ( R / C ) ratio, acceptance ratio, and runtime of the embedding problem. Jiang Liu 0010, Tao Huang 0005, Jianya Chen, Yunjie Liu 0001 |
J. Zhejiang Univ. Sci. C | 2 |
| 2007 | Bit-interleaved space-time-frequency coded modulation with general linear precodingabstractAbstract A general linear precoding space‐time‐frequency bit‐interleaved coded modulation (GLP‐STF‐BICM) is presented. By expanding the dimensions of linear precoding (LP) matrix we achieve the correlation between the adjacent code matrices, and further time diversity is realized, and by further time diversity higher diversity gain is realized which effectively avoids possible burst errors in block fading channels. Also, the Singleton bound, channel capacity and BER performances are theoretically analyzed. Then the optimum criterion of the precoding matrix to maximize both diversity and coding gain is given. Furthermore, to reduce the decoding complexity, we present an efficient sphere iterative decoding (SID) algorithm for our scheme. Finally, we consider the downlink multiuser system model for GLP‐STF‐BICM, the scheduling and capability of the three multi‐access methods TDMA, F/TDMA, and S/F/TDMA are analyzed. The simulation results prove that the performance of the new scheme improved greatly in the frequency‐selective block fading channels. Copyright © 2007 John Wiley & Sons, Ltd. Tao Huang 0005, Chaowei Yuan, Changlu Sun |
Wirel. Commun. Mob. Comput. | 1 |