Tong Zhang 0018

dblp:07/4227-18 · DBLP profile ↗
← Back
63ranked-venue papers
10as first author
41since 2021 · last 2026
0000-0003-2477-7140ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 39 · 6 first-author · 26 since 2021Systems, architecture and hardware · 19 · 4 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CSATC: Communication-Semantic-Aware Transmission Control for Large-Scale MoE Training
abstract
Transmission control is critical for efficient Mixture-of-Experts (MoE) training over RDMA networks in data centers. Existing mechanisms mainly rely on network-level signals and cannot distinguish training flows with different communication semantics. This paper proposes CSATC, a semantic-aware transmission control mechanism that combines static priority initialization with receiver-side dynamic feedback. CSATC uses MoE phase, endpoint hotness, and flow size to initialize priorities, then rewrites DSCP for subsequent packets and refines sending rates based on runtime feedback. NS-3 simulations show reduced phase completion time and tail FCT under hotspot traffic.
Weiju Li, Xiaoxiang Hua, Tong Zhang 0018
APNet3
2026 Enabling Bounded Delay in TSN Using Per-Flow Hierarchical Scheduling
abstract
Time-Sensitive Networking (TSN) enables deterministic transmission of Scheduled Traffic (ST) through mechanisms such as the Time-Aware Shaper (TAS) and Cyclic Queuing and Forwarding (CQF). However, both rely on high-precision global clock synchronization and adopt a per-class queuing paradigm, making them susceptible to clock synchronization errors. Asynchronous Traffic Shaping (ATS) ensures bounded delay without requiring global time synchronization by assigning an eligibility time to each flow during per-flow shaping, yet often exhibits larger delay jitter. In this paper, we propose a novel scheduling paradigm for TSN by introducing per-flow queues and present the Per-Flow Hierarchical Scheduling (PFHS) mechanism. PFHS divides a hyperperiod into equal-length time slots, which are then assigned to different levels. ST flows are also assigned to these levels, and each flow is scheduled to transmit using only the time slots assigned to its level. By partitioning the bandwidth among distinct ST flows’ per-flow queues in this hierarchical manner, PFHS achieves fine-grained bandwidth allocation. We conduct a theoretical worst-case delay analysis of PFHS, demonstrating its capability to ensure deterministic transmission and tolerate bounded clock synchronization errors. We evaluate the performance of PFHS across various scenarios by extensive simulations on the OMNeT++ platform. Experimental results show that PFHS guarantees the end-to-end delay of ST flows. Moreover, the proposed hierarchical division algorithm successfully schedules 93% of 1000 flows.
Tong Zhang 0018, Xiaoqin Feng, Hao Yang 0064, Fengyuan Ren
IEEE Internet Things J.2
2026 A Multiobjective Bayesian Approach for Optimizing E2E Performance of NFV-Based Tactile Internet With Subjective-Objective Evaluation
abstract
In this paper, we study a multi-objective Bayesian approach to optimize end-to-end (E2E) performance for network function virtualization-based Tactile Internet, called NFV-based TI, with joint subjective and objective evaluation. In the considered NFV-based TI system, the aim is to deploy virtual network functions (VNFs) through middleware (e.g., servers, switches, etc.) in providing service function chains (SFCs), establishing bidirectional communication links between tactile user-teleoperator pairs, accelerating the deployment of new services, and facilitating the completion of relevant tactile interaction requests. To meet the needs of an immersive user experience, we explore the joint subjective-objective evaluation of E2E performance characterization tailored for NFV-based TI and formulate a hybrid black-white box optimization problem. Due to the uncertainty of subjective evaluation in tactile interaction feedback (e.g., irreplicable subjective user ratings), E2E performance characterization is inaccurate, and addressing this problem is nontrivial. Particularly, to reduce the E2E delay and improve E2E user satisfaction, a multi-objective Bayesian approach is proposed involving joint wireless resource allocation and SFC scheduling applying to the uplink/downlink bidirectional communication for the NFV-based TI. Simulations evaluate the proposed solution and demonstrate its superiority over its counterparts.
Hao Xiang 0002, Tong Zhang 0018, Lucheng Chen, Changyan Yi
IEEE Internet Things J.2
2026 Enhancing CQF Robustness to Time Synchronization Errors Using Shadow Queues in Time-Sensitive Network
Hao Yang 0064, Tong Zhang 0018, Xiaoqin Feng, Wenxue Wu, Fengyuan Ren
IEEE Internet Things J.2
2026 Mutual Knowledge Distillation and Contrastive Learning between Multi-View Graphs for Cross-Domain Recommendation
abstract
As a powerful tool to alleviate the data sparsity and cold-start problems in traditional recommender systems, cross-domain recommendation hinges on addressing two fundamental questions: how to transfer knowledge and what to transfer. Regarding the two questions, existing methods have limitations, such as restricted domain connections, inadequate representation disentanglement, and insufficient knowledge transfer. To overcome these challenges, we propose a novel model, KDCLM, which integrates sophisticated knowledge distillation and contrastive learning mechanisms within a multi-view graph architecture. The proposed model comprises two views—a local view and a global view—both of which construct multiple graphs based on user–item interactions to establish richer domain connections. Specifically, the local view incorporates two contrastive learning mechanisms: one for aligning domain-invariant representations and another for differentiating domain-specific representations, which jointly achieve effective representation disentanglement. In addition, we employ knowledge distillation between the global heterogeneous user–item interaction graph and the homogeneous user–user and item–item relationship graphs to facilitate sufficient knowledge transfer. Through extensive experiments on real-world cross-domain recommendation tasks, our proposed KDCLM model demonstrates significant improvements over current state-of-the-art methods. We release our source code at https://github.com/fanydan/KDCLM .
Tianzi Zang, Yidan Fan, Juan Li 0011, Tong Zhang 0018, Yanmin Zhu 0006
ACM Trans. Inf. Syst.5
2026 Efficient Headroom Allocation With Two-Level Flow Control for Lossless Datacenter Networks
abstract
In datacenters, lossless network is very attractive as it can achieve ultra-low latency. In commodity Ethernet, lossless forwarding is achieved by hop-by-hop Priority-based Flow Control (PFC). To avoid buffer overflow, PFC-enabled switches need to reserve some buffer asheadroom, absorbing in-flight packets during the delay for backpressure messages to take effect. However, with the growing link speed in production networks, the buffer becomes increasingly insufficient, and the headroom can occupy a considerable fraction of buffer. As a result, the remaining buffer for absorbing normal traffic bursts is significantly squeezed, leading to frequent PFC messages that degrade the network performance. Worse yet, we find that the current static and queue-independent headroom allocation scheme is quite inefficient, resulting in significant buffer wastage. In light of this, we propose Dynamic and Shared Headroom allocation scheme (DSH), which dynamically allocates headroom to congested queues and enables sharing of allocated headroom among different queues. To achieve this, DSH first introduces port-level flow control, which performs flow control at the granularity of individual ports, guaranteeing lossless forwarding with a small fraction of per-port headroom. With this lossless guarantee, the switch is liberated for dynamic headroom adjustment. DSH dynamically allocates per-queue headroom based on the congestion status of each queue. Meanwhile, DSH preserves the queue-level flow control to protect the non-congested queues from being paused by congested queues, ensuring performance isolation on buffer sharing. Extensive experiments show that DSH can reduce the flow completion time by up to ~78.8%.
Danfeng Shan, Jinchao Ma, Yunguang Li, Boxuan Hu, Tong Zhang 0018, Yazhe Tang, Hao Li 0011, Jinyu Wang 0002, Peng Zhang 0011
IEEE Trans. Netw.5
2025 A Transmission Optimization Method with Bounded Gradient Loss Tolerance for Distributed Transformer Training
Xiaoxiang Hua, Tong Zhang 0018
APNet2
2025 Multi-Objective Bayesian Approach for Optimizing Subjective-Objective Performance of Tactile Internet
abstract
This paper proposes a multi-objective Bayesian approach to optimize the end-to-end (E2E) performance of network function virtualization (NFV)-based tactile Internet (TI) by balancing subjective and objective performances. The system aims to deploy virtual network functions (VNFs) via middleware (e.g., servers, switches) to establish service function chains (SFCs), enabling bidirectional communication between tactile users and teleoperators, accelerating service deployment, and facilitating immersive tactile interaction requests (e.g., in the Metaverse). To meet the demands of an immersive user experience, we explore tailored E2E performance metrics and formulate a hybrid black-white box optimization problem. Addressing the uncertainty in subjective feedback (e.g., non-reproducible user ratings), the proposed approach jointly optimizes wireless resource allocation and SFC scheduling for bidirectional uplink/downlink communication in NFV-based TI, reducing E2E delay and enhancing user satisfaction. Simulations demonstrate the superiority of the proposed solution over existing approaches.
Hao Xiang 0002, Tong Zhang 0018, Jiayuan Chen 0001, Changyan Yi
GLOBECOM2
2025 Joint Optimization of Feedback Signal Transmission and Reconstruction in Tactile Internet
abstract
This paper proposes a novel multi-objective Bayesian optimization approach for tactile feedback transmission and reconstruction in network function virtualization-based Tactile Internet (NFV-based TI). For such a system, guaranteeing real-time and precise feedback transmission and reconstruction is crucial for achieving seamless remote interactions. By jointly optimizing virtual network function (VNF) placement, routing, service function chain (SFC) admission control and wireless resource allocation, this work addresses dynamic adjustment strategies under uncertain network conditions. We propose a multi-objective Bayesian approach tailored for hybrid black-white box optimization problems. A hybrid kernel surrogate model along with adaptive sampling strategies are designed to handle the uncertainties in NFV-based TI environments. This enables Pareto-optimal trade-offs between feedback delay and fidelity. Simulations confirm the superiority of our approach over existing solutions.
Hao Xiang 0002, Tong Zhang 0018, Jiayuan Chen 0001, Changyan Yi
GLOBECOM2
2025 Collaborative Edge-Device DNN Inference with Dynamic Model Partitioning, Data Compression and Resource Allocation
abstract
Collaborative edge-device inference is a promising way to empower resource-constrained mobile devices to execute deep neural network (DNN)-based applications with heavy computational workloads. In particular, a DNN model is partitioned into two parts that are executed on the mobile device and the edge server, respectively. However, offline model partitioning methods suffer from poor adaptability to real computing environments, while online methods have the problem of delayed feedback. In addition, model partitioning inevitably incurs large transmission overheads of DNN's intermediate data. To tackle these challenges, we propose a collaborative edge-device inference optimization algorithm (JPCA) with Joint DNN Partitioning, data Compression, and resource Allocation. Our goal is to maximize the inference accuracy of all tasks while satisfying inference latency and energy requirements in a dynamically changing environment. In JPCA, the joint selection of model partition point and data quantization bit-width is first extracted from the original problem and we propose an improved deep reinforcement learning (DRL)-based algorithm to learn joint decisions. Optimal schemes under different bandwidth conditions are recorded, enabling mobile devices to adjust joint decisions in response to significant bandwidth changes. Furthermore, we design a dynamic edge resource allocation algorithm that makes edge resource allocation decisions for tasks arriving in real time, thereby accelerating DNN inference. The results of the testbed experiments affirm the effectiveness of our proposed algorithms in terms of inference accuracy.
Yufan Tang, Tong Zhang 0018, Kun Zhu 0001, Fengyuan Ren
HPCC2
2025 Low-Latency Microsecond Message Scheduling with Global Consistent Priorities in RDMA Networks
abstract
With the increase in computing speed and network bandwidth in data centers, microsecond-level tail latency has become a key metric for internet-based online services. However, the tail latency of messages in data centers is mainly determined by queuing delay, which is usually much larger than pure message transmission time. Existing traffic scheduling mechanisms fail to effectively coordinate end-side and in-network resources, leading to head-of-line (HOL) blocking for microsecond-level messages at both ends and switches, which severely impacts the tail latency. To address this issue, this paper proposes a network-wide Global Priority based Multi-Path message scheduling mechanism GPMP. It assigns the highest priority to microsecond-level messages at both ends and swtiches to ensure such messages can quickly acquire resources and complete quickly. Furthermore, GPMP also optimizes the transmission of low-priority large messages by introducing a multi-path method, which effectively reduces transmission time and improves the bandwidth utilization. Extensive simulation results show that GPMP not only meets the strict latency requirements of microsecond-level messages, but also significantly reduces the completion time of long messages, leading to an overall improvement in bandwidth utilization.
Qiuyu Yu, Tong Zhang 0018, Kun Zhu 0001, Fengyuan Ren, Yufan Tang, Xiaoxiang Hua
IWQoS2
2025 Hybrid Scheduling of Periodic and Burst Inference Tasks in Real-Time Edge Systems
abstract
Deep neural networks (DNNs) have revolutionized multiple generations by harnessing the power of advanced GPUs and extensive datasets. In fields such as audio and video processing, DNN models are deployed on edge servers to achieve millisecond-level latency to meet stringent service-level objectives (SLOs). However, the dynamic and unpredictable nature of real-world applications, characterized by erratic surges in inference requests, poses significant challenges to existing edge systems. To address these challenges, this study presents a novel scheduling algorithm that utilizes deep reinforcement learning to maximize the minimum margins of all GPUs. This strategic approach significantly enhances the system’s capacity to manage unexpected high-demand tasks, thereby improving resilience and overall response capabilities under diverse fluctuating workloads. The experimental results confirm that the proposed algorithm consistently outperforms traditional algorithms, demonstrating superior performance in task completion rate, resource utilization efficiency, robustness, and responsiveness across different load conditions.
Yixuan Han, Tong Zhang 0018, Kun Zhu 0001
SMC2
2025 Dynamic Slot Extension-Based High-Criticality Tasks Scheduling in TSN-Based DMCS
abstract
With the development of Industry 4.0, Distributed Mixed-Criticality Systems (DMCS) have been widely applied to handle various complex tasks in IoT and aerospace fields. To ensure the reliability and end-to-end Quality of Service (QoS) of high-criticality tasks in DMCS, Time-Sensitive Networking (TSN) with inherent determinism can be adopted to provide deterministic low-latency transmission service. However, in modern DMCS, the emergency burst high-criticality tasks and the event-triggered scheduling mechanism used on end systems (ESs) conflict significantly with the time-triggered scheduling mechanism adopted in TSN. In this paper, we modify the standard static scheduling constraints in TSN and introduce new constraints to enforce time-triggered scheduling behavior on ESs based on the original event-triggered scheduling mechanism. Additionally, we propose a Priority-based Dynamic Slot Extension (PDSE) method to handle the emergency high-criticality tasks generated in DMCS. The simulation results show that our proposed time-triggered constraint is compatible with the standard TSN scheduling, enabling time-triggered scheduling of periodic high-criticality tasks on even-triggered ESs. Moreover, the results of end-to-end delays and delay jitters indicate that compared to other methods, PDSE can better schedule emergency high-criticality tasks in DMCS while exerting a lower impact on other high-criticality tasks.
Tong Zhang 0018, Kun Zhu 0001
IEEE Internet Things J.2
2025 Improving Robustness of Time-Aware Shaper in Time-Sensitive Networking
abstract
Deterministic delivery of scheduled traffic (ST) is critical in time-sensitive networking (TSN). The time-aware shaper (TAS) defined by IEEE 802.1Qbv is the enabler to ensure deterministic end-to-end delays of ST flows. However, TAS does not consider the emergency sporadic flows that commonly exist in automotive and industrial control applications. Therefore, TAS cannot resist the interference of emergency sporadic traffic on normal scheduled traffic. In this paper, we improve the TAS’s time slot allocation and queue occupancy rule and propose a combined scheduling strategy composed of dynamic local regulation and static global planning. Dynamic local regulation executes the earliest deadline first (EDF) discipline at switching nodes to adjust ST frames’ transmission dynamically. Static global planning calculates the offset of ST flows at the source and guarantees bounded delay and jitter. Furthermore, we formalize the offset optimizing problem with considering the EDF, delay, and jitter constraints. Finally, we implement our prototype on OMNet++ 6.0.1. Simulation results further demonstrate that the combined strategy has bounded latency and jitter in the presence of emergency sporadic flows. Besides, the comparison experiments show that our design is superior to eTAS in terms of the jitter and end-to-end delay.
Tong Zhang 0018, Xiaoqin Feng, Hao Yang 0064, Fengyuan Ren
IEEE Internet Things J.2
2025 Absorbing Time Synchronization Errors Using Shadow Queues in Time-Sensitive Networking
abstract
Time-sensitive networking (TSN) is widely used in industrial automation and automotive applications due to its ability to provide deterministic transmission. To meet the stringent deterministic requirements of time-sensitive traffic, the mainstream traffic management mechanism time-aware shaper (TAS) relies on time synchronization among network devices. TAS schedules periodic traffic transmission using preallocated time windows. However, in real networks, time-sensitive traffic may not be precisely forwarded as scheduled, thus failing to achieve the expected performance. A fundamental reason is that the statically planned transmission schemes cannot accommodate dynamic time synchronization errors. To address this problem, we propose a novel TSN traffic scheduling strategy called robust TAS (RTAS), which utilizes a dual-queue gating structure to define a new scheduling rule for absorbing time synchronization errors to guarantee the real-time performance of scheduled traffic (ST). We conducted extensive simulations on OMNeT++ to study the performance of RTAS in industrial automation scenarios. The results show that RTAS enables deterministic transmission with imperfectly synchronized clocks and minimizes the impact of time synchronization errors on ST.
Hao Yang 0064, Tong Zhang 0018, Xiaoqin Feng, Fengyuan Ren
IEEE Internet Things J.2
2025 Dynamic Per-Flow Queues in Shared Buffer TSN Switches
abstract
Time-Sensitive Networking (TSN), as an enhancement based on Ethernet, can ensure deterministic traffic transmission with low delays and minimal jitters. However, TSN switches have only eight priority queues inherited from Ethernet at each egress port, which limits the flexibility and efficiency of traffic scheduling, as well as the support for developing traffic management mechanisms. Although per-flow queues boost scheduling and Quality of Service (QoS), static per-flow hardware queues in switches are considered unpractical due to resource limits. In this article, we leverage the limitation of buffer size on the number of concurrent flows in shared buffer TSN switches to design Dynamic Per-Flow Queues (DFQ). DFQ only maintains a fixed number of virtual queues determined by the buffer size and dynamically manages the mapping between virtual queues and active flows to provide the capability of per-flow queuing. By constructing Flow Mapping Table (FMT) with content-addressable memory (or hash bucket), DFQ can quickly match, create, and recycle queues to multiplex limited switch resource. We prototype DFQ on an FPGA switch and evaluate its performance in different scenarios. Experimental results show that DFQ can decrease the overhead of per-flow isolation with minimal impact on delay and throughput, indicating that DFQ is an effective per-flow queues solution.
Wenxue Wu, Tong Zhang 0018, Xiaoqin Feng, Fengyuan Ren
ACM Trans. Design Autom. Electr. Syst.2
2025 Fault-Tolerant Cyclic Queuing and Forwarding with Fast ACK in Time-Sensitive Networking
abstract
TSN is widely used in industrial automation networks because it can provide deterministic transmission services for critical data. Cyclic Queuing and Forwarding (CQF) is used to shape critical data. However, unexpected data errors may occur due to transient failures like electromagnetic interference. IEEE 802.1CB provides a solution to tolerate such failures by transmitting multiple replicas of data over disjoint paths. However, this solution introduces network resources wastage. Compared to redundant transmission, retransmission can reduce resource waste, but may violate the determinism in TSN. To address this issue, we propose a fault-tolerant mechanism for CQF that supports retransmission, called fault-tolerant CQF (FT-CQF). FT-CQF adopts the Go-Back-N concept to resist failure. Therefore, it does not violate the original transmission sequence of frames. On the basis of standard CQF, FT-CQF occupies an additional queue to cache replicas of Time-Trigger (TT) flows and reserves time slots to forward them. FT-CQF will forward or remove these replicas based on the ACK information. Non-TT flows can use this time slot to transmit when replicas are removed. We implemented FT-CQF on OMNeT++ and verified the performance of FT-CQF. Simulation experiments show that FT-CQF is effective in terms of reliability, bandwidth consumption, and delay.
Tong Zhang 0018, Xiaoqin Feng, Hao Yang 0064, Fengyuan Ren
ACM Trans. Design Autom. Electr. Syst.2
2024 Enabling Low Latency for ECQF based Flow Aggregation Scheduling in Time-Sensitive Networking
abstract
Cyclic Queuing and Forwarding (CQF) configures the same cycle length on the flow path, resulting in certain flows unschedulable. Enhanced CQF (ECQF) based flow aggregation utilizes variable cycle length to address this issue. However, it remains a conceptual model without a concrete implementation. In this paper, we propose a jointly optimize aggregation cycle and flows' offsets (JACO) mechanism to achieve ECQF-based flow aggregation. We also design an incremental heuristic algorithm for JACO. Finally, we evaluate the performance of JACO in different scenarios using OMNet++ simulation platform. Compared with ECQF, the results show that JACO reduces latency and improves resource utilization.
Tong Zhang 0018, Xiaoqin Feng, Fengyuan Ren
DAC2
2024 Dynamic Per-Flow Queues for TSN Switches
abstract
Dynamic Per-Flow Queues (DFQ) extend queues from per-class to per-flow in Time-Sensitive Networking (TSN) switches that overcome large resource consumption by dynamically mapping a fixed number of physical queues to active flows. It can implement per-flow queuing with much less on-chip resource. Compared to brute-force hardware queues, DFQ prototyped on an FPGA, can effectively manage more per-flow queues, allowing for improved priority scheduling with minimal throughput and latency impact.
Wenxue Wu, Tong Zhang 0018, Xiaoqin Feng, Xuelong Qi, Fengyuan Ren
DATE3
2024 Fault- Tolerant Cyclic Queuing and Forwarding in Time-Sensitive Networking
abstract
Time-sensitive networking (TSN) provides determin-istic time-sensitive transmission services for critical data at the link layer. Cyclic Queuing and Forwarding (CQF) defined by IEEE 802.1Qch is used for critical data transmission. However, unexpected data errors may occur due to transient faults like electromagnetic interference. At present, the solution to such faults defined in the IEEE TSN standards is to transmit multiple data copies on redundant paths, which introduces network resources wastage. Compared to redundant transmission, retransmission can reduce resource waste, but may violate the deterministic transmission guarantee in TSN. To tackle with this issue, we propose a time-redundant fault-tolerant mechanism for CQF, called fault-tolerant CQF (FT-CQF). On the basis of standard CQF, FT-CQF occupies an additional queue to cache copies of Time- Trigger (TT) flows and reserves time slots to forward them. According to the returned CRC-related messages, FT-CQF will decide whether to forward these copies. Non- Ttflows can also be transmitted during this time when copies are not required to be forwarded. We implement FT-CQF in OMNeT++, and verify the performance of FT-CQF in typical network scenarios. The extensive simulation experiments show that FT-CQF is effective in terms of fault-tolerant effects, consumed resources, delay, and jitter.
Tong Zhang 0018, Wenxue Wu, Xiaoqin Feng, Guoxi Lin, Fengyuan Ren
DATE2
2024 A Joint Multi-Dimensional Fine-Grained Pruning Method for Deep Neural Network
abstract
Existing deep neural network (DNN) pruning methods can be classified into two main categories: structured pruning and weight pruning. Structured pruning is a representative model compression technology of DNN to reduce the storage and computation requirements and accelerate inference, which mainly includes filter pruning and channel pruning. However, they both belong to coarse-grained methods, which can only decide whether to prune a whole filter or channel or not and provide limited decision space. On the other hand, structured stripe-wise pruning has finer granularity than filter pruning, and shape-wise pruning also has finer granularity than channel pruning. These two fine-grained methods are related to two dimensions: rows and columns from the general matrix multiplication (GEMM) perspective of convolution operations. Considering that combining pruning decisions in finer granularity from multiple dimensions will produce a larger solution space, in this paper we propose a joint multi-dimensional fine-grained pruning scheme (JFP) for DNN compression, which simultaneously prune elements in filters and channels. Extensive experiments on the CIFAR-10 dataset demonstrate that: (1) JFP achieves stabler pruning ratios compared to stripe-wise pruning (2) JFP effectively compresses DNN parameters and reduces calculation amount while maintaining the accuracy compared with counterparts.
Tong Zhang 0018, Kun Zhu 0001
SMC2
2024 Consistent Low-Latency Scheduling for Microsecond-Scale Tasks in Data Centers
Qiuyu Yu, Tong Zhang 0018, Changyan Yi
WASA (3)2
2024 Energy-Efficient UAV Swarm Assisted MEC With Dynamic Clustering and Scheduling
abstract
In this paper, the energy-efficient unmanned aerial vehicle (UAV) swarm assisted mobile edge computing (MEC) with dynamic clustering and scheduling is studied. In the considered system model, UAVs are divided into multiple swarms, with each swarm consisting of a leader UAV and several follower UAVs to provide computing services to end-users. Unlike existing work, we allow UAVs to dynamically cluster into different swarms, i.e., each follower UAV can change its leader based on the time-varying spatial positions, updated application placement, etc. in a dynamic manner. Meanwhile, UAVs are required to dynamically schedule their energy replenishment, application placement, trajectory planning and task delegation. With the aim of maximizing the long-term energy efficiency of the UAV swarm assisted MEC system, a joint optimization problem of dynamic clustering and scheduling is formulated. Taking into account the underlying cooperation and competition among intelligent UAVs, we further reformulate this optimization problem as a combination of a series of strongly coupled multi-agent stochastic games, and then propose a novel reinforcement learning-based UAV swarm dynamic coordination (RLDC) algorithm for obtaining the equilibrium. Simulations are conducted to evaluate the performance of the RLDC algorithm and demonstrate its superiority over counterparts.
Jialiuyuan Li, Jiayuan Chen 0001, Changyan Yi, Tong Zhang 0018, Kun Zhu 0001, Jun Cai 0001
WCNC4
2024 Switch-Assistant Loss Recovery for RDMA Transport Control
abstract
RoCEv2 (RDMA over Converged Ethernet version 2) is the canonical method for deploying RDMA in Ethernet-based datacenters. Traditionally, RoCEv2 runs over the lossless network which is in turn achieved by enabling Priority Flow Control (PFC) within the network. However, as the scale of the datacenter increases, PFC’s side effects, such as head-of-line blocking, congestion spreading, and pause frame storm, are amplified. Datacenter operators can no longer tolerate these problems. In hence, they are seeking PFC alternatives for RDMA networks. Rather than aiming at the lossless RDMA network, we instead handle packet loss effectively to support RDMA over Ethernet. In this paper, we propose Switch-assistant Loss Recovery (SLR), a switch building block to enhance RoCEv2’s loss recovery. Specifically, SLR-enabled switches send loss notifications to request fast retransmissions. To cooperate with go-back-N retransmission, SLR generates loss notifications only when expected packets (i.e., in-order packets expected by receivers) are dropped and then filters out unexpected packets, which can avoid timeouts and prevent exacerbating congestion. Further, we adapt SLR to multi-bottleneck scenarios by inferring expected packets among multiple switch views. We implement SLR prototypes on commodity programmable switches. Evaluations show that SLR reduces the 99.9th-percentile FCT slowdown by up to 21.6$\times$compared to PFC and other state-of-the-arts.
Qingkai Meng 0001, Shan Zhang 0001, Zhiyuan Wang 0004, Tong Zhang 0018, Hongbin Luo, Fengyuan Ren
IEEE/ACM Trans. Netw.5
2023 Collaborative Caching and Scheduling for Live Streaming in Mobile Edge Computing
abstract
With the vigorous increase in live video traffic today, more and more live users require low-latency high-quality video streaming. To this end, there have been many Adaptive Bitrate Streaming (ABR) algorithms to adapt video bitrate to network conditions, and most of them are implemented in the client side. Such algorithms typically can only optimize the quality of experience (QoE) for a single user, but are agnostic to the comprehensive video streaming performance of multiple users. The client-based ABR algorithms also cannot provide sufficient utilization of network resources due to the lack of multi-user perspective. Mobile edge computing (MEC) can achieve lower response latency and can obtain network states of multiple users at edge servers, which is a most applicable technology for mobile live streaming. In this paper, we propose a collaborative caching and scheduling (CCS) mechanism for live streaming services in the MEC environment, aiming to improve the overall viewing QoE for multiple users. CCS provides integrated segment scheduling and allocation of bandwidth and cache resources in the edge network to improve the utilization of resources. At the same time, CCS further explores the larger optimization space provided by scalable video coding (SVC) for enhancing the quality of caching and scheduling solutions. According to our simulation results, CCS can provide a better comprehensive QoE for users compared with counterparts.
Tong Zhang 0018, Kun Zhu 0001
ICC2
2023 A Vacation Queue Based Optimization for Dynamic Application Placement in Edge Computing
abstract
This paper studies the dynamic application placement for edge computing with a variety of random task arrivals (or task offloading requests). Since the storage capacity of the edge server is inherently limited, for better serving mobile devices with heterogeneous computation demands, the edge server is required to update its application placement, leading to a potential energy-latency paradox (i.e., frequent updates may introduce a high energy consumption while infrequent updates may result in the growth of latency). To this end, we propose a novel application placement policy, consisting of a response threshold and a waiting duration for installing and uninstalling the application for each type of task, respectively. Furthermore, we leverage the vacation queue for analyzing the performance of such system, and formulate a joint optimization problem for deriving the optimal configurations of application placement, along with the computation resource allocations, in minimizing the average service latency with a desired energy constraint. A branch-and-bound method integrating an inner convex approximation approach is proposed, and then evaluated with numerical simulations which demonstrate its superiority over counterparts.
Shanfei Shang, Changyan Yi, Tong Zhang 0018, Jun Cai 0001
ICC3
2023 Less is More: Dynamic and Shared Headroom Allocation in PFC-Enabled Datacenter Networks
abstract
In datacenters, lossless network is very attractive as it can achieve ultra-low latency. In commodity Ethernet, lossless forwarding is achieved by hop-by-hop Priority-based Flow Control (PFC). To avoid buffer overflow, PFC-enabled switches need to reserve some buffer as headroom, which is for absorbing in-flight packets during the delay for backpressure messages to take effect. However, with the growing link speed in production networks, the buffer becomes increasingly insufficient, and the headroom can occupy a considerable fraction of buffer. As a result, the remaining buffer for absorbing normal traffic bursts is significantly squeezed, leading to frequent PFC messages that degrade the network performance. However, the current static and queue-independent headroom allocation scheme is inherently inefficient in solving this problem. In light of this, we propose Dynamic and Shared Headroom allocation scheme (DSH), which dynamically allocates headroom to congested queues and enables the allocated headroom to be shared among different queues. By statistical multiplexing, DSH needs much less headroom to ensure lossless forwarding. Furthermore, DSH can be implemented on switching chips with moderate modifications. Extensive simulations show that DSH can absorb 4× more bursts without triggering PFC messages and reduce the flow completion time by up to ~31%.
Danfeng Shan, Tong Zhang 0018, Yazhe Tang, Hao Li 0011, Peng Zhang 0011
ICDCS3
2023 Hierarchical Scheduling of Hybrid DNN Tasks in Embedded Real-Time Systems
abstract
With the widespread application of deep learning (DL) technology in the modern Internet of Things (IoT) areas such as autonomous driving, smart cities and homes, embedded real-time systems are increasingly used at the edge of the network to complete various hybrid DNN tasks. Although embedded real-time systems are equipped with heterogeneous CPU and GPU cores to reduce the response time of inference jobs, the computing resources of heterogeneous devices are not fully utilized, and there is still plenty of room for schedulability to be improved. In this paper, we propose a layer-based hybrid deep neural network (DNN) tasks scheduling algorithm in embedded real-time systems (LHTS) that maps DNN layers to CPU and GPU devices and regulates their start time to avoid confliction. We evaluate LHTS through extensive simulations. The experimental results show that LHTS can achieve more sufficient use of heterogeneous CPU and GPU resources in embedded real-time systems, reduce the worst-case execution time and enhance the schedulability performance of hybrid DNN tasks.
Jiaxin Feng, Kun Zhu 0001, Tong Zhang 0018
ICPADS3
2023 Effective Coflow Scheduling in Hybrid Circuit and Packet Switching Networks
abstract
Hybrid circuit and packet switching networks combine optical circuit switching with electrical packet switching technologies. It can provide higher bandwidth at a lower cost than pure optical or electrical networks, which can meet different performance goals. Coflow is a superior traffic abstraction that captures applications' networking semantics. To improve the transmission performance at the application level, we focus on reducing coflow completion time (CCT). Unlike the problem of minimizing CCT in traditional networks, hybrid network provides larger coflow scheduling space about through which network the internal flows are transmitted, in addition to priorities and bandwidth allocations of coflows. To solve this problem, this paper proposes an Online Network State based coflow scheduling algorithm (ONS) in hybrid switching networks, which minimizes CCT by combining coflow priorities and internal flow path planning, fully considering different switches' characteristic. Extensive simulations prove that ONS has smaller CCT and higher throughput under different load levels.
Renjie Jiang, Tong Zhang 0018, Changyan Yi
ISCC2
2023 Improving Traffic Scheduling Based on Per-flow Virtual Queues in Time-Sensitive Networking
abstract
Time-sensitive networking (TSN) is a set of stan-dards designed to enhance reliable and real-time transmission for ensuring Quality of Service (QoS) for time-critical applications. The switching architecture is a fundamental component of TSN switches to ensure different requirements. However, in the existing TSN switching architecture, each port has up to 8 queues. Limited number of queues may make traffic scheduling more complex and constrained, degrading scheduling performance. In this paper, we propose a traffic scheduling method based on the dynamic hashing per-flow virtual queue (HPFS) in TSN. On this basis, we develop a new traffic scheduling strategy that fully utilizes the per-flow virtual queue. Extensive simulations show that HPFS is effective for schedulability enhancement and latency reduction while providing more concise scheduling,
Haokai Jing, Tong Zhang 0018
ISCC2
2023 Workload Re-Allocation for Edge Computing With Server Collaboration: A Cooperative Queueing Game Approach
abstract
In this paper, a long-term workload management problem for multi-server edge computing with server collaboration is studied. In the considered model, mobile users’ computation-intensive tasks are generated dynamically over the time and offloaded to associated edge servers according to pre-determined subscription agreements. Upon receiving the subscribed workload, each edge server can then decide to whether participate in server collaboration for enabling workload re-allocation (i.e., workload exchange) with other heterogeneously configured edge servers. Unlike most of the existing work, this paper takes into account both competitions and collaborations among strategic edge servers in sharing their computing capacities. To achieve the equilibrium for each edge server in minimizing its expected cost (including energy consumption, delay, transmission, configuration and pricing costs), a joint optimization is formulated for determining i) its amount of workload to undertake, ii) compensation price charged from peers, and iii) computing speed to adopt. To efficiently solve this problem, we propose a novel cooperative queueing game approach, which integrates a convex optimization, a core cost sharing scheme and a mapping rule. Theoretical analyses and extensive simulations are conducted to evaluate the performance of the proposed solution, and demonstrate its superiority over counterparts.
Changyan Yi, Jun Cai 0001, Tong Zhang 0018, Kun Zhu 0001, Bing Chen 0002, Qiang Wu 0018
IEEE Trans. Mob. Comput.3
2023 Towards Impact of Chunk-Level Characteristics on Mobile Live Streaming Performance
abstract
Today, mobile live streaming is gaining a rapid growth in use, which refers to watching the media content recorded and broadcast in real time on mobile devices. In live streaming process, each video segment must go through recording, encoding, uploading, transcoding, publishing, downloading, decoding before playback. The ingest algorithm inside the streamer decides the upload bitrate, while the adaptive bitrate (ABR) algorithm inside the player determines the download bitrate. Thanks to the chunked CMAF standard, each segment is split into smaller chunks that can be independently encoded, transferred, decoded and played. It is of great help to quantify the impact of chunk-level characteristics on mobile live streaming performance. In this paper, we establish a tandem queuing model to describe the whole streaming system. Based on the model, we respectively characterize rebuffering probability, rebuffering count, streaming latency, and average bitrate, analyzing the impact of chunk upload and download rates, upload and download time variances, startup threshold and chunk length on them. From analysis results, we propose insights and recommendations for bitrate adaptation in mobile live streaming and design simple heuristic ingest and ABR algorithms leveraging them. Extensive simulations verify the insights as well as effectiveness of designed algorithms.
Tong Zhang 0018, Zhewei Tang, Jiakun Bao, Fengyuan Ren
IEEE Trans. Mob. Comput.1
2023 Consistent Low Latency Scheduler for Distributed Key-Value Stores
abstract
Nowadays, the distributed key-value stores have become the basic building block for large-scale cloud applications. In large-scale distributed key-value stores, many key-value access operations, which will be processed in parallel on different servers, are usually generated for a single end-user request. Accordingly, the completion time of an end-user request is determined by the last completed key-value access operation. Scheduling the order of serving key-value access operations can effectively reduce the completion times of end requests, thereby improving the user experience. However, existing scheduling algorithms hardly achieve consistent low latency due to the following challenges: the large overhead of cooperating clients and servers, the time-varying load and performance of servers, the traffic distribution can be either heavy-tailed or light-tailed and both the mean and the tail completion time are expected to be low. In this paper, we formalize the problem of scheduling key-value access operations and show it is NP-hard. Furthermore, we heuristically design the distributed adaptive scheduler (DAS), which distributively combines the largest remaining processing time last and the shortest remaining process time first algorithms. Theoretical analysis shows that DAS is adaptive to the time-varying traffic and server performance and can achieve consistent low mean and tail latency regardless of traffic distributions. Extensive simulations show that DAS reduces the mean request completion time by$17 \! \sim \! 50\%$with heavy-tailed traffic and$2 \! \sim 26 \! \%$with light-tailed traffic, while keeping the smallest tail completion time, compared to the default first come first served algorithm. Moreover, DAS outperforms the existing Rein-SBF algorithm under various scenarios.
Wanchun Jiang, Haoyang Li 0006, Yulong Yan, Fa Ji, Jiawei Huang 0001, Jianxin Wang 0001, Tong Zhang 0018
IEEE Trans. Parallel Distributed Syst.7
2022 ORSM: Online Routing and Scheduling Mechanism for Mix-flows in Data Center Networks
abstract
Nowadays, diverse cloud services generate a mix of deadline and non-deadline data flows that constitute the mix-flow environment in data center networks. Deadline flows are mainly generated by user-interactive services and must be completed within deadlines, while non-deadline flows are usually produced by data-parallel services and desire a shorter completion time. To satisfy their performance requirements simultaneously, most existing mix-flow scheduling algorithms are dedicated to balancing the resource contention between these two types of flows. However, routing should also be considered in conjunction with scheduling because inappropriate routing tends to prevent scheduling from playing a valid role. In this paper, we propose ORSM, an online routing and scheduling mechanism for mix-flow transport. We formalize the mix-flow routing and scheduling problem as a mixed integer programming problem. ORSM then solves this problem in a distributed manner by leveraging the intrinsic feedback mechanism to determine the path and transmission rate for each flow. Extensive simulation results show that ORSM can effectively reduce the deadline miss ratio of deadline flows by up to 35.7% as well as the average flow completion time of non-deadline flows by up to 38.6% compared to state-of-art mechanisms.
Zhewei Tang, Tong Zhang 0018, Kun Zhu 0001
ICCCN2
2022 Demystifying and Mitigating TCP Capping
abstract
Today’s Internet user experience greatly depends on some user-perceived network metrics, such as throughput and latency. To improve these metrics, many Internet content providers build the content delivery network (CDN) to provide their services. Generally, CDNs adopt TCP as their transport protocol. A recent line of work improves TCP by proposing novel congestion control algorithms. However, we measure TCP performance in the production CDN and identify an interesting phenomenon termed TCP capping. When the flows experience TCP capping, the fixed-size receive window (rwnd) restricts these flows from fully utilizing network bandwidth. Through in-depth analysis, we demystify that the root cause of TCP capping is an inappropriate constraint on rwnd due to not considering the receiver’s processing capability. To mitigate it, this paper proposes a server-side scheme Apollo and a client-side scheme Artemis for Internet content providers and users, respectively. Apollo probes the receiver’s processing capability and assists the sender in packet sending. And Artemis adjusts the receive buffer in light of the receiver’s processing capability. In our evaluation, compared to vanilla TCP, TCP (w/ Apollo) and TCP (w/ Artemis) shorten flow completion time by up to 91.8% and 94.9%, respectively.
Qingkai Meng 0001, Fengyuan Ren, Tong Zhang 0018, Danfeng Shan, Yajun Yang
IWQoS3
2022 Reliability-Aware Comprehensive Routing and Scheduling in Time-Sensitive Networking
Tong Zhang 0018, Changyan Yi
WASA (2)2
2022 A Time Utility Function Driven Scheduling Scheme for Managing Mixed-Criticality Traffic in TSN
Jinxin Yu, Changyan Yi, Tong Zhang 0018, Jun Cai 0001
WASA (3)3
2022 Multi-Connection Based Scalable Video Streaming in UDNs: A Multi-Agent Multi-Armed Bandit Approach
abstract
Scalable video coding (SVC) has received much attention for video transmission over wireless due to its flexibility. However, most previous work only considered SVC video streaming from a single base station (BS). At present, the densification of BSs enables a user equipment (UE) to connect to multiple BSs in ultra-dense networks (UDNs). In this paper, we consider the problem of SVC video streaming in a UDN, which allows different layers of a video block to be downloaded from different BSs. An optimization problem is formulated aiming to maximize the quality of experience (QoE) of users by selecting the optimal connection strategy and optimal number of video layers. Considering the complexity, to efficiently solve the problem in a distributed manner, the problem of choosing connection strategy is formulated as a multi-agent multi-armed bandit (MA-MAB) problem with only few information exchange. Each user can adapt its connection strategy in a distributed self-learning system. To obtain the optimal arm for the MA-MAB problem, we propose a multi-user arm decision algorithm. To avoid large computation and handover costs, we adopt the same connection strategy for the entire video sequence. Then for each video block, with the given connection strategy, the number of video layers is adjusted adaptively according to dynamic network conditions. Finally, based on the above designs, we provide the SVC-based video downloading scheme to obtain an approximate optimal solution to the original optimization problem. Extensive simulations and comparisons show the feasibility and superiority of the proposed scheme.
Kun Zhu 0001, Lujiu Li, Yuanyuan Xu 0001, Tong Zhang 0018, Lu Zhou 0002
IEEE Trans. Wirel. Commun.4
2021 Cutting the Request Completion Time in Key-value Stores with Distributed Adaptive Scheduler
abstract
Nowadays, the distributed key-value stores have become the basic building block for large scale cloud applications. In large-scale distributed key-value stores, many key-value access operations, which will be processed in parallel on different servers, are usually generated for the data required by a single end-user request. Hence, the completion time of the end request is determined by the last completed key-value access operation. Accordingly, scheduling the order of key-value access operations of different end requests can effectively reduce their completion time, improving the user experience. However, existing algorithms are either hard to employ in distributed key-value stores due to the relatively large cooperation overhead for centralized information or unable to adapt to the time-varying load and server performance under different traffic patterns. In this paper, we first formalize the scheduling problem for small mean request completion time. As a step further, because of the NP-hardness of this problem, we heuristically design the distributed adaptive scheduler (DAS) for distributed key-value stores. DAS reduces the average request completion time by a distributed combination of the largest remaining processing time last and shortest remaining process time first algorithms. Moreover, DAS is adaptive to the time-varying server load and performance. Extensive simulations show that DAS reduces the mean request completion time by more than 15 ~ 50% compared to the default first come first served algorithm and outperforms the existing Rein-SBF algorithm under various scenarios.
Wanchun Jiang, Haoyang Li 0006, Yulong Yan, Fa Ji, Jianxin Wang 0001, Tong Zhang 0018
ICDCS7
2021 Optimizing the Response Time of Memcached Systems via Model and Quantitative Analysis
abstract
Memcached is a widely used in-memory caching solution in large-scale searching scenarios. The most crucial metric of Memcached systems is the response time, which is affected by various factors such as workload, service rate, unbalanced load distribution, and cache miss ratio. This article aims to quantify the influence of each factor on the response time of Memcached systems. First, we establish a theoretical model for Memcached systems that captures their main features, including burst and concurrent key arrival, unbalanced load distribution, and cache miss process. By solving this model using queuing and stochastic theories, we obtain an estimate of the response time in Memcached systems. Intensive experiments based on real-world components demonstrate that the estimate always matches perfectly with the actual value. Furthermore, we obtain a comprehensive and quantitative understanding of all factors. The main insights are threefold. 1) There exists an optimum range of utilization at Memcached servers in which the response time is kept at a low level with a small penalty. 2) The influence of the cache miss ratio on the response time is logarithmic rather than linear. 3) The number of keys generated from an end-user request has the greatest impact in Memcached systems.
Wenxue Cheng, Fengyuan Ren, Wanchun Jiang, Tong Zhang 0018
IEEE Trans. Computers4
2021 Minimizing Coflow Completion Time in Optical Circuit Switched Networks
abstract
Nowadays, optical circuit switching is becoming an increasingly favored technology in scaling data center networks for its definitive advantages in data rate, power consumption, and device cost. Concurrently, reducing coflow completion time (CCT) is of great significance for improving application-level performance. However, minimizing CCT in circuit switched networks is totally different from that in traditional packet switched networks due to port constraints and circuit reconfiguration delays. To address this issue, this article proposes Grouped Optimization-based Scheduling (GOS), a CCT minimization algorithm for circuit switched networks integrating circuit and coflow scheduling. We first formalize the CCT minimization problem into a 0-1 programming problem, then relax and solve the problem in 2 steps to obtain the coflow order and flow grouping decisions on each circuit. Thus intra-group reconfiguration delays are saved, and small coflows can be prioritized at the group level. Theoretical analysis proves GOS is a 4-approximation algorithm in average CCT. To reduce computing overheads, we further propose a heuristic approximation algorithm. Extensive simulations show that the heuristic algorithm has satisfactory CCT performance (0.12× Varys, 0.36× Sunflow) as well as high throughput (16.74× Varys, 1.32× Sunflow), and well adapts to a wide range of reconfiguration delays and algorithm decision time.
Tong Zhang 0018, Fengyuan Ren, Jiakun Bao, Ran Shu 0001, Wenxue Cheng
IEEE Trans. Parallel Distributed Syst.1
2020 One Rein to Rule Them All: A Framework for Datacenter-to-User Congestion Control
abstract
Today, considerable Internet traffic is sent from datacenter and heads for users. The network characteristics of connections served by servers in datacenters are usually diverse. As a result, a specific congestion control algorithm hardly accommodates the heterogeneity and performs well in various scenarios. In this work, we present Rein — a novel framework for Internet congestion control. With Rein, diverse congestion control algorithms can be assigned purposely to connections in one server to adapt to heterogeneity. We design and implement Rein in Linux, and the experiments validate that Rein is capable of smoothly switching among various candidate algorithms on the fly to achieve potential performance gain. Meanwhile, the overheads introduced by Rein are moderate and acceptable.
Danfeng Shan, Xiaohui Luo, Tong Zhang 0018, Yajun Yang, Fengyuan Ren
APNet4
2020 Modeling and Analyzing Live Streaming Performance
abstract
Today, live streaming is gaining a rapid growth in use, which refers to streaming the media content recorded and broadcast in real time. In live streaming, latency is of utmost importance since smaller latency means higher user engagement. HTTP adaptive streaming (HAS) is now the most popular live streaming technology, where the video client sends HTTP requests to server to download video segments. The bitrate adaptation (ABR) algorithm inside the client determines bitrate level for every segment. It is of great help for ABR algorithm to quantify the influence of different HAS factors on streaming performance. However, existing work mainly focuses on video on demand (VoD) streaming rather than live streaming. In this paper, we theoretically analyze live streaming performance. We first establish a queuing model to describe playout buffer evolution. Based on the model, we respectively characterize rebuffering probability, rebuffering count and streaming latency, and analyze the effects of chunk arrival rate, arrival interval fluctuation, startup threshold and video skipping on them. From analysis results, we propose insights and recommendations for bitrate adaptation in live streaming and design a simple heuristic ABR algorithm leveraging them. Extensive simulations verify the insights as well as effectiveness of the designed algorithm.
Tong Zhang 0018, Fengyuan Ren, Bo Wang 0066
IWQoS1
2020 Re-architecting Congestion Management in Lossless Ethernet
Wenxue Cheng, Kun Qian 0017, Wanchun Jiang, Tong Zhang 0018, Fengyuan Ren
NSDI4
2020 Towards Influence of Chunk Size Variation on Video Streaming in Wireless Networks
abstract
In recent years, the growth in popularity of mobile video streaming services is unbroken. There are tremendous demands for video streaming over wireless networks. Currently, most video streaming is over HTTP. Up to now, HTTP-based adaptive video streaming is standardized as DASH, where a client-side video player can dynamically pick the bitrate level according to the perceived network conditions. Actually, not only the available bandwidth drastically varies due to wireless network properties, but also the chunk sizes in the same bitrate level significantly fluctuate, which also influences the bitrate adaptation. However, existing bitrate adaptation algorithms mostly focus on available bandwidth but do not involve chunk size variation, leading to performance losses. In this paper, we theoretically analyze the influence of chunk size variation on bitrate adaptation performance in wireless networks. Based on DASH system features, we build a general model describing playback buffer evolution. Applying stochastic theories, we respectively analyze the influence of the chunk size variation on rebuffering probability, average bitrate, and bitrate switching interval. Furthermore, based on theoretical insights, we provide several suggestions for algorithm designing and rate encoding, and also design a simple bitrate adaptation algorithm. Extensive simulations verify our insights, suggestions, and designed algorithm effectiveness.
Tong Zhang 0018, Fengyuan Ren, Wenxue Cheng, Xiaohui Luo, Ran Shu 0001
IEEE Trans. Mob. Comput.1
2020 Towards Power Efficient High Performance Packet I/O
abstract
Recently, high performance packet I/O frameworks continue to flourish for their ability to process packets from high-speed links. To achieve high throughput and low latency, high performance packet I/O frameworks usually employ busy polling. As busy polling will burn all CPU cycles even if there's no packet to process, these frameworks are quite power inefficient. However, exploiting power management techniques such as DVFS and LPI in the frameworks is challenging, because neither the OS nor the frameworks can provide information (e.g., actual CPU utilization, available idle period, or the target frequency) required by these techniques. In this article, we establish a model that can formulate the packet processing flow of high performance packet I/O to help and address the above challenges. From the model, we can deduce the information needed for power management techniques, and gain the insights to balance the power and latency. After suggesting to use pause instruction to reduce CPU power within short idle period, we propose two approaches to conduct power conservation for high performance packet I/O: one with the aid of traffic information and the other without. Experiments with Intel DPDK show that both approaches can achieve significant power reduction with little latency increase.
Wenxue Cheng, Tong Zhang 0018, Fengyuan Ren, Bailong Yang
IEEE Trans. Parallel Distributed Syst.3
2019 FlexGate: High-performance Heterogeneous Gateway in Data Centers
abstract
Large-scale data centers support various applications and process/issue terabits per second traffic from/to Internet. On the boundary of data center, the gateway needs to execute a series of network functions for each incoming packet. The Network Function Virtualization (NFV) technology leverages commodity servers to flexibly implement network functions. This solution provides satisfying processing and storage capability. However, state-of-the-art NFV platforms can merely process network functions at the line rate of 10~40Gbps. Supporting throughput of terabits per second requires dozens or even hundreds of servers operating exclusively for network functions, which is not only expensive but also difficult to maintain. On the other hand, programmable packet processing hardwares proposed in recent years offer a new platform for implementing network functions. They can execute user-defined packet processing logics at ultra-high line rate while containing limited processing and storage resources.
Kun Qian 0017, Mao Miao, Jianyuan Lu, Tong Zhang 0018, Peilong Wang, Fengyuan Ren
APNet5
2019 Active and Adaptive Application-Level Flow Control for Latency Sensitive RPC Applications
abstract
The Remote Procedure Call (RPC) frameworks are widely deployed in industry. Applications supported by RPC frameworks are often latency-sensitive which strictly require to be responded before the deadline. For meeting this requirement, RPC frameworks adopt the application-level flow control mechanism. This mechanism gives an appropriate threshold determining the number of RPC requests that the server can process, thus avoids missing the deadline. However, this threshold at the application-level is a fixed empirical value so that it is hard to obtain respectable performance because an endpoint's processing capacity can take a huge quantity of values by varying workload and different hardware configurations. While other methods based on specialized transport protocols are adaptive, they will introduce extra costs for message reporting from server to client. Furthermore, adopting specialized transport protocols will also introduce extra transplanting efforts for TCP-based applications. In this paper, we provide an active and adaptive application level flow control mechanism at the client side. We first design an algorithm to find the appropriate threshold to achieve the desired response time. Then based on this algorithm, we control the threshold to bound the response time as expected. We implement our flow control mechanism using a memcached testbed. Experiments prove that our mechanism can accurately reduce the mean and 99th percentile response time by at least 71.3% and 69.4% respectively, while keeping a relatively high QPS. Furthermore, compared to static-threshold mechanism, our flow control mechanism is more efficient under low latency constraints.
Jing Xie 0005, Wenxue Cheng, Tong Zhang 0018, Qingkai Meng 0001, Fengyuan Ren
ICPADS3
2019 Gentle flow control: avoiding deadlock in lossless networks
abstract
Many applications in distributed systems rely on underlying lossless networks to achieve required performance. Existing lossless network solutions propose different hop-by-hop flow controls to guarantee zero packet loss. However, another crucial problem called network deadlock occurs concomitantly. Once the system traps in a deadlock, a large part of network would be disabled. Existing deadlock avoidance solutions focus all their attentions on breaking the cyclic buffer dependency to eliminate circular wait (one necessary condition of deadlock). These solutions, however, impose many restrictions on network configurations and side-effects on performance.
Kun Qian 0017, Wenxue Cheng, Tong Zhang 0018, Fengyuan Ren
SIGCOMM3
2019 Distributed Bottleneck-Aware Coflow Scheduling in Data Centers
abstract
With the booming development of data parallel frameworks, the coflow abstraction has been greatly favored by data center transport designs, for its prominent ability in capturing application-level semantics. To accelerate job completion, coflow completion time (CCT) is a most important metric, and coflow scheduling is the most effective and widely-adopted means of optimizing CCT. However, most existing coflow scheduling mechanisms neglect the ubiquitous in-network bottlenecks and schedule coflows based on non-blocking giant switch hyperthesis. Such a practice is likely to result in undesired link contention inside the fabric, finally impairing CCT performance. To address this problem, we propose the Distributed Bottleneck-Aware coflow scheduling algorithm called DBA, which approximates the minimum remaining time first (MRTF) heuristic on all fabric-wide links. In this way, core link bandwidths are allocated to coflows as expected and the CCT performance will not be violated. As an evolutionary algorithm, DBA enhances the traditional dual decomposition method thus converges to the optimal bandwidth allocation very fast. Extensive simulations verify DBA's outstanding CCT performance as well as high link utilization. Furthermore, DBA introduces very little overhead and is robust to routing strategies, parameter variations and computation delays.
Tong Zhang 0018, Ran Shu 0001, Zhiguang Shan, Fengyuan Ren
IEEE Trans. Parallel Distributed Syst.1
2018 Estimating Short Connection Capacity on High Performance User Level Network Stack
abstract
Short connections are generally used to transfer small-size messages, which contribute a large part of workload in modern applications. The maximum sustainable short connection rate, which is called short connection capacity, is an important index for admission control, Web QoS control, and energy saving. A capacity estimation mechanism aims to find the workload just saturating the server, and it relies on both workload information and system information. Past researches point out that kernel space network stack becomes the bottleneck when a huge number of concurrent short connections coexist. On the other hand, high performance user level network stacks have been proved to eliminate such bottleneck, thus become a hot research topic in both academia and industry. However, they also bring challenges for estimating short connection capacity, making traditional methods ineffective. Therefore, it is important to find a new method to estimate short connection capacity on high performance user level network stacks. In this paper, we prove that the effective CPU utilization is an adaptive index to different workload patterns and application complexities, which can reflect the server state. Then we design and implement an online capacity estimator on the Seastar platform. We conduct experiments to verify the effectiveness of our online capacity estimator. The results show that our estimator can actually estimate the capacity online. When the server is near saturated, the 90th percentile relative estimating error is no more than 9.18%. Furthermore, our capacity estimator only introduces no more than 1.38% of capacity loss in our experiments.
Jing Xie 0005, Wenxue Cheng, Tong Zhang 0018, Danfeng Shan, Fengyuan Ren
ICCCN3
2018 Power Efficient High Performance Packet I/O
abstract
Recently, high performance packet I/O frameworks are expected an extensive application for their ability to process packets from 10Gbps or higher speed links. To achieve high throughput and low latency, high performance packet I/O frameworks usually employ busy polling technique. As busy polling will burn all CPU cycles even if there's no packet to process, these frameworks are quite power inefficient. Meanwhile, exploiting power management techniques such as DVFS and LPI in high performance packet I/O frameworks is challenging, because neither the OS nor the frameworks can provide information (e.g., the actual CPU utilization, available idle period, or the target frequency) required by power management techniques. In this paper, we establish an analytical model that can formulate the packet processing flow of high performance packet I/O to help address the above challenges. From the analytical model, we can deduce the actual CPU utilization and average idle period in different traffic load, and gain the insight to choose CPU frequency that can appropriately balance the power consumption and packet latency. Then, we propose two simple but effective approaches to conduct power conservation for high performance packet I/O: one with the aid of traffic information and the other without. Experiments with Intel DPDK show that both approaches can achieve significant power reduction (35.90% and 34.43% on average respectively) while incurring < 1 μs of latency increase.
Wenxue Cheng, Tong Zhang 0018, Jing Xie 0005, Fengyuan Ren, Bailong Yang
ICPP3
2018 High Performance Userspace Networking for Containerized Microservices
Xiaohui Luo, Fengyuan Ren, Tong Zhang 0018
ICSOC3
2018 Scheduling Coflows with Incomplete Information
abstract
In recent years, the coflow abstraction has received significant attentions, for its prominent ability to capture application semantics. On this basis, multiple coflow scheduling mechanisms have been proposed to minimize the coflow completion time (CCT). Currently, existing coflow scheduling mechanisms mainly belong to two categories: information-omniscient and information-agnostic. However, in data center applications, there are still quite a few cases in between where incomplete coflow information is known, and such incomplete information makes great contributions to improving the CCT performance. To address such cases, we propose IICS, a coflow scheduling algorithm based on incomplete coflow information. IICS leverages information of a coflow's arrived parts to deduce the coflow's remaining transmission time, and uses it to approximate the Minimum Remaining Time First (MRTF) heuristic. Besides, IICS allocates bandwidth by monopolization and in a maximal manner, which achieves high bandwidth utilization. Extensive simulations under realistic settings show that IICS achieves the average CCT comparable to that of the information-omniscient algorithm and the 99th percentile CCT much smaller than both information-omniscient and information-agnostic algorithms. Furthermore, IICS holds observably higher throughput and is robust to algorithm parameters.
Tong Zhang 0018, Fengyuan Ren, Ran Shu 0001, Bo Wang 0066
IWQoS1
2018 Analysing and improving convergence of quantized congestion notification in Data Center Ethernet
Ran Shu 0001, Fengyuan Ren, Jiao Zhang 0002, Tong Zhang 0018, Chuang Lin 0002
Comput. Networks4
2018 Towards Stable Flow Scheduling in Data Centers
abstract
At present, soft real-time data center applications are in a booming development and impose stringent delay requirements on internal data transfers. In this context, many recently proposed data center transport protocols share a common goal of minimizing Flow Completion Time (FCT), and the Shortest Remaining Processing Time (SRPT) scheduling algorithm has attracted widespread attentions for its superior performance in average FCT. However, SRPT suffers from the instability problem, incurring more and more flows left uncompleted even if the traffic load is within the fabric capacity, which implies unnecessary bandwidth waste. To solve the problem, this paper proposes a backlog-aware flow scheduling algorithm (BASRPT) for both giant switch and general topologies. Because of taking into account queue backlogs other than flow sizes at scheduling, we prove that BASRPT is stable and still maintains good FCT performance. To overcome the huge computation overhead and enable distributed implementation, a fast and practical approximation algorithm called fast BASRPT is also developed. Extensive flow-level simulations show that fast BASRPT indeed stabilizes the queue length and obtains a higher throughput while being able to push the FCT arbitrarily close to the optimal value in the condition of feasible traffic loads.
Tong Zhang 0018, Fengyuan Ren, Ran Shu 0001
IEEE Trans. Parallel Distributed Syst.1
2018 MPTCP Tunnel: An Architecture for Aggregating Bandwidth of Heterogeneous Access Networks
abstract
Fixed and cellular networks are two typical access networks provided by operators. Fixed access network is widely employed; nevertheless, its bandwidth is sometimes not sufficient enough to meet user bandwidth requirements. Meanwhile, cellular access network owns unique advantages of wider coverage, faster increasing link speed, more flexible deployment, and so forth. Therefore, it is attractive for operators to mitigate the bandwidth shortage by bundling these two. Actually, there have been existing schemes proposed to aggregate the bandwidth of two access networks, whereas they all have their own problems, like packet reordering or extra latency overhead. To address this problem, we design new architecture, MPTCP Tunnel, to aggregate the bandwidth of multiple heterogeneous access networks from the perspective of operators. MPTCP Tunnel uses MPTCP, which solves the reordering problem essentially, to bundle multiple access networks. Besides, MPTCP Tunnel sets up only one MPTCP connection at play which adapts itself to multiple traffic types and TCP flows. Furthermore, MPTCP Tunnel forwards intact IP packets through access networks, maintaining the end‐to‐end TCP semantics. Experimental results manifest that MPTCP Tunnel can efficiently aggregate the bandwidth of multiple access networks and is more adaptable to the increasing heterogeneity of access networks than existing mechanisms.
Danfeng Shan, Ran Shu 0001, Tong Zhang 0018
Wirel. Commun. Mob. Comput.4
2017 Modeling and Analyzing Latency in the Memcached system
abstract
Memcached is a widely used in-memory caching solution in large-scale searching scenarios. The most pivotal performance metric in Memcached is latency, which is affected by various factors including the workload pattern, the service rate, the unbalanced load distribution and the cache miss ratio. To quantitate the impact of each factor on latency, we establish a theoretical model for the Memcached system. Specially, we formulate the unbalanced load distribution among Memcached servers by a set of probabilities, capture the burst and concurrent key arrivals at Memcached servers in form of batching blocks, and add a cache miss processing stage. Based on this model, algebraic derivations are conducted to estimate latency in Memcached. The latency estimation is validated by intensive experiments. Moreover, we obtain a quantitative understanding of how much improvement of latency performance can be achieved by optimizing each factor and provide several useful recommendations to optimal latency in Memcached.
Wenxue Cheng, Fengyuan Ren, Wanchun Jiang, Tong Zhang 0018
ICDCS4
2017 Modeling and analyzing the influence of chunk size variation on bitrate adaptation in DASH
abstract
Recently, HTTP-based adaptive video streaming has been widely adopted in the Internet. Up to now, HTTP-based adaptive video streaming is standardized as Dynamic Adaptive Streaming over HTTP (DASH), where a client-side video player can dynamically pick the bitrate level according to the perceived network conditions. Actually, not only the available bandwidth is varying, but also the chunk sizes in the same bitrate level significantly fluctuate, which also influences the bitrate adaptation. However, existing bitrate adaptation algorithms do not accurately involve the chunk size variation, leading to performance losses. In this paper, we theoretically analyze the influence of chunk size variation on bitrate adaptation performance. Based on DASH system features, we build a general model describing the playback buffer evolution. Applying stochastic theories, we respectively analyze the influence of the chunk size variation on rebuffering probability and average bitrate level. Furthermore, based on theoretical insights, we provide several recommendations for algorithm designing and rate encoding, and also propose a simple bitrate adaptation algorithm. Extensive simulations verify our insights as well as the efficiency of the proposed recommendations and algorithm.
Tong Zhang 0018, Fengyuan Ren, Wenxue Cheng, Xiaohui Luo, Ran Shu 0001
INFOCOM1
2017 Congestion control in Converged Ethernet with heterogeneous and time-varying delays
abstract
Congestion control is an indispensable mechanism in the new trend of enhanced Ethernet as a unified fabric for traditional LAN, SAN, and high-performance computing networks. A congestion management framework for Converged Ethernet (CE) networks has been standardized by IEEE 802.1 Qau work group, and QCN is recommended as the congestion control scheme in the standard draft. QCN is heuristically designed for 1/10Gbps Ethernet without considering the impact of delays. Recent work find that QCN will encounter stability issues with feedback delays, and these issues will be more serious as Ethernet extends to 40/100Gbps and the delays become heterogeneous and time-varying. This work aims to mitigate the negative impact of delays on congestion control scheme in CE. Specially, considering the delays are heterogeneous and time-varying, we build a model for Converged Ethernet with the standard congestion management framework. The model provides a new congestion detector to estimate the real congestion status under the impact of delays and regards the heterogeneous and time-varying feature as disturbances. Leveraging the new congestion detector and tolerating the disturbance through the sliding mode control method, we design the Delay-tolerant Sliding Mode (DSM) congestion control scheme. Extensive simulations show that DSM outperforms other congestion control schemes when the Ethernet ranges from 1Gbps to 100Gbps and the delays are heterogeneous and time-varying.
Wenxue Cheng, Wanchun Jiang, Tong Zhang 0018, Bo Wang 0066, Kun Qian 0017, Fengyuan Ren
IWQoS3
2017 Performance analysis of randomized data fetching in cluster computing
abstract
The shuffle transfer pattern is widely adopted in today's cluster computing applications and the completion time of each group of transmissions directly affects application performance. Because of the restriction on the number of concurrent threads and the TCP Incast problem, the randomized data fetching strategy is widely employed in this kind of communication in practice. In this paper, to assess the performance of randomized data fetching, we build a general analytical model and define two metrics - link overload probability and K-deviation load balancing probability - to evaluate the degree of link overload and load balancing respectively, since they are closely related to the transfer completion time. Leveraging our model, we theoretically analyze the transfer performance in three typical scenarios and provide recommendations for setting the number of concurrent connections per receiver. Finally, we validate the theoretical analysis as well as the recommendations through extensive simulations.
Tong Zhang 0018, Peng Cheng 0005, Wenxue Cheng, Bo Wang 0066, Fengyuan Ren
IWQoS1
2017 Awakening Power of Physical Layer: High Precision Time Synchronization for Industrial Ethernet
abstract
High-precision time synchronization is critical for nowadays industrial Ethernet systems. Most existing time synchronization mechanisms are implemented based on packet communication. This interaction pattern, however, greatly limits their synchronizing frequency. In order to achieve microsecond-level synchronization precision, expensive high-quality oscillator is necessary for maintaining low clock skew under this long synchronization period (usually several seconds). Furthermore, packet processing introduces many nondeterministic variances (e.g. network stack overhead), which needs to be carefully eliminated. In this paper, we propose the brand-new Industrial Time Protocol (ITP). We deploy the entire ITP in the physical layer, so it eliminates most time uncertainties caused by network stack processing. Furthermore, ITP leverages the InterFrame Gap (IFG), which is the inherent interval between any two Ethernet frames, to carry the synchronization message. With this novel design, ITP can synchronize peer devices at very high frequency without degrading the goodput. The accuracy of ITP is bounded by 16ns for two adjacent devices with only intrinsic cheap oscillator. Furthermore, our theoretical analysis deduces that ITP guarantees 16N-nanosecond accuracy for N-hop network. We implement ITP design with NetFPGA. Experiments show that ITP can provide about 76-nanosecond accuracy for #hops=16 network under severe congestions. In addition, the design of ITP is scalable. It only consumes about 0.67% of logic cells in the low-end FPGA for supporting every ITP-aware port increase.
Kun Qian 0017, Tong Zhang 0018, Fengyuan Ren
RTSS2
2016 Backlog-Aware SRPT Flow Scheduling in Data Center Networks
abstract
The rapidly developing soft real-time data center applications impose stringent delay requirements on internal data transfers. Therefore many recently emerged network protocols in data center share a common goal of decreasing Flow Completion Time (FCT), in which case the Shortest Remaining Processing Time (SRPT) scheduling discipline has attracted widespread attentions. However, SRPT suffers the instability issue, incurring more and more flows left uncompleted even when traffic load is within network capacity, which implies unnecessary bandwidth waste. To solve the problem, this paper proposes a backlog aware scheduling algorithm (BASRPT) that stabilizes queue length while maintaining relatively low FCT based on Lyapunov optimization. To overcome the huge computational overhead, a fast and practical approximation algorithm called fast BASRPT is also developed. Extensive flow-level simulations show that fast BASRPT indeed stabilizes switch queue and obtains a higher throughput while being able to push FCT arbitrarily close to the optimal value in the condition of feasible traffic load.
Tong Zhang 0018, Fengyuan Ren, Ran Shu 0001
ICDCS1