VLDB 2026 Research / reviewers in the wild / expert
Tao Zhang 0019
dblp:15/4777-19
· DBLP profile ↗
36ranked-venue papers
16as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 15 · 7 first-author · 11 since 2021Systems, architecture and hardware · 13 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context ReasoningabstractLong Chain-of-Thought (CoT) reasoning has significantly advanced the capabilities of Large Language Models (LLMs), but this progress is accompanied by substantial memory and latency overhead from the extensive Key-Value (KV) cache.Although KV cache quantization is a promising compression technique, existing low-bit quantization methods often exhibit severe performance degradation on complex reasoning tasks.Fixed-precision quantization struggles to handle outlier channels in the key cache, while current mixed-precision strategies fail to accurately identify components requiring high-precision representation.We find that an effective low-bit KV cache quantization strategy must consider two factors: a key channel's intrinsic quantization difficulty and its relevance to the query.Based on this insight, we propose MixKVQ, a novel quantization method that introduces a lightweight, query-aware algorithm to identify and preserve critical key channels that need higher precision, while applying per-token quantization for value cache.Experiments on complex reasoning datasets demonstrate that our approach significantly outperforms existing low-bit methods, achieving performance comparable to a fullprecision baseline at a substantially reduced memory footprint.The source code is available at https://github.com/ZeroNLP/MixKVQ. Tao Zhang 0019, Ziqian Zeng, Huiping Zhuang, Cen Chen 0002 |
ACL (1) | 1 |
| 2026 | Breaking the Accuracy-Latency Trade-Off in Sketch Compression for Network Measurement
Tao Zhang 0019, Siyuan Fan, Yunsheng Liu, Linfei Dong, Haozhi Tang, Hui Yin 0001, Fangmin Li |
IWQoS | 2 |
| 2026 | Switch-Transparent Load Balancing for RDMA Data Centers: A Host-Only Approach
Tao Zhang 0019, Haozhi Tang, Linfei Dong, Siyuan Fan, Hui Yin 0001, Fangmin Li |
IWQoS | 1 |
| 2026 | Balancing data center traffic load with speeding up flow-transmission
Tao Zhang 0019, Yunsheng Liu, Haotian Jing, Siyuan Fan, Haozhi Tang, Xidao Luan |
Future Gener. Comput. Syst. | 1 |
| 2026 | MGID-Net: Multimodal denoising for industrial GPR images with deep image-text fusion
Wentai Lei, Tao Zhang 0019 |
Pattern Recognit. | 3 |
| 2026 | An FPGA-Based Frequency-Focused Vision Transformer Accelerator for Real-Time Inference on Edge PlatformsabstractVision Transformers (ViTs) often struggle to effectively capture local features. SpectFormer addresses this limitation by incorporating spectral blocks into the shallow layers of DeiT. However, deploying SpectFormer on edge devices is challenging due to its high computational complexity, the implementation of spectral blocks, additional matrix transpositions, and intensive nonlinear operations. To address these challenges, we propose an FPGA-based Frequency-focused Vision Transformer Accelerator (FFVTA), the first FPGA-based ViT accelerator that incorporates frequency-domain optimization for real-time inference of SpectFormer on edge devices. FFVTA employs a Unified DFT-Attention Matrix (UDAM) architecture to unify the computation of spectral and self-attention blocks, significantly enhancing the hardware reuse and reducing resource usage. Additionally, FFVTA introduces a Block-Broadcast Loop (BBL) Dataflow to efficiently accelerate matrix multiplication in self-attention blocks and a Configurable Two-Stage Log-Softmax (CTS-Log-Softmax) computation process to optimize nonlinear function computations. These innovations improve computational efficiency while reducing the resource usage. When deployed on a resource-constrained FPGA platform (KV260), FFVTA achieves a 16.61× speedup compared to the baseline, with significant reductions in the hardware resource usage and power consumption. FFVTA demonstrates superior performance with an energy efficiency of 63.40 GOPs/W and a DSP efficiency of 0.46 GOPs/DSP. FFVTA not only achieves efficient acceleration for the hierarchical SpectFormer-H-S variant, but also attains competitive accuracy of 84.02% on ImageNet classification. Chengrui Tian, Congwei Liao, Tao Zhang 0019, Jiafeng Ding, Lianwen Deng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Toward QoE-Fairness for Video Streaming Over Heterogeneous Networks: An Innovative Bandwidth Allocation MechanismabstractWith the growing ubiquity of video streaming, ensuring a fair and high quality of experience (QoE) for users has emerged as a shared concern among video content providers. State-of-the-art video delivery systems achieve QoE fairness through bottleneck bandwidth allocation across multiple video streams, all based on the assumption of a unified congestion control (CC) protocol. However, the widespread use of heterogeneous CC protocols on the Internet not only disrupts QoE fairness among video streams but also poses challenges in achieving fast convergence under dynamic bandwidth. To address these issues, we propose a QoE-Fairnessawarebandwidthallocationmechanism called Fabam, which establishes a unified QoE control plane across heterogeneous CC protocols. Fabam constructs independent virtual targets based on the real-time QoE of each video stream to achieve QoE fairness, and offers rapid convergence for the underlying CC protocols to improve efficiency. In addition, we propose a Deep Neural Network (DNN)-based multi-step mapping model aimed at balancing the performance and overhead of Fabam, thereby enhancing its deployment potential in practical applications. We implement Fabam on QUIC and integrate it with Dash.js. The evaluation results demonstrate the significant superiority of Fabam over the state-of-the-art approaches, including an enhancement of 44.01% in QoE fairness and an improvement of 36.39% in QoE efficiency. Meanwhile, Fabam-DNN maintains satisfactory QoE fairness while supporting multiple users at a low cost. Qichen Su, Jiawei Huang 0001, Weihe Li, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Netw. | 5 |
| 2025 | GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language ModelsabstractLarge Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective approach to mitigate gender biases. Although proprietary LLMs have made significant strides in mitigating gender bias, their alignment datasets are not publicly available. The commonly used and publicly available alignment dataset, HH-RLHF, still exhibits gender bias to some extent. There is a lack of publicly available alignment datasets specifically designed to address gender bias. Hence, we developed a new dataset named GenderAlign, aiming at mitigating a comprehensive set of gender biases in LLMs. This dataset comprises 8k single-turn dialogues, each paired with a “chosen” and a “rejected” response. Compared to the “rejected” responses, the “chosen” responses demonstrate lower levels of gender bias and higher quality. Furthermore, we categorized the gender biases in the “rejected” responses of GenderAlign into 4 principal categories. The experimental results show the effectiveness of GenderAlign in reducing gender bias in LLMs. Tao Zhang 0019, Ziqian Zeng, YuxiangXiao YuxiangXiao, Huiping Zhuang, Cen Chen 0002, James R. Foulds, Shimei Pan |
ACL (1) | 1 |
| 2025 | Mitigating Hash Polarization with Flow-Level Load Balancing in Leaf-Spine Data Center Network
Siyuan Fan, Tao Zhang 0019, Linfei Dong, Xidao Luan, Hui Yin 0001 |
ICA3PP (4) | 4 |
| 2025 | SMAR: Short-Flow Multi-path Adaptive Routing for Heterogeneous RDMA Workloads
Tao Zhang 0019, Xidao Luan, Hui Yin 0001, Jyoti Sahni, Winston Khoon Guan Seah |
ICA3PP (8) | 1 |
| 2025 | GPR Bscan Imaging Enhancement Method for Rebar OcclusionabstractWhen using ground penetrating radar (GPR) to detect targets below shallow rebar mesh in reinforced concrete structures, the strong scattering characteristics of rebar mesh causes distortion and interference of targets echoes and leads to imaging artifacts and degradation. This letter proposes a coarse-scale and fine-scale dual-branch imaging enhancement network (CFD-IENet) to achieve target imaging under rebar mesh in reinforced concrete by combining Bscan echo data enhancement with BP imaging result enhancement. First, a residual U network (Res-U) suppresses complex background clutter in Bscan data to improve the signal-to-noise ratio. Then, a coarse-scale and fine-scale dual-branch network is constructed to enhance both Bscan and BP imaging. In the Bscan enhancement stage, strong and weak signals are trained separately, aiming for surface rebar echo interference in reconstructing weak target signals beneath the rebar mesh. In the BP imaging enhancement stage, artifacts and multipath ghosts are suppressed to enhance occluded target imaging. A bilinear fusion module (BFM) is designed to facilitate global feature interaction, promoting the fusion of Bscan and BP imaging features across scales, thereby improving reconstruction and enhancement accuracy. Experimental results on cracks occluded by rebar mesh demonstrate the method’s effectiveness, showing a 4.73dB improvement in PSNR and a 0.16 improvement in SSIM compared to the RNMF+BP+Unet enhancement method. Qiguo Xu, Tao Zhang 0019, Zebang Pang, Wentai Lei |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Subkv: Quantizing Long Context KV Cache for Sub-Billion Parameter Language Models on Edge DevicesabstractABSTRACT Background Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their substantial computational and memory requirements present significant challenges for widespread deployment on edge devices. Motivation In long‐context scenarios, even sub‐billion parameter LLMs face unavoidable memory and performance bottlenecks due to inefficient KV Cache utilization. Existing quantization methods fail to address these challenges effectively. Method This paper addresses these challenges by introducing advanced quantization techniques tailored for sub‐billion parameter LLMs. It specifically targets reducing memory consumption through the conversion of the model's KV Cache to lower‐bit integers. We present SubKV, a quantization method specifically designed to optimize the KV Cache in sub‐billion parameter LLMs. Our analysis reveals distinct distributional differences in the magnitude of key and value caches. Leveraging this insight, we apply Per‐Channel Quantization to the key cache and Per‐Token Quantization to the value cache. Furthermore, we introduce the Dynamic Window Quantization method to enhance attention computations. To mitigate the extreme sensitivity of the first token, we also introduce Attention Sink‐Aware Quantization. Results Experimental results demonstrate that SubKV significantly reduces the KV Cache size during long context inference while maintaining model performance, offering superior results to existing KV Cache quantization methods. Ziqian Zeng, Tao Zhang 0019, Zhengdong Lu, Huiping Zhuang, Hongen Shao, Sin G. Teo, Xiaofeng Zou |
Softw. Pract. Exp. | 2 |
| 2025 | Progress-Aware Transmission Protocol for Efficient In-Network Aggregation in Distributed Machine LearningabstractLarge-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortunately, the phase of gradient aggregation has become the performance bottleneck for data-parallel DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. Besides, we reveal that existing INA solutions cannot provide the fairness performance among multiple jobs with varying number of workers. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. Moreover, to ensure the fair throughput among multiple jobs, we dynamically adjust the aggregator allocation for each job by tuning the number of hash operations. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols. Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
IEEE Trans. Netw. | 3 |
| 2025 | $R^{3}$R3: A Building Block for Disordering-Tolerant Load Balancing in Data Center NetworksabstractPacket-level load balancing has shown its massive potential for long in utilizing super high bisection bandwidth of data center network (DCN). This kind of potential, however, has still not been completely transformed into huge performance enhancement of data transmission. The fundamental reason is that packet-level load balancing can fully utilize the parallel paths of underlying physical network, but suffer from the problem of packet disordering transmission, which greatly impairs the flow-level transmission performance of DCN. This paper explores the root cause of performance impairment generated by packet disordering transmission, and proposes$R^{3}$, a solution focusing on “recognizably releasing redundant acknowledgements” as a building block for data center packet-level load balancer. In$R^{3}$'s heart, the source leaf switch perceives the global packet loss information and selectively intercepts the redundant acknowledgement packets, thus avoiding the TCP-driven end-host from experiencing frequent window reductions and unnecessary packet retransmissions. Experimental results of numerous simulation tests and real implementations show that, after integrating$R^{3}$into the representative data center packet-level load balancing schemes, the transmission performances of both delay-sensitive and throughput-oriented data center flows are significantly improved. Furthermore,$R^{3}$is merely implemented by switch, leaving the end hosts and the deployed load balancing scheme totally unchanged. Tao Zhang 0019, Yuanzhen Hu, Jinbin Hu 0001, Haotian Jing, Yangfan Li 0001, Xidao Luan |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | Leveraging Packet Cloning to Achieve Fast Flow-Transmission for Data Center Load BalancingabstractModern data center network often possesses multiple end-to-end parallel paths, which undertake the crucial task of transmitting vast heterogeneous data traffic generated by a wide variety of applications. To fully utilize the offered super high bisection network bandwidth thus benefiting application performance, many data center load balancing schemes are proposed to improve path utilization for avoiding network congestion hot-spot. However, these schemes are naturally agnostic to data center traffic pattern and the diverse requirements on flow-transmission, leading to the sub-optimal network transmission performance. To address this issue, this paper presents a new data center load balancing scheme, called PCLB, which selectively generates Clone Packets by considering both flow-transmission phases and path states, thereby helping different types of flows choose more appropriate paths for speeding up their data transmission. Experimental results of numerous NS2 simulations show that PCLB significantly reduces the average and tail flow completion time for delay-sensitive flows, while the performance of throughput-oriented flows can be always maintained at high level. Haotian Jing, Tao Zhang 0019, Shaojun Zou, Xidao Luan, Hui Yin 0001, Fangmin Li |
ISPA | 4 |
| 2024 | Achieving Ultra-low Latency for Timeout-less Congestion Control in Data Center NetworksabstractModern data centers are hosting a great number of various applications (e.g. MapReduce and web search) that require a high fan-in data communication, which easily causes serious packet losses and timeouts, substantially degrading the application performance. To address this issue, various host-based and switch-based transport protocols are proposed to eliminate timeout and improve the user experience. Unfortunately, although existing transport protocols can effectively eliminate the timeout, they inevitably result in persistent queueing backlog and degrade the network performance, especially delay-sensitive short flows. To this end, we propose a general scheme with ultra-low latency called UL2to address the above problem. Concretely, the sender periodically estimates the queueing delay of each packet on the transmission path and senses the degree of congestion based on its measured result. Then the sender timely yet cautiously executes a pausing transmission operation based on measured queueing delay, guaranteeing fast elimination of queue delays and high link utilization. Our evaluation indicates that UL2can effectively eliminate queue backlog and reduce the queueing delay by more than 90%. Moreover, UL2enhances the performance of state-of-the-art transport protocols in terms of flow completion times by up to 44.98%. Shaojun Zou, Jiacheng Qu, Tao Zhang 0019, Yuanzhen Hu, Yujie Peng |
ISPA | 4 |
| 2024 | Achieving QoE Fairness in Video Streaming over Heterogeneous Congestion Control ProtocolsabstractWith the growing ubiquity of video streaming, ensuring a fair and high quality of experience (QoE) for users has emerged as a shared concern among video content providers. State-of-the-art video delivery systems achieve QoE fairness through bottleneck bandwidth allocation across multiple video streaming, all based on the assumption of a unified congestion control (CC) protocol. However, the widespread use of heterogeneous CC protocols on the Internet not only disrupts QoE fairness among video streaming but also poses challenges in achieving fast convergence under dynamic bandwidth. To address these issues, we propose a QoE-Fairness aware bandwidth allocation mechanism called Fabam, which establishes a unified QoE control plane across heterogeneous CC protocols. Fabam constructs independent virtual targets based on the real-time QoE of each video streaming to achieve QoE fairness, and offers rapid convergence for the underlying CC protocols to improve efficiency. We implement Fabam on QUIC and integrate it with Dash.js. The evaluation results demonstrate the significant superiority of Fabam over the state-of-the-art approaches, including an enhancement of 24.48% in QoE fairness and an improvement of 16.63% in QoE efficiency. Qichen Su, Jiawei Huang 0001, Weihe Li, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001 |
IWQoS | 5 |
| 2024 | HG: Leveraging Hybrid Switching Granularity to Balance Heterogeneous Data Center Traffic Load for Cloud-Based Industrial ApplicationsabstractNowadays, the deluge of heterogeneous data generated by various cloud-based industrial applications often has to be delivered to the data center for analysis and storage. To speed up data processing thus facilitating application performance, the modern data center network offers rich parallel paths and super high bisection bandwidth for data communications between servers, expecting to provide good transmission performance for the heterogeneous data traffic caused by cloud-based industrial applications. Due to high path diversities, however, balancing the heterogeneous traffic load across multiple parallel paths for fully utilizing the offered super high bisection bandwidth is full of challenges (i.e., how to achieve high path utilization without incurring adverse impact). Although prior studies demonstrate that the flowlet-based solutions are promising to fill the bill, we argue that their rerouting operations are still inappropriate in timing and manner. This article presents HG, a load balancing scheme adopting hybrid switching granularity to make traffic rerouting. HG embeds the flow-fragment-based and flowcell-based path switching into the flowlet-based path switching, and employs state-weighted path measurement to choose paths for newly appeared flow fragments, flowcells, and flowlets. The results of numerous NS2 tests show that, compared with the state-of-the-art data center load balancing schemes, HG significantly reduces the average and tail-flow completion times for delay-sensitive flows, and the throughput of throughput-oriented flows is always maintained at high level. Tao Zhang 0019, Shengli He, Ku Jin, Yuanzhen Hu, Chang Ruan, Shaojun Zou, Jinbin Hu 0001, Fangmin Li |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Straggler-Aware Gradient Aggregation for Large-Scale Distributed Deep Learning SystemabstractDeep Neural Network (DNN) is a critical component of a wide range of applications. However, with the rapid growth of the training dataset and model size, communication becomes the bottleneck, resulting in low utilization of computing resources. To accelerate communication, recent works propose to aggregate gradients from multiple workers in the programmable switch to reduce the volume of exchanged data. Unfortunately, since using synchronization transmission to aggregate data, current in-network aggregation designs suffer from the straggler problem, which often occurs in shared clusters due to resource contention. To address this issue, we propose a straggler-aware aggregation transport protocol (SA-ATP), which enables the leading worker to leverage the spare computing and storage resources to help the straggling worker. We implement SA-ATP atop clusters using P4-programmable switches. The evaluation results show that SA-ATP reduces the iteration time by up to 57% and accelerates training by up to$1.8\times $in real-world benchmark models. Yijun Li 0002, Jiawei Huang 0001, Jingling Liu, Shengwen Zhou, Tao Zhang 0019, Wanchun Jiang, Jianxin Wang 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2024 | Taming the Aggressiveness of Heterogeneous TCP Traffic in Data Center NetworksabstractTo achieve low latency and high link utilization, ECN-based transport protocols (i.e., DCTCP) are widely deployed in data center networks (DCN). In multi-tenant environment, however, the newly introduced ECN-enabled TCP greatly impairs the performance of applications with out-dated and misconfigured TCP stacks. The reason is that the ECN-enabled switch fails to treat the mixed TCP traffic fairly, resulting in the distinguished performance gap between the ECN-enabled and ECN-disabled TCPs. This paper proposes DDT (Dual Dynamic Thresholds), an active queue management algorithm (AQM) to achieve the flow-level fairness for coexisting heterogeneous TCP traffic. DDT monitors the switch queue in real time, and dynamically tunes the distance between ECN-marking and packet-dropping thresholds to mitigate the aggressiveness difference between the ECN-enabled and ECN-disabled TCPs. The results of real implementations and large-scaled simulations show that DDT elegantly fills the aggressiveness gap of heterogeneous TCP traffic without disturbing their own control loops, while only introducing acceptable deployment overhead at switch. Tao Zhang 0019, Jiawei Huang 0001, Shaojun Zou, Chang Ruan, Kai Chen 0005, Jianxin Wang 0001, Geyong Min |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | PA-ATP: Progress-Aware Transmission Protocol for In-Network AggregationabstractLarge-scale machine learning typically adopts distributed machine learning (DML) techniques to accelerate model training. Due to the large communication overhead, unfortu-nately, the phase of gradient aggregation has become the performance bottleneck for DML. To reduce traffic volume, several in-network aggregation (INA) transmission protocols are proposed to offload gradient aggregation function into the programmable switches. However, since existing INA transmission protocols use synchronous congestion control mechanism to drive each round of gradient aggregation, the straggling workers lead to long iteration time and significant performance degradation. To solve the above problem, we propose PA-ATP, a progress-aware INA transmission protocol, which adopts the progress-aware asynchronous congestion control. PA-ATP adjusts the sending rate in accordance with the transmission progress, allowing the straggling flow to grab more bandwidth than the leading flow and control the asynchronous degree of straggling job. We use a P4 programmable switch and a kernel-bypass protocol stack to implement PA-ATP. The results of testbed and large-scale NS3 simulations show that PA-ATP reduces training time by up to 62% compared to the state-of-the-art INA transmission protocols. Jiawei Huang 0001, Tao Zhang 0019, Shengwen Zhou, Qile Wang, Yijun Li 0002, Jingling Liu, Wanchun Jiang, Jianxin Wang 0001 |
ICNP | 3 |
| 2023 | Load Balancing With Deadline-Driven Parallel Data Transmission in Data Center NetworksabstractWith the explosive growth of the Internet of Things (IoT), an increasing amount of sensor data generated by soft real-time IoT applications has been moved to data centers for storage and data analysis. Large amounts of these data are required to be processed within a given deadline to ensure application performance. Therefore, meeting the transmission deadlines of data flows for soft real-time applications has always been crucial yet challenging to current data centers. Recent progress has demonstrated that adopting parallel data transmission over multipath data center network combining with effective load balancing can achieve a high bisection network bandwidth, thus speeding up the network transfer of data flows. Nevertheless, the deadline miss ratios (DMRs) of these flows are not lowered as expected since the existing load balancing schemes are naturally agnostic to the deadline requirement. They are either unable to reroute traffic flexibly or aimlessly reroute these deadline-restrained flows, regardless of their urgent levels and path conditions. To address these inefficiencies, we propose a deadline-aware load-balancing scheme, namely, DLB, which perceives the deadline requirements and helps the urgent flows to timely switch to those faster transmission paths to complete quickly. Specifically, DLB computes the urgent level for each flow in real time to judge if the switch needs to make proactive rerouting. When a flow is nonurgent, DLB does not proactively change its transmission path, leaving more available paths to those flows with higher urgent levels. When a flow becomes extremely urgent, it immediately switches to those light-loaded paths to finish its data transmission before its deadline as far as possible. Experimental results of NS2 simulations and real testbed implementations show that DLB reduces the DMRs by up to 50% compared to the state-of-the-art data center load-balancing schemes, while only induces trivial overhead during deployment. Tao Zhang 0019, Yuanzhen Hu, Yangfan Li 0001, Shaojun Zou, Qianqiang Zhang, Chang Ruan |
IEEE Internet Things J. | 1 |
| 2023 | Toward Communication-Efficient Digital Twin via AI-Powered Transmission and ReconstructionabstractDigital twin technology has recently gathered pace in engineering communities as it allows for the convergence of the real structure and its digital counterpart. 3D point cloud data is a more effective way to describe the real world and to reconstruct the digital counterpart than the conventional 2D images or 360-degree images. Large-scale, e.g., city-scale digital twins, typically collect point cloud data via internet-of-things (IoT) devices and transmit it over wireless networks. However, the existing wireless transmission technology can not carry real-time point cloud transmission for digital twin reconstruction due to mass data volume, high processing overheads, and low delay-tolerance. We propose a novel artificial intelligence (AI) powered end-to-end framework, termed AIRec, for efficient digital twin communication from point cloud compression, wireless channel coding, and digital twin reconstruction. AIRec adopts the encoder-decoder architecture. In the encoder, a novel importance-aware pooling scheme is designed to adaptively select important points with learnable thresholds to reduce the transmission volume. We also design a novel noise-aware joint source and channel coding is proposed to adaptively adjust the transmission strategy based on SNR and map the features to error-resilient channel symbols for wireless transmission to achieve a good tradeoff between the transmission rate and reconstruction quality. The decoder can accurately reconstruct the digital twins from the received symbols. Extensive experiments of typical datasets and comparison with baselines show that we achieve a good reconstruction quality under$24\times $compression ratio. Cen Chen 0002, Xulei Yang, Joey Tianyi Zhou, Tao Zhang 0019, Yangfan Li 0001 |
IEEE J. Sel. Areas Commun. | 5 |
| 2023 | REN: Receiver-Driven Congestion Control Using Explicit Notification for Data CenterabstractIn recent years, receiver-driven transport protocols have been proposed to use proactive congestion control to meet the stringent latency requirements of large-scale applications in data center. However, the receiver-driven proposals face the challenges brought by network dynamic. First, when the bursty flows start, the aggressive and blind line-rate transmission in the first RTT easily leads to persistent queue backlog. Second, when some flows finish transmissions, the remaining ones cannot increase their sending rates to seize the available bandwidth. To address these problems, this article presents a new receiver-driven congestion control design, called REN, which uses the under- and over-utilization notifications from switch to handle the dynamic traffic. With the aid of explicit feedback, REN alleviates the traffic burstiness due to aggressive start, mitigates the conservativeness in utilizing available bandwidth, and still retains the receiver-driven feature to achieve ultra-low latency. We implement the prototype of REN using DPDK. The experimental results of real testbed and large-scale NS2 simulation show that REN effectively reduces the average flow completion time (AFCT) by up to 68% over the state-of-the-art receiver-driven transmission schemes. Jiawei Huang 0001, Jinbin Hu 0001, Weihe Li, Tao Zhang 0019, Jingling Liu, Jianxin Wang 0001, Tian He 0001 |
IEEE Trans. Cloud Comput. | 5 |
| 2022 | Load balancing with traffic isolation in data center networks
Tao Zhang 0019, Qianqiang Zhang, Yasi Lei, Shaojun Zou, Fangmin Li |
Future Gener. Comput. Syst. | 1 |
| 2022 | Opportunistic Transmission for Video Streaming over Wild InternetabstractThe video streaming system employs adaptive bitrate (ABR) algorithms to optimize a user’s quality of experience. However, it is hard for ABR algorithms to choose the right bitrate consistently under highly dynamic bandwidth fluctuations in wild Internet. In this article, we propose a building block on the client side named Opportunistic Chunk Replacement Mechanism (OCRM) to help existing ABR algorithms make full use of the available bandwidth to improve the network utilization and viewing experience of users. Specifically, the servers take advantages of the spare bandwidth to opportunistically transmit high-quality chunks (called opportunistic chunks ) with low priority to the client, without incurring any extra delay. Then, the client player replaces the low-quality chunks with the opportunistic ones that have high quality. We compare OCRM with state-of-the-art ABR algorithms by using trace-driven experiments spanning a wide variety of quality of experience metrics and network conditions. The test results show that OCRM effectively achieves high network utilization and improves the user’s viewing experience by up to 35%. Jiawei Huang 0001, Qichen Su, Weihe Li, Zhuoran Liu 0003, Tao Zhang 0019, Sen Liu 0002, Ping Zhong 0002, Wanchun Jiang, Jianxin Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | GTCP: Hybrid Congestion Control for Cross-Datacenter NetworksabstractTo improve the quality of experience for worldwide users, an increasing number of service providers deploy their services on geographically dispersed data centers, which are connected by wide area network (WAN). In the cross-datacenter networks, however, the intra- and inter-datacenter parts have different characteristics, including switch buffer depth, round-trip time and bandwidth. Besides, most of intra-DC flows belong to interactive services that require low delay while inter-DC flows typically need to achieve high throughput. Unfortunately, existing sender-based and receiver-driven transport protocols do not consider the network heterogeneity between inter- and intra- DC networks so that they fail to simultaneously achieve low latency for intra-DC flows and high throughput for inter-DC flows. This paper proposes a general hybrid congestion control mechanism called GTCP to address this problem. When the inter-DC flow detects congestion inside data center, it switches to the receiver-driven mode to avoid the impact on intra-DC flows. Otherwise, it switches back to the sender-based mode to proactively explore the available bandwidth. Besides, the intra-DC flow leverages the pausing mechanism to eliminate the queue build-up. Through a series of testbed experiments and large-scale NS2 simulations, we demonstrate that GTCP reduces flow completion time by up to 79.3% compared with existing protocols. Shaojun Zou, Jiawei Huang 0001, Jingling Liu, Tao Zhang 0019, Jianxin Wang 0001 |
ICDCS | 4 |
| 2020 | Polo: Receiver-Driven Congestion Control for Low Latency over Commodity Network FabricabstractRecently, numerous novel transport protocols are proposed for the low latency of applications deployed in data center networks, e.g., web search and retail recommendation system. The state-of-art receiver-driven protocols, e.g., Homa and NDP, show the superior performance for achieving the lowest possible latency. However, Homa assumes that the core layer in data center network has no congestion, which limits its application for the existing over-subscribed networks. NDP requires to modify the switch hardware since it trims packets to headers when the packets cause the switch buffer to overflow, resulting in high deployment cost. In this paper, we present Polo to realize low latency for flows over commodity network fabric relying on Explicit Congestion Notification (ECN) and priority queues. According to packets with ECN marking, the Polo receiver obtains the congestion information such that it dynamically adjusts the number of data packets in network for maintaining the small switch queue. The adjustment is carried out periodically. The time interval is determined by keeping the extra high priority packet always in flight nor by a fine-grained timer. Further, Polo designs the packet recovery mechanisms to retransmit the lost packets as soon as possible. Simulation experiment results show that Polo outperforms the state-of-art receiver-driven protocols in a wide range of scenarios including incast. Chang Ruan, Jianxin Wang 0001, Wanchun Jiang, Tao Zhang 0019 |
ICPP | 4 |
| 2020 | Rethinking Fast and Friendly Transport in Data Center NetworksabstractThe sustainable growth of bandwidth has been an inevitable tendency in current Data Center Networks (DCN). However, the dramatic expansion of link capacity offers a remarkable challenge to the transport layer protocols of DCN, i.e., how to converge fast and enable data flow to utilize the high bandwidth effectively. Meanwhile, the new protocol should be compatible to the traditional TCP because the applications with old TCP versions are still widely deployed. Therefore, it is important to achieve a trade-off between the aggressiveness and TCP-friendliness in protocol design. In this article, we first empirically investigate why the existing typical data center TCP variants naturally fail to guarantee both fast convergence and TCP friendliness. Then, we design a new transport protocol for DCN, namely Fast and Friendly Converging (FFC), which makes independent decisions and self-adjustment through retrieving the two-dimensional congestion notification from both RTT and ECN. We further present a mathematic model to analyze its competing behavior and converging process. The results from simulation experiments and real implementation show that FFC can achieve fast convergence, thus benefiting the flow completion time. Moreover, when coexisting with the traditional TCP, FFC also presents a moderate behavior, while introducing trivial deployment overhead only at the end-hosts. Tao Zhang 0019, Jiawei Huang 0001, Kai Chen 0005, Jianxin Wang 0001, Jianer Chen, Yi Pan 0001, Geyong Min |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | DDT: Mitigating the Competitiveness Difference of Data Center TCPsabstractTo achieve better network performance, the cloud service providers are widely deploying the ECN-based transport protocols (i.e., DCTCP) in their data center networks (DCN). In multi-tenant environment, however, the newly introduced ECN-enabled TCP greatly impairs the performance of applications with out-dated and miscon figured TCP stacks. The reason is that the ECN-enabled datacenter switch fails to treat the mixed TCP traffic fairly, causing the distinguished performance gap between the ECN-enabled and ECN-disabled TCPs. This paper proposes DDT (Dual Dynamic Thresholds), an active queue management algorithm (AQM) that aims to achieve the flow-level fairness when the heterogeneous TCP traffic coexists. DDT monitors the switch queue in real time, and dynamically tunes the distance between ECN-marking and packet-dropping thresholds to mitigate the competitiveness difference between the ECN-enabled and ECN-disabled TCP. Our preliminary real implementations and testing results show that DDT elegantly fills the competitiveness gap of heterogeneous TCP traffic without disturbing their own control loops, while only introducing acceptable deployment overhead at the switch. Tao Zhang 0019, Jiawei Huang 0001, Shaojun Zou, Sen Liu 0002, Jinbin Hu 0001, Jingling Liu, Chang Ruan, Jianxin Wang 0001, Geyong Min |
APNet | 1 |
| 2019 | QoS3: Secure Caching in HTTPS Based on Fine-Grained Trust DelegationabstractWith the ever-increasing concern in network security and privacy, a major portion of Internet traffic is encrypted now. Recent research shows that more than 70% of Internet content is transmitted using HyperText Transfer Protocol Secure (HTTPS). However, HTTPS encryption eliminates the advantages of many intermediate services like the caching proxy, which can significantly degrade the performance of web content delivery. We argue that these restrictions lead to the need for other mechanisms to access sites quickly and safely. In this paper, we introduce QoS3, which is a protocol that can overcome such limitations by allowing clients to explicitly and securely re-introduce in-network caching proxies using fine-grained trust delegation without compromising the integrity of the HTTPS content and modifying the format of Transport Layer Security (TLS). In QoS3, we classify web page contents into two types: (1) public contents that are common for all users, which can be stored in the caching proxies, and (2) private contents that are specific for each user. Correspondingly, QoS3 establishes two separate TLS connections between the client and the web server for them. Specifically, for private contents, QoS3 just leverages the original HTTPS protocol to deliver them, without involving any middlebox. For public contents, QoS3 allows clients to delegate trust to specific caching proxy along the path, thereby allowing the clients to use the cached contents in the caching proxy via a delegated HTTPS connection. Meanwhile, to prevent Man-in-the-Middle (MitM) attacks on public contents, QoS3 validates the public contents by employing Document object Model (DoM) object-level checksums, which are delivered through the original HTTPS connection. We implement a prototype of QoS3 and evaluate its performance in our testbed. Experimental results show that QoS3 provides acceleration on page load time ranging between 30% and 64% over traditional HTTPS with negligible overhead. Moreover, QoS3 is deployable since it requires just minor software modifications to the server, client, and the middlebox. Abdulrahman Al-Dailami, Chang Ruan, Zhihong Bao, Tao Zhang 0019 |
Secur. Commun. Networks | 4 |
| 2018 | Designing Fast and Friendly TCP to Fit High Speed Data Center NetworksabstractThe dramatic expansion of link capacity in current data center network causes remarkable challenges to the design of new transport layer protocol, that is, how to converge as fast as possible to help data flow effectively utilize the high bandwidth. Meanwhile, the new protocol should be friendly to the traditional TCP because the non-cooperating applications with old TCP versions are widely existing. Therefore, it is important to achieve a trade-off between the aggressiveness and TCP-friendliness in protocol design. In this paper, we first empirically study why the existing typical data center TCP variants naturally fail to guarantee both fast convergence and TCP friendliness. Then, we design FFC, a transport protocol that makes independent decisions and self-adjustment through retrieving the two-dimensional congestion notification from the RTT and ECN. The results of simulation experiments and real implementations show that the fast convergence of FFC leads to the lower flow completion time compared with DX and DCTCP. Meanwhile, when coexisting with the traditional TCP, FFC also presents a moderate competitiveness, while introducing trivial deployment overhead only at the end hosts. Tao Zhang 0019, Jiawei Huang 0001, Jianxin Wang 0001, Jianer Chen, Yi Pan 0001, Geyong Min |
ICDCS | 1 |
| 2017 | Tuning the Aggressive TCP Behavior for Highly Concurrent HTTP Connections in Intra-DatacenterabstractModern data centers host diverse hyper text transfer protocol (HTTP)-based services, which employ persistent transmission control protocol (TCP) connections to send HTTP requests and responses. However, the ON/OFF pattern of HTTP traffic disturbs the increase of TCP congestion window, potentially triggering packet loss at the beginning of ON period. Furthermore, the transmission performance becomes worse due to severe congestion in the concurrent transfer of HTTP response. In this paper, we provide the first extensive study to investigate the root cause of performance degradation of highly concurrent HTTP connections in data center network. We further present the design and implementation of TCP-TRIM, which employs probe packets to smooth the aggressive increase of congestion window in persistent TCP connection and leverages congestion detection and control at end-host to limit the growth of switch queue length under highly concurrent TCP connections. The experimental results of at-scale simulations and real implementations demonstrate that TCP-TRIM reduces the completion time of HTTP response by up to 80%, while introducing little deployment overhead only at the end hosts. Tao Zhang 0019, Jianxin Wang 0001, Jiawei Huang 0001, Jianer Chen, Yi Pan 0001, Geyong Min |
IEEE/ACM Trans. Netw. | 1 |
| 2016 | Tuning the Aggressive TCP Behavior for Highly Concurrent HTTP Connections in Data CenterabstractModern data centers host diverse HTTP-based services, which employ persistent TCP connections to send HTTP requests and responses. However, the ON/OFF pattern of HTTP traffic disturbs the increase of TCP congestion window, potentially triggering packet loss at the beginning of ON period. Furthermore, the transmission performance becomes worse due to severe congestion in the concurrent transfer of HTTP response. In this work, we first reveal that the TCP's aggressive behavior in increasing congestion window causes TCP timeouts and throughput collapse. We further present the design and implementation of TCP-TRIM, which employs probe packets to smooth the aggressive increase of congestion window in persistent TCP connection, and leverages congestion detection and control at end-host to limit the growth of switch queue length under highly concurrent TCP connections. The experimental results of at-scale simulations and real implementations show that TCPTRIM reduces the completion time of HTTP response by up to 80%, while introducing little deployment overhead only at the end hosts. Jiawei Huang 0001, Jianxin Wang 0001, Tao Zhang 0019, Jianer Chen, Yi Pan 0001 |
ICDCS | 3 |
| 2016 | Adaptive marking threshold method for delay-sensitive TCP in data center network
Tao Zhang 0019, Jianxin Wang 0001, Jiawei Huang 0001, Yi Huang 0005, Jianer Chen, Yi Pan 0001 |
J. Netw. Comput. Appl. | 1 |
| 2015 | Adaptive-Acceleration Data Center TCPabstractProviding deadline-sensitive services is a challenge in data centers. Because of the conservativeness in additive increase congestion avoidance, current transmission control protocols are inefficient in utilizing the super high bandwidth of data centers. This may cause many deadline-sensitive flows to miss their deadlines before achieving their available bandwidths. We propose an Adaptive-Acceleration Data Center TCP, A2DTCP, which takes into account both network congestion and latency requirement of application service. By using congestion avoidance with an adaptive increase rate that varies between additive and multiplicative, A2DTCP accelerates bandwidth detection thus achieving high bandwidth utilization efficiency. At-scale simulations and real testbed implementations show that A2DTCP significantly reduces the missed deadline ratio compared to D2TCP and DCTCP. In addition, A2DTCP can co-exist with conventional TCP as well without requiring more changes in switch hardware than D2TCP and DCTCP. Tao Zhang 0019, Jianxin Wang 0001, Jiawei Huang 0001, Yi Huang 0005, Jianer Chen, Yi Pan 0001 |
IEEE Trans. Computers | 1 |