Huaiyi Zhao

dblp:221/1581 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0002-1818-2549ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2024 QuarkTable: Building Compact Forwarding Tables for Programmable Switches on Public Clouds
abstract
Programmable switches have been recently proposed as dataplane solutions for public clouds. However, the conflict of limited on-chip memory and massive forwarding rules in cloud networks hinders the large-scale deployments. We argue that building compact forwarding tables for programmable switches is a viable option to this problem. In this paper, as a first step, we explore the feasibility of compact data structures for VPC routing tables (VRTs) and propose QuarkTable as a solution. The idea of QuarkTable builds upon the existence of redundancy in the prefixes of VRTs, supported by extensive analysis of real-world VRTs collected from six geographically distributed regions of Alibaba Cloud. By cutting the VRT into two partitions and encoding the upper part of the longer prefixes, QuarkTable effectively shortens the length of each entry and reduces the overall memory consumption. We demonstrate the effectiveness of QuarkTable by showing its proximity to the entropy bound of real-world VRTs. Experiments on six VRTs from Alibaba Cloud show practical memory savings up to 33.5% and 30.2% for SRAM and TCAM, respectively.
Jianyuan Lu, Huaiyi Zhao, Yehao Feng, Shengru Li, Enge Song, Xionglie Wei, Biao Lyu, Rong Wen, Shunmin Zhu
APNet2
2024 TAR: Traffic Adaptive IPv6 Routing Lookup Scheme
abstract
IP lookup aims at identifying the longest prefix match within routing tables to determine the forwarding path for packets. The increase in IPv6 address length makes IP lookup particularly difficult, requiring more efficient lookup mechanisms to handle the expanded address space and ensure the timely and accurate routing of packets. Existing IPv6 lookup algorithms focus on constructing efficient data structures to enhance lookup performance. However, given the inherent characteristics of network traffic distribution and the varying hit probabilities of nodes within the lookup structure, there is still room for optimization. Our study introduces a new IPv6 routing lookup scheme that utilizes the characteristics of network traffic to guide the construction of lookup data structures. Experimental results indicate that the lookup performance of the traffic-adaptive algorithm is 1.1 to 2.2 times that of conventional algorithms. Additionally, our work designs a traffic-adaptive routing system, which includes an adaptive cache structure capable of responding to dynamic changes in network traffic. The throughput of our proposed system is 1.7 to 2.8 times that of systems implementing only basic algorithms.
Xinyi Zhang 0004, Huaiyi Zhao, Yanbiao Li 0001, Gaogang Xie
APNet3
2024 CloudSentry: Two-Stage Heavy Hitter Detection for Cloud-Scale Gateway Overload Protection
abstract
The cloud vendors provide sharing resources for millions of tenants across the world to achieve economies of scale. At the same time, the cloud network keeps the performance isolation between different tenants as if they use their private dedicated resources. However, heavy hitters caused by a single tenant at cloud gateways will break such isolation, undermining the predictable performance expected by other cloud tenants. To prevent it, heavy hitter detection becomes a key concern at the performance-critical cloud gateways but faces the dilemma between fine granularity and low overhead. In this work, we presentCloudSentry, a scalable two-stage heavy hitter detection system dedicated to multi-tenant cloud gateways against such a dilemma. CloudSentry uses CPU utilization as an indicator of heavy hitters and conducts a lightweight coarse-grained detection running 24/7 to detect such CPU spikes. Then it invokes a fine-grained detection to precisely dump and analyze the potential heavy-hitter packets at the CPU spikes. After that, a more comprehensive analysis is conducted to associate heavy hitters with the cloud service scenarios and invoke a corresponding backpressure procedure. CloudSentry significantly reduces memory, computation and storage overhead compared with existing approaches. In a gateway cluster under an average traffic throughput of 251 Gbps, CloudSentry consumes only a fraction of 2%–5% CPU utilization with 8 KB run-time memory, producing only 10 MB heavy hitter logs during one month. Additionally, as it has been deployed in Alibaba Cloud for over two years, we share case studies and a lot of deployment experiences in this article.
Jianyuan Lu, Tian Pan 0001, Mao Miao, Guangzhe Zhou, Yining Qi, Shize Zhang, Enge Song, Xiaoqing Sun, Huaiyi Zhao, Biao Lyu, Shunmin Zhu
IEEE Trans. Parallel Distributed Syst.10
2023 RecMon: A Deep Learning-based Data Recovery System for Network Monitoring
abstract
Network monitoring systems struggle with the issue that the measurement data is incomplete, with only a subset of origin-destination (OD) pairs or time slots observed, due to the high deployment and measurement cost. Recent studies show that the missing data can be inferred from partial measurements using neural network models and tensor methods. However, these recovery approaches fail to achieve accuracy, adaptability and high speed, simultaneously. In this paper, we propose RecMon, a deep learning-based data recovery system that satisfies the above three criteria. A global spatio-temporal attention mechanism and a data augmentation algorithm are proposed to improve the recovery accuracy. A semi-supervised learning-based scheme is devised for fast and effective model updates. We conduct extensive experiments on three real-world datasets to compare RecMon with four state-of-the-art methods in terms of online recovery performance. The experimental results show that RecMon can adapt to the latest state of the network and accurately recover network measurement data in less than 100 milliseconds. When 90% of the data is missing, the recovery accuracy of RecMon improves over the strongest baseline method by 22.7%, 16.0%, and 8.2% in the three datasets, respectively.
Huaiyi Zhao, Xinyi Zhang 0004, Kun Xie 0001, Dong Tian, Gaogang Xie
INFOCOM1
2023 Improving the Scalability of Distributed Network Emulations: An Algorithmic Perspective
abstract
By deploying virtualized network elements (hosts, switches, routers, links, etc.) on clusters of commodity machines, distributed network emulations (DNE) closely mimic the behaviors of network systems and provide real-time interactions and analysis for network service management. However, DNE encounters scalability challenges when faced with large network topologies. These challenges can be boiled down to the assignment problem: to which physical machine each virtualized network element should be assigned so that the largest possible network topology can be emulated? In this paper, we tackle this problem from an algorithmic perspective. We first propose TBR (topology balancing relaxation) as the relaxation of the assignment problem. TBR tries to maintain a balance of the hardware resource consumption, by minimizing the maximum inter-machine bandwidth. We further develop TBS (topology balancing solver), which combines mathematical techniques with multi-level algorithms to solve TBR efficiently. We integrate TBR and TBS into MaxiNet, a famous distributed network emulator. Experimental results show that with the same available physical resources, TBR and TBS can improve emulation scalability by up to$4.7\times $compared to baselines.
Huaiyi Zhao, Xinyi Zhang 0004, Yang Wang 0147, Zulong Diao, Yanbiao Li 0001, Gaogang Xie
IEEE Trans. Netw. Serv. Manag.1