Shize Zhang

dblp:182/6680 · DBLP profile ↗
← Back
34ranked-venue papers
5as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 21 · 2 first-author · 19 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scaling LLM Agent Tool Access at Cloud Scale
abstract
LLM agents increasingly rely on tool calling, and the Model Context Protocol (MCP) standardizes it between agents and tool providers, reducing integration cost and driving rapid growth in tool scale. Yet a standardized interface does not make tool access work at production scale: legacy services are not MCP-callable, fast protocol evolution creates compatibility cost, large tool sets exhaust the context window, and stateful sessions complicate load balancing. We solve these with a shared control point, a centralized MCP Gateway System that makes MCP operational at cloud scale. The gateway breaks the direct-connect data plane and consolidates legacy API integration, protocol bridging, access control, and session-aware routing, while scaling out elastically at low per-call overhead. It scales agent tool access to thousands of cloud operations.
Enge Song, Yueshang Zuo, Rong Wen, Jing Tie, Zhou Shao, Qiang Fu 0011, Xiaobo Xue, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Tian Pan 0001, Yang Song 0031, Xing Li 0007, Biao Lyu, Meng Li 0010, Haipeng Dai 0001, Guihai Chen, Shunmin Zhu
APNet16
2026 Bifrost: Alibaba's Next-Generation VPC Network with High-Performance Multipath Reliable Transport
Xing Li 0007, Bo Jiang 0003, Yilong Lv, Yuke Hong, Yinian Zhou, Junnan Cai, Jiayue Xu, Yunrui Hu, Zhao Gao, Enge Song, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Changgang Zheng, Yang Song 0031, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu
NSDI19
2026 CStar Gateway: Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Tian Pan 0001, Jin Ke 0005, Baohai Hu, Changgang Zheng, Enge Song, Donglin Lai, Yisong Qiao, Bengbeng Xue, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Yang Song 0031, Xionglie Wei, Biao Lyu, Rong Wen, Zhigang Zong, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu
NSDI15
2026 ZooRoute: Enhancing Cloud-Scale Network Reliability via Candidate Path Provisioning and Overlay Proactive Rerouting
Xiaoqing Sun, Xing Li 0007, Xionglie Wei, Tian Pan 0001, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Xiaobo Xue, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu
NSDI16
2026 Spillway: Orchestrating DPU and Host into a Unified vSwitching Fabric
abstract
The transition to Data Processing Unit (DPU)-centric architectures has become the de-facto standard in modern cloud networks, enabling infrastructure offload and improved host resource utilization. However, the fixed hardware limits of DPUs increasingly fail to keep pace with the rapid growth of host compute density and network-intensive workloads. As a result, when DPU resources are saturated, host compute capacity often remains underutilized due to insufficient network provisioning.
Xiaochong Jiang, Yilong Lv, Naixuan Guan, Qiming Zhao, Sihan Fu, Xuyang Ge, Denghui Wu, Yibin Shen, Guochun Hong, Yijian Dong, Yiquan Chen, Shaoliang An, Zhixiong Guo, Yisong Qiao, Hongwei Ding 0004, Shize Zhang, Rong Wen, Yang Song 0031, Zhigang Zong, Xing Li 0007, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
SIGCOMM25
2026 FPGA-Friendly Architecture of Processing Elements for Efficient and Accurate Quantized CNNs
abstract
An FPGA-friendly processing element based on the small logarithmic floating-point (SLFP) format is proposed. The proposed processing elements not only support inner product but also perform various nonlinear activation functions (NAF), which consume 674× LUT6s and 7× DSPs and operate at 450MHz in a pipeline manner for Zynq-7000. In addition, as the distribution of SLFP numbers is not uniform, this brief revises the weight decay scheme in the quantization aware training process to explore the optimum quantized weights. Compared with INT8 based design, the proposed method balances the resource usage between lookup tables and digital signal processing blocks. The accuracy loss of the quantized model based on the 8-bit SLFP is also small due to the high dynamic range of SLFP format. Moreover, since the proposed method can support different NAFs, this brief improves the quantized model accuracy by selecting an appropriate NAF from Swish, GELU, Mish and PReLU. Compared to the baseline (parameters are FP32, NAF is ReLU), the accuracy of quantized ResNet-50 and MobileNet is increased by 2.65% and -0.33%.
Botao Xiong, Shize Zhang, Xingyu Shao, Xintong He, Yuchun Chang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 Understanding the Long Tail Latency of TCP in Large-Scale Cloud Networks
Enge Song, Bo Jiang 0003, Yang Song 0031, Yuke Hong, Yilong Lv, Yinian Zhou, Junnan Cai, Chao Wang 0128, Yi Wang 0004, Yehao Feng, Shize Zhang, Xiaoqing Sun, Jianyuan Lu, Xing Li 0007, Biao Lyu, Zhigang Zong, Shunmin Zhu
APNet15
2025 Hermes: Enhancing Layer-7 Cloud Load Balancers with Userspace-Directed I/O Event Notification
abstract
Layer-7 load balancers (L7 LBs) improve service performance, availability, and scalability in public clouds. They rely on I/O event notification mechanisms such as epoll to dispatch connections from the kernel to userspace workers. However, early epoll versions suffered from the thundering herd problem. Epoll exclusive (available since Linux 4.5) mitigates this but introduces LIFO wakeups, causing connection concentration on a few workers. Reuseport (Linux 3.9) hashes connections across workers but suffers from hash collisions and lacks awareness of worker load. Since each worker serves multi-tenant traffic, inter-worker load balancing is critical to avoid worker overload and preserve tenant performance isolation.
Tian Pan 0001, Enge Song, Yueshang Zuo, Shaokai Zhang, Yang Song 0031, Jiangu Zhao, Wengang Hou, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu
SIGCOMM10
2025 Nezha: SmartNIC-based Virtual Switch Load Sharing
abstract
Cloud providers use SmartNIC-accelerated virtual switches (vSwitches) to offer rich network functions (NFs) for tenant VMs. Constrained by limited SmartNIC resources, it is a challenge to provide sufficient network performance for high-demand VMs. Meanwhile, we observed a significant number of idle vSwitches in the data center, which led us to consider leveraging them to build a remote resource pool for high-demand virtual NICs (vNICs). In this work, we propose Nezha, a distributed vSwitch load sharing system. Nezha reuses the existing idle SmartNICs to handle the excess load from the local SmartNIC without adding new devices. Nezha offloads stateless rule/flow tables to the remote, while keeping states locally. This eliminates the need for state synchronization, facilitating load sharing and failover. The deployment cost of Nezha is only a small fraction of that required to deploy new devices. Data collected from production show that our CPS capability bottleneck has shifted from the vSwitch to the VM kernel stack, with #concurrent flows and #vNICs increased by up to 50.4x and 40x, respectively.
Xing Li 0007, Enge Song, Tian Pan 0001, Qiang Fu 0011, Yang Song 0031, Yilong Lv, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Zhigang Zong, Qinming He, Shunmin Zhu
SIGCOMM11
2025 Albatross: A Containerized Cloud Gateway Platform with FPGA-accelerated Packet-level Load Balancing
abstract
Alibaba Cloud's centralized gateways relied heavily on high-capacity switching ASICs, but the abrupt halt of Tofino chip evolution in Jan 2023 forced us to seek alternatives that can meet the requirements of performance, supply-chain security, code reuse, and resource efficiency. After evaluating multiple options, we developed Albatross, our 3rd gen cloud gateway based on FPGA and x86 CPUs. Albatross delivers FPGA-based packet-level load balancing to the host CPUs to prevent CPU core overload, manages large reorder buffers under high-latency jitters (100μs) during complex cloud service processing, and resolves head-of-line (HOL) blocking from packet losses or software exceptions in CPUs. To avoid being overloaded by heavy hitters due to anomalies or attacks, it also implements a two-stage rate limiter for millions of tenants with only 2MB of FPGA memory. To maximize resource utilization, Albatross uses containerization to host multiple gateway instances and designs a BGP proxy to lessen the BGP peering overhead on uplink switches caused by high-density container deployments. After hundreds of man-months of development, a single Albatross node can process 80~120Mpps of cloud network traffic with an average latency of 20μs, reducing gateway and sandbox infra costs by 50%.
Jianyuan Lu, Shunmin Zhu, Tian Pan 0001, Yisong Qiao, Yang Song 0031, Wenqiang Su, Yanqiang Li, Enge Song, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Xing Li 0007
SIGCOMM12
2025 ZooRoute: Enhancing Cloud-Scale Network Reliability via Overlay Proactive Rerouting
abstract
This paper presents ZooRoute, a tenant-transparent, fast failure recovery service that requires no modifications to physical devices. ZooRoute leverages the overlay layer and enables traffic flows to bypass failures by altering source ports (srcPorts) in packet headers during encapsulation. To enable deployment in large-scale cloud networks, ZooRoute proposes: 1) On-demand probing to efficiently monitor a vast number of hosts while minimizing telemetry costs. 2) Table compression to record the states of numerous paths with limited on-chip resources. 3) A device-sensing mechanism to prevent unnecessary reconnections in stateful forwarding. Deployed in Alibaba Cloud for 18 months, ZooRoute has significantly improved network reliability, reducing cumulative outage time by 92.71%.
Xiaoqing Sun, Xionglie Wei, Xing Li 0007, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Tian Pan 0001, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu
SIGCOMM14
2025 Design of Low-Cost and High-Accurate 8-bit Logarithmic Floating-Point Arithmetic Circuits
abstract
Recent studies suggest that the 8-bit floating-point (FP) format plays an important role in the deep learning, where the$E4M3$(4-bit exponent, 3-bit mantissa) is suited for the natural language processing model and the$E3M4$is better on computer vision task. In this brief, the logarithmic number system (LNS) is used to simplify the design of FP8 multipliers and dividers because the multiplication and division can be performed by the addition and subtraction in the logarithmic domain. Furthermore, this brief finds that the 3- and 4-bit logarithmic and anti-logarithmic (Antilog) converters can be effectively realized by {x,$x+1$} and {x,$x-1$}. As a result, compared to the standard$E4M3$and$E3M4$multipliers, the cell area can be reduced by 32% and 40%. Compared to the standard$E4M3$and$E3M4$divider, the cell area can be reduced by 61% and 67%. In addition, compared with the INT8-based design, the area of convolution core using proposed multiplier is reduced by 33%. The accuracy loss of the quantized ResNet-50, MobileNet, and ViT-B based on the proposed convolution core are −0.12%, +0.38%, and +0.8%, which are better than the INT8-based design. In the end, the proposed divider can be used in the image change detection. The false rate is slightly reduced from 2.97% to 2.95% compared to the standard$E3M4$divider.
Botao Xiong, Xingyu Shao, Shize Zhang, Yuchun Chang 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2024 CloudPlanner: Minimizing Upgrade Risk of Virtual Network Devices for Large-Scale Cloud Networks
abstract
Cloud networks continuously upgrade softwarized virtual network devices (VNDs) to meet evolving tenant demands. However, such upgrades may result in unexpected failures. An intuitive idea to prevent upgrade failures is to resolve all compatibility issues before deployment, but it is impractical to replicate all deployed VND cases and test them with lots of replayed real traffic for the VND developers. As a result, the operations team takes upgrade risk to test upgrades by gradually deploying them. Although careful upgrade schedule planning is the most common method to minimize upgrade risk, to the best of our knowledge, no VND upgrade schedule planning scheme has been adequately studied for large-scale cloud networks. To fill this gap, we propose CloudPlanner, the first VND upgrade schedule planning scheme aiming to minimize the VND upgrade risk for large-scale cloud networks. CloudPlanner prioritizes upgrading VNDs that are more likely to trigger failures based on expert knowledge and historical failure-trigger VND properties and limits the number of tenants associated with simultaneously upgraded VNDs. We also propose a heuristic solver which can quickly and greedily plan schedules. Using real-world data from production environments, we demonstrate the benefits of CloudPlanner through extensive experiments.
Enhuan Dong, Jiahai Yang 0001, Shize Zhang, Zejie Wang, Xiaoqing Sun, Enge Song, Jianyuan Lu, Biao Lyu, Shunmin Zhu
INFOCOM4
2024 POSEIDON: A Consolidated Virtual Network Controller that Manages Millions of Tenants via Config Tree
Biao Lyu, Enge Song, Tian Pan 0001, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Chenxiao Wang, Xiuheng Chen, Yandong Duan, Weisheng Wang, Jinpeng Long, Kunpeng Zhou, Zhigang Zong, Xing Li 0007, Guangwang Li, Peng Cheng 0001, Jiming Chen 0001, Shunmin Zhu
NSDI5
2024 Proactive Telemetry in Large-Scale Multi-Tenant Cloud Overlay Networks
abstract
At present, public clouds have served millions of tenants. To provide reliable services, cloud vendors need to perceive health status of the cloud network by building a telemetry system to detect possible network failures. While telemetry systems for physical networks have been extensively studied, research on telemetry systems for virtual networks is still insufficient. Different from physical networks, we conclude that building a virtual network telemetry system faces new challenges of feasibility, efficiency, and effectiveness. Specifically, we need to 1) protect privacy of tenants and adapt to heterogeneous middleboxes at the data plane; 2) handle frequent virtual network topology updates and compress large-scale measurement paths for millions of tenants at the control plane; 3) analyze telemetry results to locate network failures at the analysis plane. To address these challenges, we present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. At the data plane, Zoonet uses host agent and arp-ping to protect tenants’ privacy and defines an elegant generalization of ping and traceroute, which can work on heterogeneous middleboxes. At the control plane, Zoonet conducts update batch processing and substantial probing path pruning to lessen the overhead. At the analysis plane, Zoonet reduces noises and aggregates alerts based on temporal and spatial correlation and conducts the hop-by-hop telemetry mode to locate failures. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting.
Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Shize Zhang, Xiaoqing Sun, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Xihui Yang, Jiahai Yang 0001
IEEE/ACM Trans. Netw.5
2024 CloudSentry: Two-Stage Heavy Hitter Detection for Cloud-Scale Gateway Overload Protection
abstract
The cloud vendors provide sharing resources for millions of tenants across the world to achieve economies of scale. At the same time, the cloud network keeps the performance isolation between different tenants as if they use their private dedicated resources. However, heavy hitters caused by a single tenant at cloud gateways will break such isolation, undermining the predictable performance expected by other cloud tenants. To prevent it, heavy hitter detection becomes a key concern at the performance-critical cloud gateways but faces the dilemma between fine granularity and low overhead. In this work, we presentCloudSentry, a scalable two-stage heavy hitter detection system dedicated to multi-tenant cloud gateways against such a dilemma. CloudSentry uses CPU utilization as an indicator of heavy hitters and conducts a lightweight coarse-grained detection running 24/7 to detect such CPU spikes. Then it invokes a fine-grained detection to precisely dump and analyze the potential heavy-hitter packets at the CPU spikes. After that, a more comprehensive analysis is conducted to associate heavy hitters with the cloud service scenarios and invoke a corresponding backpressure procedure. CloudSentry significantly reduces memory, computation and storage overhead compared with existing approaches. In a gateway cluster under an average traffic throughput of 251 Gbps, CloudSentry consumes only a fraction of 2%–5% CPU utilization with 8 KB run-time memory, producing only 10 MB heavy hitter logs during one month. Additionally, as it has been deployed in Alibaba Cloud for over two years, we share case studies and a lot of deployment experiences in this article.
Jianyuan Lu, Tian Pan 0001, Mao Miao, Guangzhe Zhou, Yining Qi, Shize Zhang, Enge Song, Xiaoqing Sun, Huaiyi Zhao, Biao Lyu, Shunmin Zhu
IEEE Trans. Parallel Distributed Syst.7
2024 CouldPin-Fast: Effient and Effective Root Cause Localization for Shared Bandwidth Package Traffic Anomalies in Public Cloud Networks
abstract
As cloud services become increasingly widespread, many public cloud tenants opt for Shared Bandwidth Package (sBwp) services for inbound/outbound communication. The sBwp service allows tenants to purchase shared bandwidth for multiple virtual machines (VMs) instead of buying it individually, which is a convenient and cost-effective traffic management mode. However, the sBwp service presents new challenges for operators to identify the root cause of abnormal sBwp traffic, especially in large-scale, globally distributed public clouds with millions of users. Developing a localization system in public cloud faces several challenges, including dynamic scalability, hyper-scale data efficiently obtaining, and complex application scenarios. To address these challenges, we propose a two-stage localization method calledCloudPin-Fast. First,CloudPin-Fastemploys a cold-start mode to meet dynamic requirements. Second,CloudPin-Fastimplements a pre-filter to reduce the transmission and processing of hyper-scale data. Finally,CloudPin-Fastuses an anomaly localization algorithm based on multi-dimensional statistics fusion in the second stage to cover complex scenarios. The evaluation results on four production datasets have shown superior efficiency and effectiveness. We also share lessons learned from deployingCloudPin-Fastfor over a year in a world-renowned public cloud vendor.
Shize Zhang, Jianyuan Lu, Biao Lyu, Shunmin Zhu, Enhuan Dong, Jiahai Yang 0001
IEEE Trans. Serv. Comput.1
2023 Multi-stage Location for Root-Cause Metrics in Online Service Systems
abstract
The failure of the online service system will seriously affect the user experience and bring huge economic losses. Therefore, the operators usually monitor service-level metrics and machine-level metrics to help quickly find failures, locate root-cause metrics, and reduce MTTR(mean time to repair). Many methods have emerged in recent years to automatically locate root-cause metrics. However, the existing methods cannot meet the requirements of efficiency, accuracy, and ease of deployment at the same time, and are difficult to use in practice. To overcome their limitations, we propose MetricMiner- a multi-stage location method for root-cause metrics in online service systems. Our approach is based on a key observation from numerous real-world cases: root-cause metrics tend to be unique in both the time dimension and the machine dimension. Therefore, we divide the root-cause metrics localization into three stages: first, quickly filter out normal metrics with limited historical data; second, obtain sufficient historical data to eliminate abnormal metrics; finally, according to the clustering of abnormal metrics between machines to sort and locate root-cause metrics. Experimental results on two real-world datasets with 194 cases show that our method can significantly outperform the state-of-the-art methods. Moreover, MetricMiner has been deployed to multiple banking services for more than six months, and we also shared some lessons learned from real deployment.
Wenchi Zhang, Shize Zhang, Kaixin Sui, Enhuan Dong, Jiahai Yang 0001
NOMS4
2023 AutoIoT: Automatically Updated IoT Device Identification With Semi-Supervised Learning
abstract
IoT devices bring great convenience to a person's life and industrial production. However, their rapid proliferation also troubles device management and network security. Network administrators usually need to know how many IoT devices are in the network and whether they behave normally. IoT device identification is the first step to achieving these goals. Previous IoT device identification methods reach high accuracy in a closed environment. But they are not applicable in the continuously changing environment. When new types of devices are plugged in, they cannot update themselves automatically. Besides, they usually rely on supervised learning and need lots of labeled data, which is costly. To solve these problems, we propose a novel IoT device identification model namedAutoIoT, updating itself automatically when new types of devices are plugged in. Besides, it only needs a few labeled data and identifies IoT devices with high accuracy. The evaluation on two public datasets shows thatAutoIoTcan identify new device types only using 1.5$\sim$2.5 hours’ traffic and still have high accuracy after updating. Moreover, it has a better performance than other works when there are only a few labeled data, especially in an environment with scanning traffic.
Linna Fan, Lin He 0004, Yichao Wu, Shize Zhang, Jia Li 0033, Jiahai Yang 0001, Chaocan Xiang, Xiaoqian Ma
IEEE Trans. Mob. Comput.4
2022 Few-shot Learning for Trajectory-based Mobile Game Cheating Detection
abstract
With the emerging of smartphones, mobile games have attracted billions of players and occupied most of the share for game companies. On the other hand, mobile game cheating, aiming to gain improper advantages by using programs that simulate the players' inputs, severely damages the game's fairness and harms the user experience. Therefore, detecting mobile game cheating is of great importance for mobile game companies. Many PC game-oriented cheating detection methods have been proposed in the past decades, however, they can not be directly adopted in mobile games due to the concern of privacy, power, and memory limitations of mobile devices. Even worse, in practice, the cheating programs are quickly updated, leading to the label scarcity for novel cheating patterns. To handle such issues, we in this paper introduce a mobile game cheating detection framework, namely FCDGame, to detect the cheats under the few-shot learning framework. FCDGame only consumes the screen sensor data, recording users' touch trajectories, which is less sensitive and more general for almost all mobile games. Moreover, a Hierarchical Trajectory Encoder and a Cross-pattern Meta Learner are designed in FCDGame to capture the intrinsic characters of mobile games and solve the label scarcity problem, respectively. Extensive experiments on two real online games show that FCDGame achieves almost 10% improvements in detection accuracy with only few fine-tuned samples.
Yueyang Su, Di Yao 0001, Xiaokai Chu, Wenbin Li 0012, Jingping Bi, Runze Wu 0001, Shize Zhang, Jianrong Tao
KDD8
2022 FingFormer: Contrastive Graph-based Finger Operation Transformer for Unsupervised Mobile Game Bot Detection
abstract
This paper studies the task of detecting bots for online mobile games. Considering the fact of lacking labeled cheating samples and restricted available data in the real detection systems, we aim to study the finger operations captured by screen sensors to infer the potential bots in an unsupervised way. In detail, we introduce a Transformer-style detection model, namely FingFormer. It studies the finger operations in the format of graph structure in order to capture the spatial and temporal relatedness between the two hands’ operations. To optimize the model in an unsupervised way, we introduce two contrastive learning strategies to refine both finger moving patterns and players’ operation habits. We conduct extensive experiments under different experimental environments, including the synthetic dataset, the offline dataset, as well as the large-scale online data flow from three mobile games. The multi-facet experiments illustrate the proposed model is both effective and general to detect the bots for different mobile games.
Wenbin Li 0012, Xiaokai Chu, Yueyang Su, Di Yao 0001, Runze Wu 0001, Shize Zhang, Jianrong Tao, Jingping Bi
WWW7
2022 IntStream: Towards Flexible, Expressive, and Scalable Network Telemetry
abstract
Due to the complexity of the network structure and the high growth of the transmission speed, the measurement and management of the network are facing serious challenges. The traditional bottom-up network telemetry methods are no longer applicable to complex network scenarios. To bridge this gap, we propose IntStream, a flexible, expressive and scalable network telemetry framework to allow network operators to measure and analyze network through passive stream processing and active probing. However, there are three key challenges to building an intent-based telemetry system: (1) The diversity of network data sources. (2) The complexity of the measurement tasks. (3) The low overhead requirements of the telemetry system. IntStream introduces a lightweight component to extract and parse data from various data sources and divides the data stream processing into local and global stages. IntStream provides a set of rich expressive primitives to support operators to write telemetry tasks based on intent. By performing part of the telemetry task on the local stage, the transmission overhead of intermediate data can be effectively reduced. The evaluation results conducted on a large campus network show that IntStream can support a wide range of telemetry tasks while reducing the intermediate data transmission overhead by 99.31% on average.
Xin Cheng 0022, Shize Zhang, Jiahai Yang 0001
IEEE Trans. Netw. Serv. Manag.3
2021 MineHunter: A Practical Cryptomining Traffic Detection Algorithm Based on Time Series Tracking
abstract
With the development of cryptocurrencies’ market, the problem of cryptojacking, which is an unauthorized control of someone else’s computer to mine cryptocurrency, has been more and more serious. Existing cryptojacking detection methods require to install anti-virus software on the host or load plug-in in the browser, which are difficult to deploy on enterprise or campus networks with a large number of hosts and servers. To bridge the gap, we propose MineHunter, a practical cryptomining traffic detection algorithm based on time series tracking. Instead of being deployed at the hosts, MineHunter detects the cryptomining traffic at the entrance of enterprise or campus networks. Minehunter has taken into account the challenges faced by the actual deployment environment, including extremely unbalanced datasets, controllable alarms, traffic confusion, and efficiency. The accurate network-level detection is achieved by analyzing the network traffic characteristics of cryptomining and investigating the association between the network flow sequence of cryptomining and the block creation sequence of cryptocurrency. We evaluate our algorithm at the entrance of a large office building in a campus network for a month. The total volumes exceed 28 TeraBytes. Our experimental results show that MineHunter can achieve precision of 97.0% and recall of 99.7%.
Shize Zhang, Jiahai Yang 0001, Xin Cheng 0022, Xiaoqian Ma, Hui Zhang 0052, Bo Wang 0066, Zimu Li
ACSAC1
2021 IntStream: An Intent-driven Streaming Network Telemetry Framework
abstract
Due to the complexity of the network structure and the high growth of the transmission speed, the measurement and management of the network are facing serious challenges. The traditional bottom-up network telemetry methods are no longer applicable to complex network scenarios. To bridge this gap, we propose IntStream, an intent-driven streaming network telemetry framework to allow network operators to measure and analyze network traffic. However, there are three key challenges to building an intent-based telemetry system: (1) The diversity of network data sources. (2) The complexity of the measurement tasks. (3) The low overhead requirements of the telemetry system. IntStream introduces a lightweight component to extract and parse data from various types of data sources to form a data stream and divides the data stream conversation process into local and global stages. IntStream provides a set of rich expressive primitives to support users to write telemetry tasks based on intent. By performing part of the telemetry task on the local stage, the transmission overhead of intermediate data can be effectively reduced. The evaluation results conducted on a large campus network show that IntStream can support a wide range of telemetry tasks while reducing the intermediate data transmission overhead by 99.64% on average.
Xin Cheng 0022, Shize Zhang, Jiahai Yang 0001
CNSM3
2021 Slider: Towards Precise, Robust and Updatable Sketch-based DDoS Flooding Attack Detection
abstract
Distributed Denial of Service (DDoS) flooding attacks have been a severe threat to the Internet for decades. These attacks usually are launched by exhausting bandwidth, network resources or server resources. Since most of these attacks are launched abruptly and severely, it is crucial to develop an efficient DDoS flooding attack detection system. In this paper, we present Slider, an online sketch-based DDoS flooding attack detection system. Slider utilizes a new type of sketch structure, namely Rotation Sketch, to effectively detect DDoS flooding attacks and efficiently identify the malicious hosts. Meanwhile, Slider also learns the characteristics of the current network during the time specified by the network operator to periodically update the parameters of its detection model. We have developed a prototype of Slider and the evaluation results on real-world traffic and public DDoS/DoS attack datasets demonstrate that Slider can effectively detect various DDoS flooding attacks with high precision and robustness.
Xin Cheng 0022, Shize Zhang, Jia Li 0033, Jiahai Yang 0001
GLOBECOM3
2021 Unsupervised IoT Fingerprinting Method via Variational Auto-encoder and K-means
abstract
With the rapid growth of the number of IoT devices on the Internet, security problems of IoT devices are becoming more and more serious, which bring more challenges to network administrators. The first task to solve these problems for network administrators is being aware of IoT devices in the network. Previous IoT device identification methods typically use supervised machine learning methods, which require a large amount of labeled sample data. However, it is difficult to obtain a large number of labeled samples effectively. In order to address this problem, we propose an unsupervised IoT device fingerprinting method at the network level, which can effectively cluster IoT devices without labeled samples. We deeply analyze the temporal and spatial dimension characteristics of network traffic, which can adequately reflect the differences between different IoT devices. By using these features, we develop a clustering framework based on variational autoencoder and K-means algorithms. We conduct evaluation experiments on a public dataset including 24 different IoT devices. The experimental results show that our clustering algorithm can achieve accuracy of 86.7% outperforming a k-NN based state-of-art supervised approach.
Shize Zhang, Jiahai Yang 0001, Dongbin Bai, Fuliang Li, Zimu Li
ICC1
2021 PINBALL: Universal and Robust Signature Extraction for Smart Home Devices
Chenxin Duan, Shize Zhang, Jiahai Yang 0001, Yang Yang 0004, Jia Li 0033
IM2
2021 CloudPin: A Root Cause Localization Framework of Shared Bandwidth Package Traffic Anomalies in Public Cloud Networks
abstract
Due to the sharing nature of public cloud, most of the cloud services use a sharing bandwidth package (sBwp) model to conduct inbound/outbound communication. The sBwp model allows users to purchase a sharing bandwidth for plenty of virtual machines instead of purchasing bandwidth for each virtual machine separately. The advantage of sBwp is that it can provide users with convenient configuration and lower economic cost. However, the sBwp model brings new challenges for operators to localize the root cause of traffic anomalies of a sharing bandwidth, especially for a globally distributed large-scale public cloud with millions of users. In this paper, we first formalize the sBwp problem on the cloud and propose CloudPin, a root cause localization framework for this problem. Our framework solves all the challenges by employing a multi-dimensional algorithm with three sub-models of prediction deviation, anomaly ampli-tude, and shape similarity, and an overall ranking algorithm. Evaluations on real-world data, from one of the world-renowned public cloud vendors, show that our algorithm precision reaches 97.8% for the top 1 of the ranking list, outperforming multiple baseline algorithms.
Shize Zhang, Jianyuan Lu, Biao Lyu, Shunmin Zhu, Jiahai Yang 0001, Lin He 0004
ISSRE1
2021 Distributed and Adaptive Traffic Engineering with Deep Reinforcement Learning
abstract
Lots of studies focus on distributed traffic engineering (TE) where routers make routing decisions independently. Existing approaches usually tackle distributed TE problems through traditional optimization methods. However, due to the intrinsic complexity of the distributed TE problems, routing decisions cannot be obtained efficiently, which leads to significant performance degradation, especially for highly dynamic traffic. Emerging machine learning technologies like deep reinforcement learning (DRL) provide a new choice to address TE problems in an experience-driven method. In this paper, we propose DATE, a distributed and adaptive TE framework with DRL. DATE distributes well-trained agents to the routers in the located network. Each agent makes local routing decisions independently based on link utilization ratios flooded by each router periodically. To coordinate the distributed agents to achieve the global optimization in different traffic conditions, we construct candidate paths, develop the agents carefully, and realize a virtual environment to train the agents with a DRL algorithm. We do extensive simulations and experiments using real-world network topologies with both real and synthetic traffic traces. The results show that DATE outperforms some existing approaches and yields near-optimal performance with superior robustness.
Nan Geng, Mingwei Xu 0001, Yuan Yang 0001, Chenyi Liu, Jiahai Yang 0001, Qi Li 0002, Shize Zhang
IWQoS7
2020 An IoT Device Identification Method based on Semi-supervised Learning
abstract
With the rapid proliferation of IoT devices, device management and network security are becoming significant challenges. Knowing how many IoT devices are in the network and whether they are behaving normally is significant. IoT device identification is the first step to achieve these goals. Previous IoT identification works mainly use supervised learning and need lots of labeled data. Considering collecting labeled data is time-consuming and cannot be scaled, in this paper, we propose an IoT identification model based on semi-supervised learning. The model can differentiate IoT and non-IoT and classify specific IoT devices based on time interval features, traffic volume features, protocol features and TLS related features. The evaluation in a public dataset shows that our model only needs 5% labeled data and gets accuracy over 99%.
Linna Fan, Shize Zhang, Yichao Wu, Chenxin Duan, Jia Li 0033, Jiahai Yang 0001
CNSM2
2019 Measurement and Analysis of Adult Websites in IPv6 Networks
abstract
The Internet is in the transition from IPv4 to IPv6. At present, researches on IPv6 networks mainly focus on architectural issues, such as routing, addressing, and security; there are few studies on the operational issues of IPv6 networks. Our preliminary observation shows that there are a large amount adult websites and traffic in IPv6 networks. Adult websites can damage health of teenagers and bring operational issues in IPv6 networks. This paper conducts a comprehensive measurement and analysis of the adult websites and traffic in IPv6 networks to help solve these operational issues. The data used in this paper is the raw packet traffic from CNGI-CERNET2 which is a pure IPv6 academic network in China. The duration of the data is from July 2017 to January 2018 and the total amount is 40+ terabytes. We detected about 3000 adult websites in the global IPv6 network. This paper analyzes these adult websites and traffic from the perspectives of websites, users and ISPs respectively. We find that adult websites are still in the developing stage in IPv6 networks and only 30% adult websites with full resources can be accessed in IPv6-only networks. But due to the IPv6-first policy in RFC 4038, adult traffic will continue to migrate to IPv6 networks from IPv4 networks. On the other hand, we find that CDN vendors promote the development of adult websites in IPv6 networks and many adult website owners use muti-domain policies to escape ISPs restricting. Our findings may help ISPs effectively understand adult websites and enhance the restriction of adult content in IPv6 networks.
Shize Zhang, Hui Zhang 0052, Jiahai Yang 0001, Guanglei Song
APNOMS1
2019 MVAN: Multi-view Attention Networks for Real Money Trading Detection in Online Games
abstract
Online gaming is a multi-billion dollar industry that entertains a large, global population. However, one unfortunate phenomenon known as real money trading harms the competition and the fun. Real money trading is an interesting economic activity used to exchange assets in a virtual world with real world currencies, leading to imbalance of game economy and inequality of wealth and opportunity. Game operation teams have been devoting much efforts on real money trading detection, however, it still remains a challenging task. To overcome the limitation from traditional methods conducted by game operation teams, we propose, MVAN, the first multi-view attention networks for detecting real money trading with multi-view data sources. We present a multi-graph attention network (MGAT) in the graph structure view, a behavior attention network (BAN) in the vertex content view, a portrait attention network (PAN) in the vertex attribute view and a data source attention network (DSAN) in the data source view. Experiments conducted on real-world game logs from a commercial NetEase MMORPG( JusticePC) show that our method consistently performs promising results compared with other competitive methods over time and verifiy the importance and rationality of attention mechanisms. MVAN is deployed to several MMORPGs in NetEase in practice and achieving remarkable performance improvement and acceleration. Our method can easily generalize to other types of related tasks in real world, such as fraud detection, drug tracking and money laundering tracking etc.
Jianrong Tao, Jianshi Lin, Shize Zhang, Sha Zhao, Runze Wu 0001, Changjie Fan, Peng Cui 0001
KDD3
2017 Robust regression for anomaly detection
abstract
In our previous work, we have applied ordinary linear regression equation to network anomaly detection. However, the performance of ordinary linear regression equation is susceptible to outliers. Unfortunately, it is almost impossible to obtain a “clean” traffic data set for ordinary regression model due to the burstiness of network traffic and the pervasive network attacks. In this paper, we make use of robust regression techniques to mitigate the impact of outliers in the training data set. The experiment results show that the robust regression based method is more reliable than the ordinary regression based method in the face of outliers.
Ziyu Wang 0007, Jiahai Yang 0001, Shize Zhang
ICC3
2016 Towards online anomaly detection by combining multiple detection methods and Storm
abstract
In this paper, we illustrate the significance and advantage of combining the results of multiple detection methods. We implement these methods as bolts in a Apache Storm cluster which is a famous real-time computation framework. We simulate two kinds of anomalies — one involving large number of small network flows and the other involving small number of large network flows. The experiments show that combining multiple methods outperforms any single detection method from the point of view of statistics. Besides, we observe that all the results are outputted in real time without delay, which means that our detection platform is indeed an effective online system.
Ziyu Wang 0007, Jiahai Yang 0001, Hui Zhang 0052, Shize Zhang, Hui Wang 0011
NOMS5