Chenhao Jia

dblp:209/5418 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-4616-9410ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ZooRoute: Enhancing Cloud-Scale Network Reliability via Candidate Path Provisioning and Overlay Proactive Rerouting
Xiaoqing Sun, Xing Li 0007, Xionglie Wei, Tian Pan 0001, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Xiaobo Xue, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu
NSDI11
2026 Distributed Rate Limiting Under Decentralized Cloud Networks
abstract
The rapid expansion of cloud applications has led to unprecedented increases in network traffic volume, diversity, and complexity. As Cloud Service Providers (CSPs) adopt decentralized, geographically distributed data centers, effective traffic management across these environments has become critical. Distributed Rate Limiting (DRL) has emerged as an essential tool to manage the complex traffic dynamics of decentralized networks, yet traditional centralized rate limiting methods fall short, facing limitations in scalability, adaptability to bursty traffic, and efficiency. This paper presents C3PDAR (Cloud Control with Constant Probabilities and Dynamic Adjustment Range), a novel DRL algorithm tailored for decentralized cloud infrastructures. C3PDAR introduces three key innovations: (1) CPS-BPS DualPoint Rate Limiting and Parent-Child Token Bucket mechanisms, which effectively mitigate burst traffic and short-lived connections while improving bandwidth fairness and inter-tenant isolation; (2) A vSwitch-CGW Cascade Rate Limiting architecture, which reduces CPU overhead in CGW clusters and accelerates convergence by 42%–78%; (3) Virtual Extensible Local Area Network (VXLAN) Padding scheme, which embeds rate-limiting information in existing traffic instead of transmitting new data packets, reducing the communication overhead of the C3PDAR algorithm by over 40%. By integrating these advancements, C3PDAR delivers a scalable, robust solution that outperforms traditional DRL approaches in performance, fault tolerance, and resource efficiency. C3PDAR uniquely empowers CSPs to manage complex, high-volume traffic dynamics in decentralized cloud environments, offering both theoretical insights and practical optimizations for next-generation network control.
Tianyu Xu 0007, Lilong Chen, Xiaochong Jiang, Liming Ye, Yilong Lv, Chenhao Jia, Yongwang Wu, Zhigang Zong, Xing Li 0007, Bingqian Lu, Shunmin Zhu, Chengkun Wei, Wenzhi Chen
IEEE Trans. Mob. Comput.9
2025 ZooRoute: Enhancing Cloud-Scale Network Reliability via Overlay Proactive Rerouting
abstract
This paper presents ZooRoute, a tenant-transparent, fast failure recovery service that requires no modifications to physical devices. ZooRoute leverages the overlay layer and enables traffic flows to bypass failures by altering source ports (srcPorts) in packet headers during encapsulation. To enable deployment in large-scale cloud networks, ZooRoute proposes: 1) On-demand probing to efficiently monitor a vast number of hosts while minimizing telemetry costs. 2) Table compression to record the states of numerous paths with limited on-chip resources. 3) A device-sensing mechanism to prevent unnecessary reconnections in stateful forwarding. Deployed in Alibaba Cloud for 18 months, ZooRoute has significantly improved network reliability, reducing cumulative outage time by 92.71%.
Xiaoqing Sun, Xionglie Wei, Xing Li 0007, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Tian Pan 0001, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu
SIGCOMM10
2024 An Automatic Search Method for 4-Bit Optimal S-Boxes Towards Considering Cryptographic Properties and Hardware Area Simultaneously
Chenhao Jia, Sijia Gong, Ting Wu 0001, Tingting Cui
Inscrypt (2)2
2024 FTA-detector: Troubleshooting Gray Link Failures Based on Fault Tree Analysis
abstract
Detecting link failures is critical to ensuring the operation of data center networks (DCNs). However, some gray link failures may go undetected by switches, leading to silent packet drops. In this paper, we propose FTA-detector, a gray link failure detection and localization approach leveraging Fault Tree Analysis (FTA), a technique previously applied in the field of reliability engineering. On the data plane, we collect fine-grained hop-by-hop information through In-band Network Telemetry (INT), detect the bidirectional connectivity of end-to-end paths through a novel aging mechanism, and implement fast reroute in response to gray link failures. On the control plane, we introduce a faulty link localization algorithm based on FTA to recommend the most likely faulty links. Specifically, we use Top K and progressive failure repair to discover and repair link faults as early as possible during failure inference, significantly reducing the overall computation complexity of sequential root cause analysis. For large-scale network topology, we propose a divide and conquer optimization scheme for scalability. To verify the efficiency of our system, we build a virtual network test platform with P4 switch software and Redis database. The test results show that FTA-detector can troubleshoot multi-point failures in DCNs in a very short time with high accuracy.
Yan Zou, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Qingqiang Yi, Ying Wan 0001, Jiao Zhang 0002, Tao Huang 0005
NOMS4
2024 Structure attack on full-round DBST
Chenhao Jia, Ting Wu 0001, Tingting Cui
Frontiers Comput. Sci.1
2024 Proactive Telemetry in Large-Scale Multi-Tenant Cloud Overlay Networks
abstract
At present, public clouds have served millions of tenants. To provide reliable services, cloud vendors need to perceive health status of the cloud network by building a telemetry system to detect possible network failures. While telemetry systems for physical networks have been extensively studied, research on telemetry systems for virtual networks is still insufficient. Different from physical networks, we conclude that building a virtual network telemetry system faces new challenges of feasibility, efficiency, and effectiveness. Specifically, we need to 1) protect privacy of tenants and adapt to heterogeneous middleboxes at the data plane; 2) handle frequent virtual network topology updates and compress large-scale measurement paths for millions of tenants at the control plane; 3) analyze telemetry results to locate network failures at the analysis plane. To address these challenges, we present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. At the data plane, Zoonet uses host agent and arp-ping to protect tenants’ privacy and defines an elegant generalization of ping and traceroute, which can work on heterogeneous middleboxes. At the control plane, Zoonet conducts update batch processing and substantial probing path pruning to lessen the overhead. At the analysis plane, Zoonet reduces noises and aggregates alerts based on temporal and spatial correlation and conducts the hop-by-hop telemetry mode to locate failures. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting.
Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Shize Zhang, Xiaoqing Sun, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Xihui Yang, Jiahai Yang 0001
IEEE/ACM Trans. Netw.7
2024 INT-Label: Lightweight In-Band Network-Wide Telemetry via Distributed Labeling
abstract
In-band Network Telemetry (INT) enables hop-by-hop device-internal state exposure for maintaining and troubleshooting data center networks. To achievenetwork-widetelemetry coverage, orchestration on top of the INT primitive is required. A straightforward solution would flood the network with INT probe packets for maximum measurement coverage, which leads to a huge bandwidth overhead. A refined solution leverages the SDN controller to collect the network topology information and carry out centralized probing path planning, which, however, is inefficient in reacting to topology changes. To tackle the above problems, we proposeINT-label, a lightweight In-band Network-Wide Telemetry architecture via the distributed labeling approach. INT-label periodically labels the sampled packets with device-internal states. It is cost-effective with a minor bandwidth overhead and able to seamlessly adapt to topology changes. In order to reduce the number of labeled packets, we introduce a times-based probabilistic labeling algorithm, which allows fewer packets to carry more INT information than the interval-based algorithm. In addition, to counteract the degradation of telemetry resolution due to loss of labeled packets, we design a feedback mechanism which can adaptively change the instant labeling frequency. We provide theoretical proof that INT-label can achieve network-wide telemetry. We analyze the impact of transmission delay on coverage rate and labeling times distribution under the INT-label architecture. Evaluation on software P4 switches suggests that INT-label can achieve 99.72% measurement coverage under the labeling frequency of 20 times per second. With the adaptive labeling enabled, even if 60% of the packets are lost, the coverage can still reach 92%.
Enge Song, Tian Pan 0001, Haoyu Song 0001, Qiang Fu 0011, Yingjiang Liu, Chenhao Jia, Chuanying Yuan, Minglan Gao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001
IEEE Trans. Parallel Distributed Syst.6
2022 Zoonet: a proactive telemetry system for large-scale cloud networks
abstract
We present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. The requirements are to (1) cover hyper-scale virtual networks with millions of tenants and millions of VMs for top tenants; (2) handle frequent virtual topology changes due to tenants' configuration through flexible APIs; (3) adapt to heterogeneous middleboxes along the probing paths; (4) achieve VM-to-VM telemetry without breaking tenant privacy; (5) differentiate virtual and physical network problems. We argue existing physical network telemetry solutions fail to satisfy our needs due to either incomplete telemetry coverage or outrageous telemetry overhead. Zoonet sets an ambitious goal to provide VM-to-VM hop-by-hop telemetry for each tenant, which is achieved based on self-developed, customizable middleboxes via hundreds of person-months under close team collaboration. At the data plane, Zoonet defines an elegant generalization of ping and traceroute, but made to work on multi-tenant clouds with heterogeneous middleboxes. At the control plane, Zoonet conducts substantial probing path pruning and update batch processing to lessen the overhead. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting.
Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Jiahai Yang 0001
CoNEXT5
2022 MIMIC: SmartNIC-aided Flow Backpressure for CPU Overloading Protection in Multi-Tenant Clouds
abstract
In multi-tenant clouds, off-the-shelf x86 boxes are widely deployed as middleboxes. With the rapid growth of cloud traffic and the migration to NFV deployment in recent years, CPU overloading at middleboxes becomes more of an issue. From our data centers, we observed that the CPU overloading was caused by heavy hitters. To address this issue, we propose MIMIC, a cloud-scale flow backpressure system, implemented onto our existing SmartNIC with FPGA acceleration. MIMIC rate-limits the selected heavy hitters through a new per-flow backpressure protocol and a new heavy-hitter detection system, to protect the other tenants. The detection system is based on hierarchical memory design, leveraging on-chip SRAM and off-chip DRAM, which can handle highly concurrent cloud traffic without the losses of flow information. We extend the design by adding a pre-filtering procedure for rapid detection. To avoid CPU being flooded by FPGA through frequent heavy-hitter reporting, due to their performance disparity, the CPU queries the FPGA on demand. The backpressure protocol is non-invasive to protect tenant privacy and allows controllable rate-limiting through the novel use of ECN and meter tables. The SmartNIC acts as a man in the middle to facilitate heavy-hitter detection and per-flow backpressuring. In a production setting, we observe that MIMIC can react quickly and bring down CPU load to the normal level within 10ms without packet losses.
Enge Song, Nianbing Yu, Tian Pan 0001, Qiang Fu 0011, Xionglie Wei, Yisong Qiao, Jianyuan Lu, Yijian Dong, Mingxu Xie, Jinkui Mao, Zhengjie Luo, Chenhao Jia, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Shunmin Zhu
ICNP14
2022 WebQMon.ai: Gateway-Based Web QoE Assessment Using Lightweight Neural Networks
Enge Song, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001
ICSOC4
2021 INT-probe: Lightweight In-band Network-Wide Telemetry with Stationary Probes
abstract
Visibility is essential for operating and troubleshooting intricate networks. In-band Network Telemetry (INT) has been embedded in the latest merchant silicons to offer high-precision device and traffic state visibility. INT is actually an underlying technique and each INT instance covers only one monitoring path. The network-wide measurement coverage therefore requires a high-level orchestration to provision multiple INT paths. An optimal path planning is expected to produce a minimum number of paths with a minimum number of overlapping links. Eulerian trail has been used to solve the general problem. However, in production networks, the vantage points where one can deploy probes to start and terminate INT paths are constrained. In this work, we propose an optimal path planning algorithm, INT-probe, which achieves the network-wide telemetry coverage under the constraint of stationary probes. INT-probe formulates the constrained path planning into an extended multi-depot k-Chinese postman problem (MDCPP-set) and then reduces it to a solvable minimum weight perfect matching problem. We analyze algorithm's theoretical bound and the complexity. Extensive evaluation on both wide area networks and data center networks with different scales and topologies are conducted. We show INT-probe is efficient, high-performance, and practical for real-world deployment. For a large-scale data center networks with 1125 switches, INT-probe can generate 112 monitoring paths (reduced by 50.4 %) by allowing only 1.79% increase of the total path length, promptly resolving link failures within 744.71ms.
Tian Pan 0001, Xingchen Lin, Haoyu Song 0001, Enge Song, Zizheng Bian, Hao Li 0011, Jiao Zhang 0002, Fuliang Li, Tao Huang 0005, Chenhao Jia, Bin Liu 0001
ICDCS10
2021 INT-label: Lightweight In-band Network-Wide Telemetry via Interval-based Distributed Labelling
abstract
The In-band Network Telemetry (INT) enables hop-by-hop device-internal state exposure for reliably maintaining and troubleshooting data center networks. For achieving network-wide telemetry, orchestration on top of the INT primitive is further required. One straightforward solution is to flood the INT probe packets into the network topology for maximum measurement coverage, which, however, leads to huge bandwidth overhead. A refined solution is to leverage the SDN controller to collect the topology and carry out centralized probing path planning, which, however, cannot seamlessly adapt to occasional topology changes. To tackle the above problems, in this work, we propose INT-label, a lightweight In-band Network-Wide Telemetry architecture via interval-based distributed labelling. INT-label periodically labels device-internal states onto sampled packets, which is cost-effective with minor bandwidth overhead and able to seamlessly adapt to topology changes. Furthermore, to avoid telemetry resolution degradation due to loss of labelled packets, we also design a feedback mechanism to adaptively change the instant label frequency. Evaluation on software P4 switches suggests that INT-label can achieve 99.72% measurement coverage under a label frequency of 20 times per second. With adaptive labelling enabled, the coverage can still reach 92% even if 60% of the packets are lost in the data plane.
Enge Song, Tian Pan 0001, Chenhao Jia, Wendi Cao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001
INFOCOM3
2021 Sailfish: accelerating cloud-scale multi-tenant multi-service gateways with programmable switches
abstract
The cloud gateway is essential in the public cloud as the central hub of cloud traffic. We show that horizontal scaling of software gateways, once sustainable for years, is no longer future-proof facing the massive scale and rapid growth of today's cloud. The root cause is the stagnant performance of the CPU core, which is prone to be overloaded by heavy hitters as traffic growth goes far beyond Moore's law. To address this, we propose \emph{Sailfish}, a cloud-scale multi-tenant multi-service gateway accelerated by programmable switches. The new challenge is that large forwarding tables due to multi-tenancy cannot be fit into the limited on-chip memories. To this end, we devise a multi-pronged approach with (1) hardware/software co-design for table sharing, (2) horizontal table splitting among gateway clusters, (3) pipeline-aware table compression for a single node. Compared with the x86 gateway of a similar price, Sailfish reduces latency by 95% (2μs), improves throughput by more than 20x in bps (3.2Tbps) and 71x in pps (1.8Gpps) with packet length < 256B. Sailfish has been deployed in Alibaba Cloud for more than two years. It is the first P4-based cloud gateway in the industry, of which a single cluster carries dozens of Tbps traffic, withstanding peak-hour traffic in large online shopping festivals.
Tian Pan 0001, Nianbing Yu, Chenhao Jia, Jianwen Pi, Yisong Qiao, Jianyuan Lu, Enge Song, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu
SIGCOMM3
2020 INT-filter: Mitigating Data Collection Overhead for High-Resolution In-band Network Telemetry
abstract
In-band Network Telemetry (INT) enables fine-grained network monitoring to ease the management of large-scale networks, which, however, relies on the real-time collection of a huge amount of telemetry data through the southbound interface. For example, the INT telemetry data upload rate of a 28-pod FatTree topology reaches 3Tbps under a probe frequency of 100 times/s, which is rather unacceptable since the controller-switch link bandwidth is limited. To mitigate the telemetry data collection overhead, in this work, we propose INT-filter, a novel measurement architecture that deploys the same prediction algorithm on both the data plane and the control plane to predict the traffic state in the near future instead of uploading all the telemetry data. Such prediction-based approach leverages the observation that there is considerable redundancy in the telemetry data sequence. In addition, we design an integration mechanism that conducts predictions using multiple methods simultaneously and uploads the predicted result from the least-error method to further decrease the upload volume. Extensive evaluation suggests that INT-filter can achieve at least 33.6% data collection decrease under a 10ms probe interval. With prediction integration, the upload reduction can further reach 58.5%.
Enge Song, Tian Pan 0001, Chenhao Jia, Wendi Cao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001
GLOBECOM3
2020 Rapid Detection and Localization of Gray Failures in Data Centers via In-band Network Telemetry
abstract
Network reliability becomes increasingly important in modern data center networks (DCNs). The DCNs are expected to work sustainably under internal failures and assist network operators in troubleshooting them rapidly. However, some network failures will happen silently with packets discarded without producing any explicit notification before causing tremendous damage to the network. To troubleshoot these "gray failures", in this work, we present a rapid gray failure detection and localization mechanism based on the recently proposed In-band Network Telemetry (INT). Specifically, we leverage simplified INT probe packets to conduct network-wide telemetry to help the servers under ToR switches obtain all the feasible paths between sources and destinations. Once a network failure occurs, the affected thus unavailable paths will immediately be detected and flushed out of the path information table at each server by a timeout mechanism. Hence, servers can proactively perform source routing-based fast traffic reroute to avoid massive packet loss and retain uninterrupted quality of experience. At the meantime, all the aged path entries will be uploaded to a remote controller for centralized failure localization by identifying common path elements. To verify the feasibility of our design, we build a virtual network testbed with software P4 switches and a Redis database. Evaluation shows that our system can successfully detect network gray failures and reroute the affected traffic in no time while complete failure localization within only a few seconds.
Chenhao Jia, Tian Pan 0001, Zizheng Bian, Xingchen Lin, Enge Song, Tao Huang 0005, Yunjie Liu 0001
NOMS1
2020 Threshold-oblivious on-line web QoE assessment using neural network-based regression model
abstract
The evaluation of the web‐browsing quality of experience (QoE) is difficult to complete through traditional methods (e.g. deducing formulas or setting thresholds) due to the diversity of websites and their contents. To evaluate web‐browsing QoE through a general way, the authors propose a web QoE evaluation architecture based on machine learning, consisting of two parts: traffic classification sub‐system and QoE prediction sub‐system. When evaluating user experience, traffic classification sub‐system first classifies the packets generated by visiting a website into a flowthrough some fields in the packet header, to model each website separately. The traffic classification accuracy of packets over six websites reaches 96.63%. Then, in the network layer, the traffic metric cumulative traffic volume is generated from the size and arrival time of packets. When a user visits a web page, their regression model predicts the above‐the‐fold time (ATF) and thus QoE. The output of the regression model is an exact ATF value that is mapped to user experience. In addition, reversing input variables further improves the model, which is evaluated on two popular websites. The QoE prediction results of the improved method for 5400 visits are obtained within 0.0975 s, reaching 0.9 .
Enge Song, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Wendi Cao, Tao Huang 0005
IET Commun.5
2019 Concept Extraction and Prerequisite Relation Learning from Educational Data
abstract
Prerequisite relations among concepts are crucial for educational applications. However, it is difficult to automatically extract domain-specific concepts and learn the prerequisite relations among them without labeled data.In this paper, we first extract high-quality phrases from a set of educational data, and identify the domain-specific concepts by a graph based ranking method. Then, we propose an iterative prerequisite relation learning framework, called iPRL, which combines a learning based model and recovery based model to leverage both concept pair features and dependencies among learning materials. In experiments, we evaluated our approach on two real-world datasets Textbook Dataset and MOOC Dataset, and validated that our approach can achieve better performance than existing methods. Finally, we also illustrate some examples of our approach.
Weiming Lu 0001, Jiale Yu, Chenhao Jia
AAAI4