Changgang Zheng

dblp:289/0081 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0003-1894-722XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scaling LLM Agent Tool Access at Cloud Scale
abstract
LLM agents increasingly rely on tool calling, and the Model Context Protocol (MCP) standardizes it between agents and tool providers, reducing integration cost and driving rapid growth in tool scale. Yet a standardized interface does not make tool access work at production scale: legacy services are not MCP-callable, fast protocol evolution creates compatibility cost, large tool sets exhaust the context window, and stateful sessions complicate load balancing. We solve these with a shared control point, a centralized MCP Gateway System that makes MCP operational at cloud scale. The gateway breaks the direct-connect data plane and consolidates legacy API integration, protocol bridging, access control, and session-aware routing, while scaling out elastically at low per-call overhead. It scales agent tool access to thousands of cloud operations.
Enge Song, Yueshang Zuo, Rong Wen, Jing Tie, Zhou Shao, Qiang Fu 0011, Xiaobo Xue, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Tian Pan 0001, Yang Song 0031, Xing Li 0007, Biao Lyu, Meng Li 0010, Haipeng Dai 0001, Guihai Chen, Shunmin Zhu
APNet18
2026 Integrating AI Clusters into Virtual Private Cloud
abstract
While commodity NIC-based back-end AI networks offer ultra-high intra-cluster bandwidth for distributed training, their limited programmability and on-chip resources hinder the implementation of advanced VPC features such as fine-grained isolation and stateful security policies. Furthermore, access to resources within the VPC needs to be routed through the front-end DPU, which is shared by the scale-up domain. The mismatch between the front-end DPU’s bandwidth and the back-end requirements causes GPU underutilization when intensive VPC communication is required for content recommendation, AIGC, and federated learning workloads. We propose an architecture that decouples complex policy enforcement from high-speed packet forwarding to support VPC semantics on back-end NICs and enable front-end/back-end integration. Evaluations show near-full GPU utilization in our analytical model and 71 μ s P999 extra latency of the first packet, suggesting that commodity hardware can support both high-throughput AI training and flexible VPC features.
Xing Li 0007, Enge Song, Changgang Zheng, Shengyao Gao, Juncheng Xiang, Junnan Cai, Haoxiang Pan, Yang Song 0031, Yilong Lv, Qiang Fu 0011, Zhigang Zong, Shunmin Zhu
APNet5
2026 Single-Core Hotspots on Your VNF? Break Them Up!
abstract
Current NFVs assign packets to CPU cores at flow granularity, where each flow is pinned to a single CPU. This approach is efficient under most scenarios but has exposed limitations when handling elephant flows. These “heavy hitters” overwhelm single cores, creating bottlenecks that affect overall throughput and degrade service quality. As networks scale to higher-speed links and core-rich CPUs, these imbalances become more severe. In this paper, we propose ParaFlowO, an architecture that Parallelizes processing elephant Flows across multiple CPU cores while preserving in-Order delivery. ParaFlowO breaks elephant flows into flowlets and dynamically rotates them across multiple cores. It integrates a lightweight reordering mechanism to preserve packet order and controls parallelism to mitigate contention on shared state. Preliminary evaluations show that ParaFlowO offers a practical solution to mixed-grained parallelism in stateful middleboxes.
Changgang Zheng, Jin Ke 0005, Enge Song, Yilong Lv, Yisong Qiao, Donglin Lai, Bengbeng Xue, Yang Song 0031, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu
APNet1
2026 Bifrost: Alibaba's Next-Generation VPC Network with High-Performance Multipath Reliable Transport
Xing Li 0007, Bo Jiang 0003, Yilong Lv, Yuke Hong, Yinian Zhou, Junnan Cai, Jiayue Xu, Yunrui Hu, Zhao Gao, Enge Song, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Changgang Zheng, Yang Song 0031, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu
NSDI22
2026 CStar Gateway: Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Tian Pan 0001, Jin Ke 0005, Baohai Hu, Changgang Zheng, Enge Song, Donglin Lai, Yisong Qiao, Bengbeng Xue, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Yang Song 0031, Xionglie Wei, Biao Lyu, Rong Wen, Zhigang Zong, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu
NSDI5
2026 Design, Implementation, and Deployment of Multi-Task Neural Networks in Programmable Data-Planes
abstract
The increasing demand for real-time inference on high-volume network traffic has led to the rise of in-network machine learning, where programmable switches execute various models directly in the data-plane at line rate. Effective network management often involves multiple prediction tasks, such as predicting bit rate, flow size, or traffic class; however, existing solutions deploy separate models for each task, placing a significant burden on the data-plane and leading to substantial resource consumption when deploying multiple tasks. To address this limitation, we introduce MUTA, a novel in-network multi-task learning framework that enables concurrent inference of multiple tasks in the data-plane, without exhausting available resources. MUTA builds a multi-task neural network to share feature representations across tasks and introduces a data-plane mapping methodology to fit it within network switches. Additionally, MUTA enhances scalability by supporting distributed deployment, where different layers of a multi-task model can be offloaded across multiple switches. An orchestrator employs multi-objective optimization to determine optimal model placement in multi-path networks. MUTA is deployed on P4 hardware switches, and is shown to reduce memory requirements by ×10.5, while at the same time improving accuracy by up to 9.14% using limited training data, compared with state-of-the-art single-task learning solutions.
Kaiyi Zhang 0005, Changgang Zheng, Nancy Samaan, Ahmed Karmouch, Noa Zilberman
IEEE Trans. Netw. Serv. Manag.2
2025 MUTA: Enabling Multi-Task Neural Network Inference in Programmable Data-Planes
abstract
The need for real-time inference of large volumes of data led to the development of in-network machine learning. Programmable network switches can now execute various machine learning models in the data-plane at line rate. While a stream of data may require several prediction tasks, such as predicting bit rate, flow size, or traffic class, current solutions only support separate models for each task. This places a significant burden on the data-plane and leads to substantial resource consumption when deploying multiple tasks. To solve this problem, we introduce MUTA; a novel in-network multi-task learning solution. MUTA enables executing multiple inference tasks concurrently in the data-plane, without exhausting available resources. It introduces a data-plane mapping methodology to fit non-binarized multi-task neural networks within network switches. MUTA is deployed on P4-based hardware switches, and is shown to reduce memory requirements by × 10.5 and improve accuracy by up to 9.14% using limited training data, compared with state-of-the-art single-task learning solutions.
Kaiyi Zhang 0005, Changgang Zheng, Nancy Samaan, Ahmed Karmouch, Noa Zilberman
HPSR2
2024 In-Network Machine Learning for Real-Time Transaction Fraud Detection
abstract
Machine learning (ML) has become a mainstream approach in the fight against transaction fraud for its intelligence. For financial institutions and businesses, low-latency detection of fraudulent transactions in real-time is highly important as it enables rapid identification and prevention. Concurrently mitigating fraudulent transactions by using ML while also reducing latency remains a challenging endeavor, for which performing inference within programmable network devices offers a potential solution. In this paper, we introduce MIND, conducting ML-based fraud detection within programmable devices. MIND is prototyped on both software and hardware network devices, including BMv2, Intel Tofino, and NVIDIA BlueField-2 DPU, and is evaluated with three publicly available transaction datasets. Experimental results demonstrate that MIND detects transaction fraud in real-time, with a throughput of 6.4 terabits per second and microsecond-scale latency. Compared with server-based solutions, MIND can process over ×800 more transactions per second, along with a latency reduction of over ×1300 per transaction. At the same time, MIND attains 99.94% of server-based benchmarks’ accuracy and 93.66% of their F1-score, exhibiting only marginal degradation in classification performance. Therefore, MIND offers substantial savings in the number of servers, leading to reduced costs and energy consumption, while providing a better customer experience.
Xinpeng Hong, Changgang Zheng, Noa Zilberman
ECAI2
2024 Accelerating Machine Learning for Trading Using Programmable Switches
abstract
High-frequency trading (HFT) employs cutting-edge hardware for rapid decision-making and order execution but often relies on simpler algorithms that may miss deeper market trends. Conversely, lower-frequency algorithmic trading uses machine learning (ML) for better market predictions but higher latency can negate its strategic benefits. To achieve the best of both worlds, we present an in-network ML solution that embeds ML processes into programmable network devices, accelerating feature engineering and extraction as well as ML inference. In this paper, we design and develop a solution that supports both stock mid-price and volatility movement forecasting using commodity switches. Our approach achieves microsecond-scale, ultra-low latency, significantly lowering it by 64% to 97% compared to previous works, while upholding the same level of ML performance as server models. Additionally, by combining network hardware and servers, a hybrid deployment strategy can keep the misclassification rate change below 0.8% relative to the server baseline while processing 49% of the traffic directly on the switch and achieving a 45% average reduction in end-to-end latency.
Xinpeng Hong, Changgang Zheng, Stefan Zohren, Noa Zilberman
ECAI2
2024 Toward Continuous Threat Defense: in-Network Traffic Analysis for IoT Gateways
abstract
The widespread use of IoT devices has unveiled overlooked security risks. With the advent of ultrareliable low-latency communications (URLLCs) in 5G, fast threat defense is critical to minimize damage from attacks. IoT gateways, equipped with wireless/wired interfaces, serve as vital frontline defense against emerging threats on IoT edge. However, current gateways struggle with dynamic IoT traffic and have limited defense capabilities against attacks with changing patterns. In-network computing offers fast machine learning (ML)-based attack detection and mitigation within network devices, but leveraging its capability in IoT gateways requires new continuous learning capability and runtime model updates. In this work, we present P4Pir, a novel in-network traffic analysis framework for IoT gateways. P4Pir incorporates programmable data plane into IoT gateway, pioneering the utilization of in-network ML inference for fast mitigation. It facilitates continuous and seamless updates of in-network inference models within gateways. P4Pir is prototyped in P4 language on raspberry pi and Dell Edge Gateway. With ML inference offloaded to gateway’s data plane, P4Pir’s in-network approach achieves swift attack mitigation and lightweight deployment compared to prior ML-based solutions. Evaluation results using three public data sets show that P4Pir accurately detects and fastly mitigates emerging attacks (>30% accuracy improvement and submillisecond mitigation time). The proposed model updates method allows seamless runtime updates without disrupting network traffic.
Mingyuan Zang, Changgang Zheng, Lars Dittmann, Noa Zilberman
IEEE Internet Things J.2
2024 Federated In-Network Machine Learning for Privacy-Preserving IoT Traffic Analysis
abstract
The expanding use of Internet-of-Things (IoT) has driven machine learning (ML)-based traffic analysis. 5G networks’ standards, requiring low-latency communications for time-critical services, pose new challenges to traffic analysis. They necessitate fast analysis and response, preventing service disruption or security impact on network infrastructure. Distributed intelligence on IoT edge has been studied to analyze traffic, but introduces delays and raises privacy concerns. Federated learning can address privacy concerns, but does not meet latency requirements. In this article, we propose FLIP4: an efficient federated learning-based framework for in-network traffic analysis. Our solution introduces a lightweight federated tree-based model, offloaded and running within network devices. FLIP4 consumes less resources than previous solutions and reduces communication overheads, making it well-suited for IoT edge traffic analysis. It ensures prompt mitigation and minimal impact on services in the presence of false alerts using two approaches (metering and dropping), thereby balancing learning accuracy and privacy requirements.
Mingyuan Zang, Changgang Zheng, Tomasz Koziak, Noa Zilberman, Lars Dittmann
ACM Trans. Internet Techn.2
2024 IIsy: Hybrid In-Network Classification Using Programmable Switches
abstract
The soaring use of machine learning leads to increasing processing demands. As data volume keeps growing, providing classification services with good machine learning performance, high throughput, low latency, and minimal equipment overheads becomes a challenge. Offloading machine learning tasks to network switches can be a scalable solution to this problem, providing high throughput and low latency. However, network devices are resource constrained, and lack support for machine learning functionality. In this paper, we introduce IIsy -a novel mapping tool of machine learning classification models to off-the-shelf switches. Using an efficient encoding algorithm, enables fitting a range of classification models on switches, co-existing with standard switch functionality. To overcome resource constraints, adopts a hybrid approach for ensemble models, running a small model on a switch and a large model on the backend. The evaluation shows that achieves near-optimal classification results, within minimum resource overheads, and while reducing the load on the backend by 70% for data-intensive use cases.
Changgang Zheng, Zhaoqi Xiong, Thanh T. Bui, Siim Kaupmees, Riyad Bensoussane, Antoine Bernabeu, Shay Vargaftik, Yaniv Ben-Itzhak, Noa Zilberman
IEEE/ACM Trans. Netw.1
2023 LOBIN: In-Network Machine Learning for Limit Order Books
abstract
Machine learning is driving the evolution of algorithmic trading, but the demands for fast execution speed remain. Although both aim to increase profitability, embedding more powerful machine learning approaches and lowering trading latencies are hard to achieve simultaneously. Offloading machine learning inference to programmable network devices, also referred to as in-network machine learning, provides a delicate balance between the two ends of this trade-off. In this paper, we present LOBIN, providing machine learning based market prediction using high-frequency market data feeds. LOBIN builds limit order books and conducts inference within programmable switches. Compared with server-based solutions, LOBIN predicts future stock price movements with lower latency, higher throughput, and a minor impact on machine learning performance.
Xinpeng Hong, Changgang Zheng, Stefan Zohren, Noa Zilberman
HPSR2
2022 Reducing variations in multi-center Alzheimer's disease classification with convolutional adversarial autoencoder
Cobbinah Bernard Mawuli, Christian Sorg, Qinli Yang, Arvid Ternblom, Changgang Zheng, Wei Han 0009, Liwei Che, Junming Shao
Medical Image Anal.5
2022 Event-driven temporal models for explanations - ETeMoX: explaining reinforcement learning
abstract
Abstract Modern software systems are increasingly expected to show higher degrees of autonomy and self-management to cope with uncertain and diverse situations. As a consequence, autonomous systems can exhibit unexpected and surprising behaviours. This is exacerbated due to the ubiquity and complexity of Artificial Intelligence (AI)-based systems. This is the case of Reinforcement Learning (RL), where autonomous agents learn through trial-and-error how to find good solutions to a problem. Thus, the underlying decision-making criteria may become opaque to users that interact with the system and who may require explanations about the system’s reasoning. Available work for eXplainable Reinforcement Learning (XRL) offers different trade-offs: e.g. for runtime explanations, the approaches are model-specific or can only analyse results after-the-fact. Different from these approaches, this paper aims to provide an online model-agnostic approach for XRL towards trustworthy and understandable AI. We present ETeMoX, an architecture based on temporal models to keep track of the decision-making processes of RL systems. In cases where the resources are limited (e.g. storage capacity or time to response), the architecture also integrates complex event processing, an event-driven approach, for detecting matches to event patterns that need to be stored, instead of keeping the entire history. The approach is applied to a mobile communications case study that uses RL for its decision-making. In order to test the generalisability of our approach, three variants of the underlying RL algorithms are used: Q-Learning, SARSA and DQN. The encouraging results show that using the proposed configurable architecture, RL developers are able to obtain explanations about the evolution of a metric, relationships between metrics, and were able to track situations of interest happening over time windows.
Juan Marcelo Parra-Ullauri, Antonio García-Domínguez, Nelly Bencomo, Changgang Zheng, Zhen Chen 0025, Juan Boubeta-Puig, Guadalupe Ortiz 0001, Shufan Yang
Softw. Syst. Model.4
2021 Modular neural network via exploring category hierarchy
Wei Han 0009, Changgang Zheng, Rui Zhang 0070, Jinxia Guo, Qinli Yang, Junming Shao
Inf. Sci.2