Peizhuang Cong

dblp:260/2955 · DBLP profile ↗
← Back
12ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-0563-745XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Mico: efficient query scheduling for multi-cloud deployed LLM inference service
Peizhuang Cong, Tong Yang 0003, Yuchao Zhang 0004, Wendong Wang 0003, Ke Xu 0002
Sci. China Inf. Sci.1
2026 SEC: Enabling MLLMs for Low-Latency IoT Video Analysis via Semantic-Aware Edge-Cloud Collaboration
abstract
The rapid proliferation of IoT-enabled cameras has driven increasing demand for low-latency, intelligent video understanding in real-world applications such as smart cities and industrial automation. While Multimodal Large Language Models (MLLMs) offer unprecedented capabilities in semantic reasoning and natural language-based video comprehension, their deployment in latency-sensitive IoT environments remains challenging due to high computational costs and sequential decoding bottlenecks. Moreover, conventional edge-cloud video analysis frameworks often rely on semantic-agnostic frame sampling, leading to information loss or redundant data transmission. In this paper, we proposeSEC, a semantic-aware edge-cloud collaborative framework for efficient and accurate video analysis.SECintroduces a task-aware key frame selection mechanism at the edge to maximize semantic relevance while minimizing bandwidth usage, and a novel adaptive speculative decoding framework with tree-based parallel generation on the cloud to accelerate MLLM inference. Extensive experiments under realistic edge-cloud deployment settings on four video understanding benchmarks demonstrate that the proposedSECframework achieves superior performance, significantly reducing the end-to-end inference latency while improving the accuracy, a rare win-win in latency-critical IoT systems.
Mengyu Yang, Ye Tian 0008, Peizhuang Cong, Lanshan Zhang, Gongli Xi, Song Wang 0006, Wendong Wang 0003
IEEE Internet Things J.3
2026 I2BGP: A Privacy-Preserving Intra-AS State-Assisted Inter-AS Routing Scheme
abstract
BGP is the most widely employed inter-AS routing protocol, connecting millions of ASes worldwide. While it is possible to select the egress for outgoing flows based on administrators’ configurations, such schemes are localized due to the privacy of the intra-AS network state. TheAS_Pathfield of BGP records all crossed ASes, which can be used to prevent routing loops and select paths,i.e., selecting the minimum AS-hop path among available paths. Although this scheme is simple, effective, and offers a certain degree of global perspective, selecting paths at AS granularity ignores the transmission performance within each intra-AS, which may result in selecting non-optimal routing paths. To enable the use of private intra-AS data for inter-AS routing, we proposed a privacy-preserving intra-AS state-assisted inter-AS routing scheme, which can select optimal inter-AS paths without disclosing specific intra-AS state data. Specifically, we added an additional BGP header field to carry path performance features and designed a three-step data masking mechanism to protect intra-AS state data, enabling the selection of inter-AS paths with intra-AS state awareness. I2BGP has been deployed in the Greater Bay Area Future Network and a large-scale network simulator based on real network topologies. The results show that I2BGP outperforms BGP in terms of specified forwarding hops, delay, and bandwidth metrics.
Peizhuang Cong, Yuchao Zhang 0004, Jun Wang 0178, Wendong Wang 0003, Tong Yang 0003, Dan Li 0001, Ke Xu 0002
IEEE Trans. Netw.1
2025 BATON: Enhancing Batch-wise Inference Efficiency for Large Language Models via Dynamic Re-batching
abstract
The advanced capabilities of Large Language Models (LLMs) have inspired the development of various interactive web services or applications, such as ChatGPT, which offer query inference services for users. Unlike traditional DNN model, the inference of LLM entails different iterations of forward computation for different queries, which result in efficiency challenges for existing run-to-completion batch-wise inference. Hence, some methods refine batch-wise inference to iteration-level by duplicating all nonlinear layers of LLM. However, this approach not only increases resource usage but also introduces idle computations to the batch due to the prefilling of newly added queries. Therefore, we propose BATON, an efficient batch-wise LLM inference scheme by dynamically adjusting processing batch, which can achieve near-zero idle computations without incurring additional resource consumption. To do so, BATON 1) shapes the vectors involved in the inference of the newly inserted query and processing batch to align dimensions and generates a new attention mask based on vector shaping to ensure inference correctness, which enables query inserting without consuming additional resource; 2) embeds prefilled Keys and Values of the new query into the KV_Cache of the processing batch by leveraging the prefilling and decoding separation mechanism, eliminating idle computations to the batch introduced by the prefilling process of the new query. Experimental results show that compared to the state-of-the-art solution Orca, BATON outperforms improves query processing by up to 1.75x.
Peizhuang Cong, Tong Yang 0003
WWW1
2023 Grandet: Cost-aware Traffic Scheduling without Prior Knowledge in SD-WAN
abstract
The rapid growth of traffic demands on wide-area networks (WANs) has resulted in escalated transmission costs for cross-national enterprises. Many researchers have proposed traffic scheduling methods that can effectively reduce transmission costs and improve network performance. However, the majority of research in this field assumes that traffic demands and network link quality are known in advance, disregarding the impact of information agnostic. While some works try to obtain this knowledge through prediction, they lack awareness of prediction errors, which makes it difficult for their scheduling strategies to achieve theoretical results. In this paper, we propose a novel scheduler Grandet that aims to reduce transmission costs without any prior knowledge. First, instead of requiring prior knowledge or accurate prediction, Grandet determines the intervals of flow sizes and link quality parameters through confidence-based Bootstrap method combined with neural network model, thus quantifying the uncertainty of these information. Then, we design a cost-aware online traffic scheduling framework using the uncertainty intervals from interval determination to optimize the cost minimization problem. Through rigorous theoretical analysis, we prove the approximate optimality of Grandet in minimizing transmission costs. Trace-driven and large-scale simulations show that Grandet successfully reduces transmission costs by over 23%, reduces deadline miss rate by over 31%, and reduces Service Level Agreement (SLA) dissatisfaction rate by over 37%.
Yuchao Zhang 0004, Huahai Zhang, Peizhuang Cong, Wendong Wang 0003, Ke Xu 0002
IWQoS3
2023 DIT and Beyond: Interdomain Routing With Intradomain Awareness for IIoT
abstract
Along with the ever-increasing amount of data generated from industrial devices, the cross domain [also known as autonomous systems (ASs)] data transmission problem has attracted more and more attention in the Industrial Internet of Things (IIoT). As mature and widely used interdomain routing protocols, border gateway protocol-based solutions often take the number of domains (i.e., AS hops) of each path as a criterion to make routing decisions, which is simple and effective. However, such protocols can only meet the reachability requirements while ignoring the performance requirements. That is, the path with the minimum AS hops will be selected to carry flows, even if the actual performance of this path does not meet the transmission requirements due to the unawareness of intradomain information on that path. But it is not impractical to directly access intradomain information for making better routing decisions given data privacy concerns. In this article, we propose M-DIT, which can make interdomain routing decisions with the assistance of desensitized intradomain information for multiple-requirement transmissions. To do so, we design a homomorphic encrypted-based private number comparison scheme to export intradomain information securely and, thus, assist in routing decisions. The results of some experiments based on five real topologies (ATMnet,Claranet,Compuserve,NSFnet, andPeer1) with thousands of interdomain flows demonstrate that M-DIT reduced flow completion time by about 60% or selected high bandwidth paths flexibly for interdomain routing for IIoT scenarios.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Xiangyang Gong, Tong Yang 0003, Dan Li 0001, Ke Xu 0002
IEEE Internet Things J.1
2022 Break the Blackbox! Desensitize Intra-domain Information for Inter-domain Routing
abstract
Along with the ever-increasing amount of data generated from edge networks, cross domain (also known as Autonomous Systems, AS) transmission problem has attracted more and more attention. As mature and widely used inter-domain routing protocols, BGP-based solutions often use the number of domains (i.e. AS hops) of each path to make inter-domain routing decisions, which is simple and effective, but usually can not get the optimal routing results due to the lack of real state/information within ASes. These protocols choose the path with less AS hops as the forwarding path, even if the total latency or cost of the domains on this path is higher. While to solve this problem, directly access to intra-domain information as the assistance to make routing decisions is impractical due to data privacy.In this paper, we propose DIT, which makes near-optimal inter-domain routing decisions with desensitized intra-domain information. To do so, we design a homomorphic encrypted-based private number comparison scheme to export intra-domain information securely and thus assist in routing decisions. We conduct a series of experiments according to five real network topologies with nearly 900 simulated flows, and the results show that DIT reduces the number of forwarding hops by about 45% in average and reduces flow completion time by about 60%.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Xiangyang Gong, Tong Yang 0003, Dan Li 0001, Ke Xu 0002
IWQoS1
2022 A&B: AI and Block-Based TCAM Entries Replacement Scheme for Routers
abstract
With the ever-increasing deployment of 5G and IoT, the number of end-hosts/terminals is increasing rapidly, so that routers have to cache more and more forwarding entries to guarantee communication reachability of these terminals, which makes Ternary Content Addressable Memory (TCAM)-based routers keep expanding resource requirements. However, the design and implementation of large-capacity TCAM-based routers are faced with such challenges: difficult circuit design, high production cost and energy consumption, thereby posing an urgent requirement on a lightweight TCAM that can still maintain those massive communication connections. In this paper, we aim to design a lightweight router with small storage requirement while still retaining the original communication connection performance, which is not straightforward due to the following two challenges: First, under the condition of massive sequential flow data, it’s difficult to accurately and timely select the entries to cache for a small capacity TCAM. Second, given the strict prefix matching principle, how to efficiently insert the selected entries into TCAM is also challenging. To address these problems, we propose A&B: an AI-based Routing entry prediction strategy (AIR) and a Block-based entry Insertion Tactic (BIT). AIR can precisely select entries by conducting accurate entry predictions, which converts dynamic flow-based prediction into stable and parallelizable entry-based prediction by decoupling spatio-temporal characteristics. BIT optimizes entry insertion by isolating TCAM into several blocks, thus eliminating the time-consuming entry movements. The experiment results based on real backbone traffic show that our lightweight A&B achieves comparable performance compared to the traditional schemes by using only 1/8 TCAM storage.
Peizhuang Cong, Yuchao Zhang 0004, Bin Liu 0001, Wendong Wang 0003, Zehui Xiong, Ke Xu 0002
IEEE J. Sel. Areas Commun.1
2021 A Deep Reinforcement Learning-based Routing Scheme with Two Modes for Dynamic Networks
abstract
With the development of communication and transmission technologies, more and more applications, like Internet of vehicles and tele-medicine, become more sensitive to network latency and accuracy, which requires routing schemes to be more efficient. In order to meet such urgent need, learning-based routing strategies emerges, with the advantages of high flexibility and accuracy. These strategies can be divided into two categories, centralized and distributed, enjoying the advantages of high precision and high efficiency, respectively. However, routing become more complex in dynamic network, where the link connections and access states are time-varying, so these learning-based routing mechanisms are required to be able to adapt to network changes in real time. In this paper, we designed and implemented both two of centralized and distributed reinforcement learning-based routing schemes (RLR-T). By conducting a series of experiments, we deeply analyzed the results and gave the conclusion that the centralized is better to cope with dynamic networks due to its faster reconvergence, while the distributed is better to handle with large-scale networks by its high scalability.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Ke Xu 0002, Ruidong Li 0001, Fuliang Li
ICC1
2021 AIR: An AI-based TCAM Entry Replacement Scheme for Routers
abstract
Ternary Content Addressable Memory (TCAM) is an important hardware used to store route entries in routers, which is used to assist routers to make fast decision on forwarding packets. In order to cope with the explosion of route entries due to massive IP terminals brought by 5G and the Internet of Things (IoT), today’s commercial TCAM has to keep the corresponding growth in capacity. But large TCAM capacity is causing many problems such as circuit design difficulties, production costs, and high energy consumption, so it is urgent to design a lightweight TCAM with small capacity while still maintains the original query performance.Designing such a TCAM faces two fundamental challenges. Firstly, it is essential to accurately predict the incoming flows in order to cache correct entries in limited TCAM capacity, but prediction on aggregated time-sequential data is challenging in the massive IoT scenarios. Secondly, the prediction algorithm needs to be real-time as the lookup process is in line-rate. In order to address the above two challenges, in this paper, we proposed a lightweight AI-based solution, called AIR, where we successfully decoupled the route entries and designed a parallel-LSTM prediction method. The experiment results under real backbone traffic showed that we successfully achieved comparable query performance by using just 1/8 TCAM size.
Yuchao Zhang 0004, Peizhuang Cong, Bin Liu 0001, Wendong Wang 0003, Ke Xu 0002
IWQoS2
2021 A deep reinforcement learning-based multi-optimality routing scheme for dynamic IoT networks
Peizhuang Cong, Yuchao Zhang 0004, Zheli Liu, Thar Baker, Hissam Tawfik, Wendong Wang 0003, Ke Xu 0002, Ruidong Li 0001, Fuliang Li
Comput. Networks1
2021 DND: Driver Node Detection for Control Message Diffusion in Smart Transportations
abstract
Along with the development of IoT and mobile edge computing in recent years, smart transportation holds great potential to improve road safety and efficiency. The network that carries smart transportation service is highly dynamic. Controllability has long been recognized as one of the fundamental properties of such temporal networks, which can provide valuable insights for the construction of new infrastructures, and thus is in urgent need to be explored. In this article, under the smart transportation scenario, we first disclose the controllability problem in Internet of Vehicles (IoV), and then design DND (Driver Node Detection) algorithm based on Kalman's controllability rank condition to analyze the controllability and control message diffusion in such a dynamic temporal network. Moreover, we use the control message diffusion efficiency as a metric to assist in selecting suitable driver nodes. At last, we conduct a series of experiments to analyze the controllability of the IoV network, and the results show the effects of vehicle density, speed, coverage radius on network controllability, and the efficiency of the control message diffusion algorithm and its feedback effect on driver nodes selection. These insights are critical for varieties of applications in the future smart transportation.
Peizhuang Cong, Yuchao Zhang 0004, Wendong Wang 0003, Ning Zhang 0007
IEEE Trans. Netw. Serv. Manag.1