Yingpu Nian

dblp:377/8265 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-0001-9540ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Digital Twin-assisted Optimization of 6G Wireless Networks: Ensuring Deterministic Communication
Yingpu Nian, Bo Yi 0002, Xingwei Wang 0001, Sajal K. Das 0001
Comput. Networks1
2026 SFTRAP: Satisfying Fidelity Threshold Routing and Adaptive Purification for Throughput Maximum in Quantum Network
abstract
The core function of quantum networks is to establish high-fidelity quantum entanglement for long-distance communication. However, the main challenge is to efficiently allocate resources under limited conditions, maximize throughput, satisfy end-to-end (E2E) fidelity requirements, and prevent quantum decoherence caused by inefficient routing algorithms. Current research focuses on optimizing either throughput or fidelity, with a lack of approaches that optimize both simultaneously; furthermore, existing algorithms suffer from high computational complexity. To tackle these challenges, this study proposes a Satisfying Fidelity Threshold Routing and Adaptive Purification Strategy (SFTRAP). SFTRAP maximizes throughput for each request by selecting multiple paths and dynamically choosing links for entanglement purification based on the current state of link resources, thus minimizing throughput loss while satisfying fidelity threshold. The strategy also adaptively adjusts the number of purification rounds according to the fidelity threshold, thereby optimizing the time required for deep purification and enhancing algorithmic efficiency. For multi-request scenarios, SFTRAP employs a priority sorting mechanism that takes into account both path cost and path freedom, which refines request scheduling and path selection to create more efficient request combinations, thus further boosting the overall network throughput. Simulation results indicate that SFTRAP surpasses state-of-the-art methods in terms of both throughput and algorithmic efficiency, highlighting its potential for optimizing resources in quantum networks.
Zhi Wang 0029, Yingpu Nian, Bo Yi 0002, Xingwei Wang 0001, Xinhao Zhou, Jianhui Lv, Geyong Min, Keqin Li 0001
IEEE Trans. Commun.3
2026 XAForward: Accelerating Distributed Large-Scale Language Model Training Through Fast eXpress Data Path
abstract
With the rapid development of Artificial Intelligence Generated Content (AIGC), single data centers are increasingly unable to meet the growing demands for data and computational resources in distributed large-scale language model (LLM) training. In this context, distributed training across heterogeneous data centers has become a necessary choice to enhance computational power and flexibility. However, the networks in heterogeneous data centers are polymorphic, with diverse communication protocols and network architectures. This heterogeneity renders traditional routing devices ineffective in recognizing and processing gradient data. Moreover, frequent copying and excessive parsing of gradient data by routing devices across heterogeneous data centers significantly increase model training time. To address these challenges, we propose XAForward, a method for accelerating distributed LLM in heterogeneous data centers using eXpress Data Path (XDP). Specifically, XAForward introduces a polymorphic-compatible protocol that reconstructs the header of gradient data packets to enable efficient data forwarding across different communication protocols in heterogeneous data centers. Additionally, to accelerate distributed LLM computing and reduce gradient data copying and excessive parsing during training, XAForward leverages kernel-bypass techniques based on XDP for packet processing and kernel-level data forwarding using network index identifiers. Experimental results show that, compared to state-of-the-art methods, XAForward reduces the distributed LLM training time by approximately 35% to 40%.
Yingpu Nian, Baishun Zhou, Zhi Wang 0029, Bo Yi 0002, Xinhao Zhou, Yuan Yang 0001, Xingwei Wang 0001, Geyong Min, Keqin Li 0001
IEEE Trans. Netw. Serv. Manag.1
2026 PAHInA: Precision-Aware Hierarchical In-Network Aggregation for Edge Distributed Training
abstract
The rise of edge intelligence is driving distributed machine learning toward a new paradigm of edge-collaborative computing. To overcome the severe communication bottleneck in this paradigm, In-Network Aggregation is a critical enabling technology. However, its effectiveness is fundamentally undermined by the profound resource heterogeneity of edge networks. Specifically, edge devices, adapting to hardware constraints, operate at varying numerical precisions, leading to significant data inflation as gradients are aggregated. Compounding this, unevenly distributed network resources and traditional, precision-oblivious routing strategies often misallocate critical, high-precision gradients to low-quality paths. This mismatch creates severe network congestion, crippling the efficiency of distributed training. To address this, we propose the Precision-Aware Hierarchical In-Network Aggregation (PAHInA) framework, the first, to our knowledge, to perform routing optimization for in-network aggregation that explicitly considers precision heterogeneity. The core of PAHInA is an intelligent control-plane scheduler that co-optimizes for gradient priority and path cost, dynamically planning the most cost-effective aggregation strategy for each flow. This fine-grained scheduling guarantees that high-priority gradients are routed through premium, low-latency paths, minimizing global communication overhead. On the data plane, we leverage the eXpress Data Path (XDP) for high-performance packet processing to reduce aggregation-induced overhead. Extensive simulations show that, compared to state-of-the-art baselines, PAHInA significantly mitigates network congestion, reducing end-to-end communication time by up to 33% and boosting overall training throughput by approximately 30%.
Yingpu Nian, Bo Yi 0002, Qiang He 0002, Xingwei Wang 0001, Geyong Min, Keqin Li 0001, Sajal K. Das 0001
IEEE Trans. Netw.1
2025 PCSR: A Low-Latency Routing Protocol for Polymorphic Networks in Real-Time Embodied AI
abstract
The rise of Embodied AI, including autonomous robots, cooperative autonomous driving, and augmented reality agents, is driving a deep integration of intelligent systems with the physical world, imposing stringent demands on the underlying network for real-time, low-latency interaction. However, these Embodied AI systems typically operate in complex polymorphic network environments, simultaneously handling heterogeneous identifiers such as content, IP, and geographic location. This causes traditional routing mechanisms to suffer from significant latency overhead and compatibility bottlenecks due to protocol conversion and adaptation, severely limiting the performance and responsiveness of Embodied AI applications. To address this challenge, we propose the Polymorphic Compatible Segment Routing (PCSR) protocol. PCSR adopts an innovative paradigm of decoupling the protocol from the infrastructure. It dynamically maps native protocol semantics to lightweight 8 -byte identifiers and utilizes compatibility logic encapsulated at the packet tail to achieve smooth compatibility with traditional networks without requiring large-scale modification of existing equipment. Furthermore, we built a zero-copy forwarding engine using the kernel eXpress Data Path (XDP) technology to fundamentally optimize data transmission efficiency. Experimental validation on a 10-node heterogeneous testbed shows that PCSR reduces end-toend latency by 32.7% and maintains a high throughput rate in hybrid network environments. This work demonstrates that PCSR provides an efficient and deployable routing solution for latency-critical, cross-domain collaborative services required by Embodied AI.
Yingpu Nian, Bo Yi 0002, Zhi Wang 0029, Yuan Yang 0001, Xingwei Wang 0001, Keqin Li 0001
ICPADS1
2025 Deep Customized Network Slicing and Efficient Routing for IoT Applications in B5G-Enabled Edge Computing Networks
abstract
Beyond 5G-enabled edge computing networking (ECN) will further deploy computing and communication resources to the edge of the networks. Then, edge service demands for Internet of Things (IoT) applications are becoming more and more diverse, while the corresponding routing service capability is limited and not flexible enough to deal with the demands of ECN, which then leads to reducing the inherent routing capability of ECN. It becomes extremely difficult for ECN to support diversified demands and provide diverse IoT applications quickly and flexibly. In this article, we propose a novel and customized deep routing mechanism for IoT applications in ECN, in which the network slicing and deep learning methods are jointly applied and leveraged. First, we design a new ECN architecture that formulates four kinds of network slices to cope with various IoT scenarios, which are eMBB, uRLLC, mMTTC, and backup slices. Second, using these slices, we can customize the ECN environment flexibly, based on which we propose the corresponding routing method for the purpose of fast and efficient service delivery. In particular, the mapping between network slices and the infrastructure is established with the object of maximizing the resource utilization. Then, the routing is designed and customized by using the deep learning model. Lastly, the experimental results show that the deep customized mechanism designed in this article can reduce the average loss rate of the model, decrease the average delay, as well as improve the average resource utilization compared with the existing studies.
Xingchi Chen, Bo Yi 0002, Qing Li 0006, Fa Zhu, Yingpu Nian, Achyut Shankar, Michele Nappi, Amr Tolba
IEEE Internet Things J.5