Yilong Lv

dblp:291/5892 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 14 · 14 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Integrating AI Clusters into Virtual Private Cloud
abstract
While commodity NIC-based back-end AI networks offer ultra-high intra-cluster bandwidth for distributed training, their limited programmability and on-chip resources hinder the implementation of advanced VPC features such as fine-grained isolation and stateful security policies. Furthermore, access to resources within the VPC needs to be routed through the front-end DPU, which is shared by the scale-up domain. The mismatch between the front-end DPU’s bandwidth and the back-end requirements causes GPU underutilization when intensive VPC communication is required for content recommendation, AIGC, and federated learning workloads. We propose an architecture that decouples complex policy enforcement from high-speed packet forwarding to support VPC semantics on back-end NICs and enable front-end/back-end integration. Evaluations show near-full GPU utilization in our analytical model and 71 μ s P999 extra latency of the first packet, suggesting that commodity hardware can support both high-throughput AI training and flexible VPC features.
Xing Li 0007, Enge Song, Changgang Zheng, Shengyao Gao, Juncheng Xiang, Junnan Cai, Haoxiang Pan, Yang Song 0031, Yilong Lv, Qiang Fu 0011, Zhigang Zong, Shunmin Zhu
APNet14
2026 Single-Core Hotspots on Your VNF? Break Them Up!
abstract
Current NFVs assign packets to CPU cores at flow granularity, where each flow is pinned to a single CPU. This approach is efficient under most scenarios but has exposed limitations when handling elephant flows. These “heavy hitters” overwhelm single cores, creating bottlenecks that affect overall throughput and degrade service quality. As networks scale to higher-speed links and core-rich CPUs, these imbalances become more severe. In this paper, we propose ParaFlowO, an architecture that Parallelizes processing elephant Flows across multiple CPU cores while preserving in-Order delivery. ParaFlowO breaks elephant flows into flowlets and dynamically rotates them across multiple cores. It integrates a lightweight reordering mechanism to preserve packet order and controls parallelism to mitigate contention on shared state. Preliminary evaluations show that ParaFlowO offers a practical solution to mixed-grained parallelism in stateful middleboxes.
Changgang Zheng, Jin Ke 0005, Enge Song, Yilong Lv, Yisong Qiao, Donglin Lai, Bengbeng Xue, Yang Song 0031, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu
APNet9
2026 Bifrost: Alibaba's Next-Generation VPC Network with High-Performance Multipath Reliable Transport
Xing Li 0007, Bo Jiang 0003, Yilong Lv, Yuke Hong, Yinian Zhou, Junnan Cai, Jiayue Xu, Yunrui Hu, Zhao Gao, Enge Song, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Changgang Zheng, Yang Song 0031, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu
NSDI6
2026 Spillway: Orchestrating DPU and Host into a Unified vSwitching Fabric
abstract
The transition to Data Processing Unit (DPU)-centric architectures has become the de-facto standard in modern cloud networks, enabling infrastructure offload and improved host resource utilization. However, the fixed hardware limits of DPUs increasingly fail to keep pace with the rapid growth of host compute density and network-intensive workloads. As a result, when DPU resources are saturated, host compute capacity often remains underutilized due to insufficient network provisioning.
Xiaochong Jiang, Yilong Lv, Naixuan Guan, Qiming Zhao, Sihan Fu, Xuyang Ge, Denghui Wu, Yibin Shen, Guochun Hong, Yijian Dong, Yiquan Chen, Shaoliang An, Zhixiong Guo, Yisong Qiao, Hongwei Ding 0004, Shize Zhang, Rong Wen, Yang Song 0031, Zhigang Zong, Xing Li 0007, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
SIGCOMM3
2026 Distributed Rate Limiting Under Decentralized Cloud Networks
abstract
The rapid expansion of cloud applications has led to unprecedented increases in network traffic volume, diversity, and complexity. As Cloud Service Providers (CSPs) adopt decentralized, geographically distributed data centers, effective traffic management across these environments has become critical. Distributed Rate Limiting (DRL) has emerged as an essential tool to manage the complex traffic dynamics of decentralized networks, yet traditional centralized rate limiting methods fall short, facing limitations in scalability, adaptability to bursty traffic, and efficiency. This paper presents C3PDAR (Cloud Control with Constant Probabilities and Dynamic Adjustment Range), a novel DRL algorithm tailored for decentralized cloud infrastructures. C3PDAR introduces three key innovations: (1) CPS-BPS DualPoint Rate Limiting and Parent-Child Token Bucket mechanisms, which effectively mitigate burst traffic and short-lived connections while improving bandwidth fairness and inter-tenant isolation; (2) A vSwitch-CGW Cascade Rate Limiting architecture, which reduces CPU overhead in CGW clusters and accelerates convergence by 42%–78%; (3) Virtual Extensible Local Area Network (VXLAN) Padding scheme, which embeds rate-limiting information in existing traffic instead of transmitting new data packets, reducing the communication overhead of the C3PDAR algorithm by over 40%. By integrating these advancements, C3PDAR delivers a scalable, robust solution that outperforms traditional DRL approaches in performance, fault tolerance, and resource efficiency. C3PDAR uniquely empowers CSPs to manage complex, high-volume traffic dynamics in decentralized cloud environments, offering both theoretical insights and practical optimizations for next-generation network control.
Tianyu Xu 0007, Lilong Chen, Xiaochong Jiang, Liming Ye, Yilong Lv, Chenhao Jia, Yongwang Wu, Zhigang Zong, Xing Li 0007, Bingqian Lu, Shunmin Zhu, Chengkun Wei, Wenzhi Chen
IEEE Trans. Mob. Comput.8
2025 Understanding the Long Tail Latency of TCP in Large-Scale Cloud Networks
Enge Song, Bo Jiang 0003, Yang Song 0031, Yuke Hong, Yilong Lv, Yinian Zhou, Junnan Cai, Chao Wang 0128, Yi Wang 0004, Yehao Feng, Shize Zhang, Xiaoqing Sun, Jianyuan Lu, Xing Li 0007, Biao Lyu, Zhigang Zong, Shunmin Zhu
APNet7
2025 Nezha: SmartNIC-based Virtual Switch Load Sharing
abstract
Cloud providers use SmartNIC-accelerated virtual switches (vSwitches) to offer rich network functions (NFs) for tenant VMs. Constrained by limited SmartNIC resources, it is a challenge to provide sufficient network performance for high-demand VMs. Meanwhile, we observed a significant number of idle vSwitches in the data center, which led us to consider leveraging them to build a remote resource pool for high-demand virtual NICs (vNICs). In this work, we propose Nezha, a distributed vSwitch load sharing system. Nezha reuses the existing idle SmartNICs to handle the excess load from the local SmartNIC without adding new devices. Nezha offloads stateless rule/flow tables to the remote, while keeping states locally. This eliminates the need for state synchronization, facilitating load sharing and failover. The deployment cost of Nezha is only a small fraction of that required to deploy new devices. Data collected from production show that our CPS capability bottleneck has shifted from the vSwitch to the VM kernel stack, with #concurrent flows and #vNICs increased by up to 50.4x and 40x, respectively.
Xing Li 0007, Enge Song, Tian Pan 0001, Qiang Fu 0011, Yang Song 0031, Yilong Lv, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Zhigang Zong, Qinming He, Shunmin Zhu
SIGCOMM8
2025 Tai Chi: A General High-Efficiency Scheduling Framework for SmartNICs in Hyperscale Clouds
abstract
Cloud service providers increasingly adopt SmartNICs to offload data-plane services (e.g., DPDK and SPDK) and control-plane tasks (such as disk and NIC initialization). Our analysis of production environments reveals that data-plane services statically provision CPUs for peak load, resulting in 67.5% idle CPU cycles during 99% of their runtime in IaaS clouds, leading to wasted CPU resources. On the other hand, control-plane tasks fail to meet critical Service Level Objectives (SLOs), such as virtual machine startup time. Unfortunately, achieving control-plane SLO improvements through co-scheduling with idle data-plane services remains highly challenging, due to the combined effects of intrinsic scheduling latency and the substantial architectural complexity inherent to control-plane ecosystems.
Bang Di, Kaijie Guo, Yibin Shen, Sanchuan Cheng, Fudong Qiu, Xiaokang Hu, Naixuan Guan, Dongdong Huang, Jinhu Li, Yi Wang 0004, Yifang Yang, Yilong Lv, Zhenwei Lu, Jiesheng Wu
SOSP18
2025 DrKD: Decoupling response-based distillation for object detection
Yilong Lv, Yancheng Cai
Pattern Recognit.1
2025 OAPR: An Offset-Aware Progressive Regression Object Detector
abstract
Object detection generally involves two main components: classification and regression. Despite the impressive performance achieved by recent refinement localization works, there is still room for improvement due to the limitations of current multistep regression strategies and task misalignment. To overcome these challenges, we propose a novel offset-aware progressive regression detector (OAPR) comprising an offset-aware head and a progressive regression predictor. Initially, we develop a head network incorporating our innovative plug-and-play offset-aware module. By utilizing the offset from one task to guide feature learning in another task, we intuitively achieve task alignment to address feature misalignment. We subsequently employ a progressive regression predictor to locate objects. In first-step regression, the aim is to identify a region within the object rather than the object itself. This is followed by second-step regression to locate the object precisely. Extensive experiments conducted on MS COCO datasets demonstrate the superior performance of our OAPR compared with recent state-of-the-art detectors with various backbones, including ATSS (~3.0 AP), GFL (~2.0 AP), BorderDet (~2.0 AP), and VFNet (~1.0 AP). Our code will be released.
Yilong Lv, Yujie He 0001, Min Li 0030
IEEE Trans. Circuits Syst. Video Technol.1
2024 Triton: A Flexible Hardware Offloading Architecture for Accelerating Apsara vSwitch in Alibaba Cloud
abstract
Apsara vSwitch (AVS) is a per-host deployed forwarding component for instance network connectivity in the Alibaba Cloud. To meet the growing performance demands, we accelerated AVS by adopting the most widely used "Sep-path" offloading architecture, which introduces a separate hardware data path to speed up popular traffic. However, the deployment results prove that it is difficult to bridge the gap in performance and programming flexibility of the software and hardware data paths, resulting in unpredictable performance and low iteration velocity.
Xing Li 0007, Xiaochong Jiang, Lilong Chen, Yi Wang 0004, Chao Wang 0128, Chao Xu 0017, Yilong Lv, Taotao Wu, Haifeng Gao, Yisong Qiao, Hongwei Ding 0004, Yijian Dong, Jianming Song, Jianyuan Lu, Chengkun Wei, Wenzhi Chen, Qinming He, Shunmin Zhu
SIGCOMM8
2024 Proactive Telemetry in Large-Scale Multi-Tenant Cloud Overlay Networks
abstract
At present, public clouds have served millions of tenants. To provide reliable services, cloud vendors need to perceive health status of the cloud network by building a telemetry system to detect possible network failures. While telemetry systems for physical networks have been extensively studied, research on telemetry systems for virtual networks is still insufficient. Different from physical networks, we conclude that building a virtual network telemetry system faces new challenges of feasibility, efficiency, and effectiveness. Specifically, we need to 1) protect privacy of tenants and adapt to heterogeneous middleboxes at the data plane; 2) handle frequent virtual network topology updates and compress large-scale measurement paths for millions of tenants at the control plane; 3) analyze telemetry results to locate network failures at the analysis plane. To address these challenges, we present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. At the data plane, Zoonet uses host agent and arp-ping to protect tenants’ privacy and defines an elegant generalization of ping and traceroute, which can work on heterogeneous middleboxes. At the control plane, Zoonet conducts update batch processing and substantial probing path pruning to lessen the overhead. At the analysis plane, Zoonet reduces noises and aggregates alerts based on temporal and spatial correlation and conducts the hop-by-hop telemetry mode to locate failures. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting.
Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Shize Zhang, Xiaoqing Sun, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Xihui Yang, Jiahai Yang 0001
IEEE/ACM Trans. Netw.10
2023 Anchor-Intermediate Detector: Decoupling and Coupling Bounding Boxes for Accurate Object Detection
abstract
Anchor-based detectors have been continuously developed for object detection. However, the individual anchor box makes it difficult to predict the boundary’s offset accurately. Instead of taking each bounding box as a closed individual, we consider using multiple boxes together to get prediction boxes. To this end, this paper proposes the Box Decouple-Couple(BDC) strategy in the inference, which no longer discards the overlapping boxes, but decouples the corner points of these boxes. Then, according to each corner’s score, we couple the corner points to select the most accurate corner pairs. To meet the BDC strategy, a simple but novel model is designed named the Anchor-Intermediate Detector(AID), which contains two head networks, i.e., an anchor-based head and an anchor-free Corner-aware head. The corner-aware head is able to score the corners of each bounding box to facilitate the coupling between corner points. Extensive experiments on MS COCO show that the proposed anchor-intermediate detector respectively outperforms their baseline RetinaNet and GFL method by ∼2.4 and ∼1.2 AP on the MS COCO test-dev dataset without any bells and whistles.
Yilong Lv, Min Li 0030, Yujie He 0001, Zhuzhen He, Shao-peng Li 0002, Aitao Yang
ICCV1
2023 Poster: Triton: Accelerating vSwitch with Flexibility through Hardware Assisting not Bypassing Software
abstract
The vSwitch, as a critical component for Virtual Machine (VM) network connectivity in cloud environments, has prompted increasing attention towards its forwarding performance. While software optimization schemes have limitations in meeting the expanding network capacity demands [11, 12, 15, 17, 18], hardware offloading architectures leveraging SoC, FPGA, and ASIC have been proposed to transfer the match-action workload [1, 3, 6, 7, 13, 16], addressing the growing need for network capacity.
Xing Li 0007, Xiaochong Jiang, Lilong Chen, Tianyu Xu 0007, Chao Xu 0017, Longbiao Xiao, Fengmin Shi, Yi Wang 0004, Taotao Wu, Yilong Lv, Hangfeng Gao, Yisong Qiao, Hongwei Ding 0004, Yijian Dong, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
SIGCOMM11
2023 Achelous: Enabling Programmability, Elasticity, and Reliability in Hyperscale Cloud Networks
abstract
Cloud computing has witnessed tremendous growth, prompting enterprises to migrate to the cloud for reliable and on-demand computing. Within a single Virtual Private Cloud (VPC), the number of instances (such as VMs, bare metals, and containers) has reached millions, posing challenges related to supporting millions of instances with network location decoupling from the underlying hardware, high elastic performance, and high reliability. However, academic studies have primarily focused on specific issues like high-speed data plane and virtualized routing infrastructure, while existing industrial network technologies fail to adequately address these challenges.
Chengkun Wei, Xing Li 0007, Xiaochong Jiang, Tianyu Xu 0007, Taotao Wu, Chao Xu 0017, Yilong Lv, Haifeng Gao, Zeke Wang, Shunmin Zhu, Wenzhi Chen
SIGCOMM9
2023 Rethinking cross-domain semantic relation for few-shot image generation
Yao Gou, Min Li 0030, Yilong Lv, Yusen Zhang 0008, Yuhang Xing, Yujie He 0001
Appl. Intell.3
2023 An Effective Instance-Level Contrastive Training Strategy for Ship Detection in SAR Images
abstract
Existing ship detection approaches in SAR images often suffer from inadequate learning of the detector and sub-optimal detection performance. To this end, Based on self-supervised contrastive learning, this letter consider using the relationship between samples to develop a more effective training strategy. First, the Instance-based RoI encode head is proposed, named InsRen head, a simple yet effective network structure. Its purpose is to encode the samples into a contrastive feature space, facilitating the measurement of contrastive learning. Furthermore, to adapt contrastive learning to ship detection, we have redefined some basic terms, such as query, positive key, and negative key, which can help the model build the training pipeline. Finally, we design the Instance-based Contrastive loss that does not require label supervision, named InsCon loss. With the penalty of the InsCon loss, the queries and positive key can learn more similar representations in the contrastive feature space. Simultaneously, the query and negative key are as far away as possible to increase the difference. With the help of InsRen head and InsCon loss, the training of the detection model is more effective. Experimental results demonstrate the superiority of our method.
Yilong Lv, Min Li 0030, Yujie He 0001
IEEE Geosci. Remote. Sens. Lett.1
2023 GTFN: GCN and Transformer Fusion Network With Spatial-Spectral Features for Hyperspectral Image Classification
abstract
Transformer has been widely used in classification tasks for hyperspectral images (HSI) in recent years. Because it can mine spectral sequence information to establish long-range dependence, its classification performance can be comparable with the convolutional neural network (CNN). However, both CNN and Transformer focus excessively on spatial or spectral domain features, resulting in an insufficient combination of spatial-spectral domain information from HSI for modeling. To solve this problem, we propose a new end-to-end graph convolutional network (GCN) and Transformer fusion network with the spatial-spectral feature extraction (GTFN) in this paper, which combines the strengths of GCN and Transformer in both spatial and spectral domain feature extraction, taking full advantage of the contextual information of classified pixels while establishing remote dependencies in the spectral domain compared with previous approaches. In addition, GTFN uses Follow Patch as an input to the GCN and effectively solves the problem of high model complexity while mining the relationship between pixels. It is worth noting that the spectral attention module is introduced in the process of GCN feature extraction, focusing on the contribution of different spectral bands to the classification. More importantly, to overcome the problem that Transformer is too scattered in the frequency domain feature extraction, a neighborhood convolution module is designed to fuse the local spectral domain features. On Indian Pines, Salinas, and Pavia University datasets, the overall accuracies (OAs) of our GTFN are 94.00%, 96.81%, and 95.14%, respectively. The core code of GTFN is released at https://github.com/1useryang/GTFN.
Aitao Yang, Min Li 0030, Yao Ding 0010, Danfeng Hong, Yilong Lv, Yujie He 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Zoonet: a proactive telemetry system for large-scale cloud networks
abstract
We present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. The requirements are to (1) cover hyper-scale virtual networks with millions of tenants and millions of VMs for top tenants; (2) handle frequent virtual topology changes due to tenants' configuration through flexible APIs; (3) adapt to heterogeneous middleboxes along the probing paths; (4) achieve VM-to-VM telemetry without breaking tenant privacy; (5) differentiate virtual and physical network problems. We argue existing physical network telemetry solutions fail to satisfy our needs due to either incomplete telemetry coverage or outrageous telemetry overhead. Zoonet sets an ambitious goal to provide VM-to-VM hop-by-hop telemetry for each tenant, which is achieved based on self-developed, customizable middleboxes via hundreds of person-months under close team collaboration. At the data plane, Zoonet defines an elegant generalization of ping and traceroute, but made to work on multi-tenant clouds with heterogeneous middleboxes. At the control plane, Zoonet conducts substantial probing path pruning and update batch processing to lessen the overhead. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting.
Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Jiahai Yang 0001
CoNEXT8
2022 TAFDet: A Task Awareness Focal Detector for Ship Detection in SAR Images
Yilong Lv, Min Li 0030, Yujie He 0001
PRCV (4)1
2021 C2QoS: CPU-Cycle based Network QoS Strategy in vSwitch of Public Cloud
Haiyang Jiang 0001, Yulei Wu, Yilong Lv, Xing Li 0007, Gaogang Xie
IM4
2021 S2H: Hypervisor as a setter within Virtualized Network I/O for VM isolation on cloud platform
Haiyang Jiang 0001, Guangxing Zhang, Xin Wang 0001, Yilong Lv, Xing Li 0007, Serge Fdida, Gaogang Xie
Comput. Networks5
2021 C2QoS: Network QoS guarantee in vSwitch through CPU-cycle management
Haiyang Jiang 0001, Yulei Wu, Chunjing Han, Yilong Lv, Xing Li 0007, Serge Fdida, Gaogang Xie
J. Syst. Archit.5