Gaofeng Lv

dblp:126/5231 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
14since 2021 · last 2026
0009-0004-3653-8432ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 8 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AeriSC: A Dual-Branch Semantic Communication System for UAV Aerial Image Transmission
Shaohua Lu, Zhongpei Liu, Gongyu Yang, Changshuai Zhan, Fengze Dong, Gaofeng Lv
ICIC (7)6
2026 ReMu: Bridging Fidelity and Flexibility in High-Mobility Network Emulation at Microsecond Scale
Mingtai Lv, Xuyan Jiang, Huan Zhou 0006, Gaofeng Lv, Jinshu Su, Xiangrui Yang 0002
IWQoS5
2026 AdaptTree: A Practical and Adaptive Packet Classification Scheme on FPGA
Jincheng Zhong, Gaofeng Lv, Shuhui Chen
SECON2
2026 Doubling the speed of large-scale packet classification through compressing decision tree nodes
Jincheng Zhong, Gaofeng Lv, Shuhui Chen
Comput. Networks3
2025 Memory-Efficient Packet Classification at High-Speed: The pRFC Architecture with Heuristic Partitioning
abstract
Packet classification is essential for modern networked systems, the rapid growth of rule sets and strategies in SDN and NFV environments demands higher performance and better memory-efficient solutions than ever. Existing RFC-based approaches, such as HybridRFC, suffer trade-offs between speed and memory usage. This paper presents pRFC, a partitioningenhanced recursive flow classification architecture that improves classification performance while significantly reducing memory consumption. By introducing a prefix-length-guided partitioning strategy and a lightweight compression mechanism, pRFC mitigates cross-product explosion and reduces bitwise processing overhead. Compared to uniform partitioning, it achieves up to 16.86% lower memory usage and 34.19% faster construction. Evaluations on ClassBench show that pRFC reduces memory usage by up to 80%, accelerates construction by up to 97%, and improves throughput by$4.0 \times$over standard RFC. Against HybridRFC, it achieves 72% lower memory consumption, 10% faster construction, and$2.45 \times$higher software throughput. An FPGA prototype demonstrates that pRFC fits entirely within on-chip memory and supports 100 Gbps line-rate classification via pipelining. These results highlight the effectiveness and practicality of pRFC for large-scale rule classification in resourceconstrained programmable networks.
Yuanfeng Chen, Xiangrui Yang 0002, Xuyan Jiang, Jincheng Zhong, Gaofeng Lv
IWQoS5
2025 GENDN: A Geospatially Enhanced NDN Framework for Location-Related Pub/Sub Services in NTN-Enabled IoT
abstract
Leveraging satellites and aerials vehicle, nonterrestrial network (NTN)-enabled IoT networks enhance coverage and reliability, enabling global data connections in remote and underserved regions. A key application within these networks is the location-related publish/subscribe service (LPSS), which is geospatial location sensitive, real time, and energy efficient, supporting disaster early warning and environmental monitoring. We demonstrate that, compared to IP technology, named data networking (NDN) is more suited to supporting LPSS. However, current NTN-enabled IoT networks lack mechanisms to utilize geospatial characteristics effectively. Additionally, interactions between IoT devices and aerial vehicles or satellites face challenges, such as low bandwidth, high latency, and intermittent connectivity, which hinder the efficiency of LPSS. We propose geospatially enhanced NDN (GENDN), an adapted NDN framework for supporting LPSS. GENDN incorporates Geohash encoding in content names, allowing flexible use of geospatial characteristics in data subscription. GENDN enhances request aggregation, enabling a single Interest packet (I-pkt) to subscribe to all data in adjacent areas without sequential matching and retrieval. Simulation experiments demonstrate that, compared to traditional NDN, GENDN: 1) effectively leverages geospatial data characteristics, increasing the hit rate of I-pkts in LPSS; 2) reduces the PIT size and network communication overhead, enhancing real-time performance and energy efficiency; and 3) shows potential for large-scale deployment in NTN-enabled IoT environments.
Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002, Gaofeng Lv
IEEE Internet Things J.6
2024 GeoNDN: Naming the localized data with the Geohash-based Scheme for NDN optimization
abstract
Named Data Networking (NDN) is an emerging paradigm for future networks that facilitates content retrieval by managing data directly through their names. However, in Non-Terrestrial Networks (NTN), including satellites and UAVs, data often have strong geospatial characteristics due to dynamic topologies and geographic dependencies. The current NDN protocol struggles to leverage these geospatial characteristics due to the mismatch between two-dimensional geographical coordinates and one-dimensional data names, resulting in inefficient data retrieval. We propose GeoNDN, a Geohash-based naming scheme that effectively utilizes geospatial characteristics during data retrieval. GeoNDN enhances the aggregation of Interest packets when consumers retrieve data from nearby areas, allowing all data from a target area to be obtained with a single Interest packet instead of sequentially matching each piece of data. Simulation experiments tailored to NTN wireless scenarios demonstrate that GeoNDN reduces the Pending Interest Table (PIT) scale by approximately 25%, decreases network communication overhead by 21%, and increases the Interest packet hit rate by about 12% compared to the traditional NDN naming. The experimental results also show that GeoNDN’s optimization effects are proportional to request density, highlighting its potential for large-scale deployment.
Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002, Gaofeng Lv
ISPA6
2022 QUIC Cryption Offloading Based on NanoBPF
abstract
QUIC is a new transmission protocol parallel with TCP. Compared to TCP, QUIC has advantages, but still, apparent bottlenecks need to be optimized. The optimization method follows the TCP research route. The mainstream is the hardware offloading technology, which offloads the computing-intensive functional modules to the network equipment, and the hardware processing replaces the host CPU for computing. However, the performance of hardware offloading is high, but the versatility and programmability are not guaranteed. To overcome the limitation above, we proposed an offloading model named NanoBPF, based on the RISC multicore DPU. The model modified the boot code of the Bootloader, guided and activated the BPF code as a runtime environment, and offloaded the QUIC's cryption module, which is high CPU occupancy. The model prototype is verified by dual host interconnection and Docker-based simulation topology. Experimental results showed that the offloading of en/decryption improved the throughput by nearly 13% and guaranteed fairness with TCP under certain conditions.
Jichang Wang, Gaofeng Lv, Zhongpei Liu, Xiangrui Yang 0002
APNOMS2
2022 Memory-efficient RMT Matching Optimization Based on MBitTree
abstract
Reconfigurable match tables (RMT) is a pro-grammable pipeline architecture for packet processing. The ar-chitecture searches for action instructions by matching keywords in the packet header vector to modify the packet header. Among them, exact matching uses hash matching, while mask matching is currently more widely implemented using the Ternary Content Addressable Memory (TCAM). TCAM has high classification performance, but its high cost and power consumption make it difficult to scale to large-scale rule sets. MBitTree, a decision tree based on multi-bit cutting implemented on FPGA, is considered to be one of the most scalable packet classification algorithms due to its fast classification speed and low memory footprint. Therefore, MBitTree is applied in the matching action stage of RMT to improve the mask matching and reduce the memory overhead of RMT. According to the characteristics of RMT pipeline, MBitTree is mapped and optimized to improve pipeline efficiency and make full use of hardware resources. In addition, for the first time, we propose to move the key extractor in each stage of RMT to the action engine of the previous stage to save the memory overhead and processing time caused by the key extractor in each stage. We implement a prototype RMT based on MBitTree matching on FPGA, and the implementation results show that our method can achieve a throughput of over 200 Gbps for 10K rule sets and greatly reduce the memory overhead.
Zhongpei Liu, Gaofeng Lv, Jichang Wang, Xiangrui Yang 0002
FPT2
2021 Joint Channel and Power Allocation Algorithm for Flying Ad Hoc Networks Based on Bayesian Optimization
Wei Peng 0005, Gaofeng Lv
AINA (1)4
2021 High-performance pipeline architecture for packet classification accelerator in DPU
abstract
Packet classification is a fundamental problem in the network. With the rapid growth of network bandwidth, wire-speed packet classification has become a key challenge for next-generation network processors. In this paper, we propose a decision-tree-based, multi-pipeline architecture for packet classification accelerator in Data Processing Unit (DPU). Our solution is based on MBitTree, a memory-efficient decision tree algorithm for packet classification. First, we present a parallel architecture composed of multiple linear pipelines for efficiently mapping the decision tree built by MBitTree. Second, a special logic is designed to quickly traverse the decision tree, reducing the logic delay of the pipeline stage. Finally, several pipeline optimization techniques are proposed to improve the performance of the architecture. The implementation results show that our architecture can achieve more than 250 Gbps throughput for the 64-byte minimum Ethernet packets, and can store 100K rules in the on-chip memory of a single NetFPGA_SUME.
Gaofeng Lv, Yanni Ma, Guanjie Qiao
FPT2
2021 MBitTree: A fast and scalable packet classification for software switches
abstract
Packet classification is a key building block of many network services, such as Quality of Service and network security. These network services require packet classification to be as fast as possible while using less memory and supporting scalability. Moreover, software-defined networking switches pose new challenges to packet classification in terms of the high dimensionality and large scale of rulesets. In this paper, we propose a new solution called MBitTree, which includes two major improvements over existing decision tree algorithms. First, we introduce a new ruleset partitioning technique to achieve adaptive and fast ruleset partitions. Second, a new multi-bit cutting scheme is used to build short trees while rarely causing rule replication. MBitTree can provide high classification speed and has good scalability. Experimental results show that compared to CutSplit, MBitTree achieves up to 6.8 times less memory consumption, as well as up to 1.7 times reduction on the number of memory accesses. Additionally, we implement the prototype of MBitTree on an FPGA, and the implementation results show that our approach can achieve more than 100 Gbps throughput for 10K rulesets and can handle over 100K rulesets on NetFPGA.
Gaofeng Lv, Guanjie Qiao
HOTI2
2021 The Design and Implementation of an Efficient Quaternary Network Flow Watermark Technology
abstract
With the increasing demand for network security, because passive traffic analysis technology has the characteristics of low efficiency, high overhead, and susceptibility to interferences, network flow watermark technology has emerged. As an active traffic analysis method, network flow watermark technology can effectively track malicious anonymous communication users and real attackers behind the stepping-stone chain, which has the advantages of high accuracy and low overhead. However, the existing flow watermark codec is embedded in the software, which is only suitable for low-rate or small-sample network traffic. In the face of modern network high-rate data stream (such as 100gbps per port), it has exceeded the software and traditional switch processing power. In order to improve the coding efficiency and processing capacity of network flow watermark, this paper combines flow watermark technology with smart NIC (Network Interface Card), and proposes an efficient quaternary network flow watermark technology, deployed in the switch to form an efficient dynamic watermark mechanism. Theoretical analysis and experimental results show that the network flow watermark technology can efficiently process high-rate data stream, with higher coding efficiency and good robustness to disturbances.
Lusha Mo, Gaofeng Lv, Guanjie Qiao
MSN2
2021 Hotcount: A High-Precision Traffic Statistics for Multi-Tenants
abstract
Traffic statistics in large-scale data streams play an important role in the network community, and can be used for congestion control, anomaly detection, heavy hitter detection, etc. However, in the context of multi-tenancy, accomplishing the traffic statistics task of multiple tenants with limited resources (CPU, memory, etc) has become one of the major challenges. To solve the challenges above, this paper proposes a multi-tenant-oriented traffic statistics structure, named Hotcount, which can grantee the dynamic allocation of resources in multi-tenancy scenarios. At the same time, Hotcount is also able to realize multi-tenant traffic statistics task, and further improve the accuracy of the traffic statistics task. Hotcount separates the cold and hot flows based on statistics. It uses the hot/cold part to record the hot flow size with high precision and cold flow size with lower precision. With extensive experiment, we proved that the processing speed of Hotcount is similar to that of the original classic algorithm. Meanwhile, it greatly improves the accuracy of traffic statistics tasks. In the per-flow size statistics task, the accuracy is improved by 6.7 to 49.5 times than the original algorithm. In the heavy hitter detection task, the accuracy is improved by 11.3 times to 2065.6 times than the original algorithm even with memory-size constraints.
Guanjie Qiao, Gaofeng Lv, Lusha Mo
Networking2
2019 A Heterogeneous Parallel Packet Processing Architecture for NFV Acceleration
abstract
Network function virtualization (NFV) offers a new way to design, deploy and manage networking services. It is of vital importance to exploit heterogeneous parallelism between hardware and software, in order to improve virtulization performance and quality of virtualized network services. In this poster, we propose a novel heterogeneous parallel architecture that highly exploits the parallelism inside packet processing, and implementation efficacy with hardware processing engines and software threads. We present two packet processing pipelines with three implemented VNF instances to better demonstrate the efficiency of heterogeneous parallelism in accelerating NFV. We show the performance of our proposed architecture with various virtualized requirements and traffics in a well-deployed network environment. Experimental results reveal that it can achieve accelerated NFV performance, as well as provide a wide class of VNFs to improve the quality of virtualized network services.
Jinshu Su, Biao Han 0003, Gaofeng Lv, Tao Li 0008, Zhigang Sun 0002
ICNP3
2018 Demonstration of Path-Based Packet Batcher for Accelerating Vectorized Packet Processing
abstract
Recently, a major challenge on generic multi-core network processing platforms is how to improve packet processing performance. Vector packet processor (VPP) is a modularized and high- performance software framework for building network dataplane applications. The key idea of VPP is to reduce instruction cache (i-cache) misses with vectorized packet processing. However, the packets in a vector may traverse different processing paths in some scenarios. In such case, the vector is split into several smaller vectors, and the per- packet overhead would increase. In this paper, we propose a Path-based Packet Batcher (PPB) to accelerate VPP. PPB is transparent to VPP, and it requires no modification to VPP. Before VPP processes packets, PPB batches the packets based on the processing paths they will traverse. We build a prototype based on FPGA to evaluate the performance optimizations to VPP with PPB. Experiment results show that the reduction of i-cache misses can be up to 57.6% when the batch size is 128.
Jinli Yan, Tao Li 0008, Gaofeng Lv, Zhigang Sun 0002
SECON4
2015 FRINGE: Improving the scalability of Ethernet DCN via efficient software-defined edge control
abstract
This paper introduces a topology-independent software-defined edge control framework named FRINGE to scale out the Ethernet Datacenter Network (DCN). FRINGE exploits programmable OpenFlow-enabled switches deployed at the edge of DCN to aggregate the forwarding rules without introducing extra packet headers. We implement the proposed FRINGE framework in an SDN prototyping environment and validate it under three typical DCN topologies including Multi-Root Tree, HyperX and Jellyfish, where three different types of DCN workloads are applied. Evaluation results reveal that FRINGE can significantly reduce the total number of rules in all network devices and suppress most of the useless broadcast packets in the DCN.
Jianbiao Mao, Biao Han 0003, Gaofeng Lv, Zhigang Sun 0002, Xicheng Lu
IWQoS3
2014 Demostration of Self-Described Buffer for Accelerating Packet Forwarding on Multi-core Servers
abstract
Network processing platform based on the multi-core CPU becomes more and more prevailing in nowadays. Buffer allocation/deallocation operations consume a large number of CPU cycles in packet I/O process. The problem becomes even worse in the scenario of packet forwarding, as buffer allocation/deallocation operations are more frequent than the host-based network applications. We thus propose a novel data structure for packet buffer management on multi-cores, named Self-Described Buffer (SDB), which merges the separated descriptor and metadata into packet buffer. SDB management overhead can be greatly reduced by utilizing the compact data structure, and zero-overhead buffer management can be further achieved by offloading SDB allocation/deallocation operations to NIC. We have prototyped SDB enabled NIC, named BcNIC, on NetFPGA-10G. In the demo, we will illustrate the advantages of the SDB scheme by comparing the performance of BcNIC with the traditional NIC on multi-core platforms.
Zhigang Sun 0002, Tao Li 0008, Biao Han 0003, Gaofeng Lv
CloudCom5