EDBT 2026 Demo / reviewers in the wild / expert
Xiangrui Yang 0002
dblp:213/1007-2
· DBLP profile ↗
21ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0002-8041-9957ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 2 first-author · 11 since 2021Systems, architecture and hardware · 6 · 5 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReMu: Bridging Fidelity and Flexibility in High-Mobility Network Emulation at Microsecond Scale
Mingtai Lv, Xuyan Jiang, Huan Zhou 0006, Gaofeng Lv, Jinshu Su, Xiangrui Yang 0002 |
IWQoS | 7 |
| 2026 | EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu 0003, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Zhan Wang 0003, Yuchao Zhang 0004, Yang Liu 0038, Xiangrui Yang 0002, Xiaohe Hu, Limin Xiao 0001, Weifeng Zhang 0003, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu |
SIGCOMM | 19 |
| 2025 | Megabits Down to Kilobits: Memory-Efficient Time-Aware Shaping for TSNabstractTime-Sensitive Networking (TSN) provides bounded latency and low jitter for cyber-physical systems, such as industrial control. As a key component of TSN, the Time-Aware Shaper (TAS) applies gate control rules to control the transmission time of frames in critical flows. TAS stores the gate control rules for each frame in the gate control table. However, in typical industrial setups, the memory usage of the table could reach over tens of megabits and even exceed the total memory capacity of TSN switches.To address this issue, we propose a memory-efficient TAS design named METAS. It transitions from a per-frame to a per-flow approach. METAS stores one persistent rule for a flow and dynamically generates a temporary rule for a frame only when the frame arrives. We prototyped METAS on an FPGA, and experimental results show that METAS reduces memory usage from 14.34 Mbits to 288 Kbits when supporting 1,024 flows, using just 1.56% of the FPGA’s logic resources while maintaining microsecondlevel latency and nanosecond-level jitter for critical flows. Xuyan Jiang, Wenwen Fu, Xiangrui Yang 0002, Wenfei Wu |
DAC | 3 |
| 2025 | Memory-Efficient Packet Classification at High-Speed: The pRFC Architecture with Heuristic PartitioningabstractPacket classification is essential for modern networked systems, the rapid growth of rule sets and strategies in SDN and NFV environments demands higher performance and better memory-efficient solutions than ever. Existing RFC-based approaches, such as HybridRFC, suffer trade-offs between speed and memory usage. This paper presents pRFC, a partitioningenhanced recursive flow classification architecture that improves classification performance while significantly reducing memory consumption. By introducing a prefix-length-guided partitioning strategy and a lightweight compression mechanism, pRFC mitigates cross-product explosion and reduces bitwise processing overhead. Compared to uniform partitioning, it achieves up to 16.86% lower memory usage and 34.19% faster construction. Evaluations on ClassBench show that pRFC reduces memory usage by up to 80%, accelerates construction by up to 97%, and improves throughput by$4.0 \times$over standard RFC. Against HybridRFC, it achieves 72% lower memory consumption, 10% faster construction, and$2.45 \times$higher software throughput. An FPGA prototype demonstrates that pRFC fits entirely within on-chip memory and supports 100 Gbps line-rate classification via pipelining. These results highlight the effectiveness and practicality of pRFC for large-scale rule classification in resourceconstrained programmable networks. Yuanfeng Chen, Xiangrui Yang 0002, Xuyan Jiang, Jincheng Zhong, Gaofeng Lv |
IWQoS | 2 |
| 2025 | GENDN: A Geospatially Enhanced NDN Framework for Location-Related Pub/Sub Services in NTN-Enabled IoTabstractLeveraging satellites and aerials vehicle, nonterrestrial network (NTN)-enabled IoT networks enhance coverage and reliability, enabling global data connections in remote and underserved regions. A key application within these networks is the location-related publish/subscribe service (LPSS), which is geospatial location sensitive, real time, and energy efficient, supporting disaster early warning and environmental monitoring. We demonstrate that, compared to IP technology, named data networking (NDN) is more suited to supporting LPSS. However, current NTN-enabled IoT networks lack mechanisms to utilize geospatial characteristics effectively. Additionally, interactions between IoT devices and aerial vehicles or satellites face challenges, such as low bandwidth, high latency, and intermittent connectivity, which hinder the efficiency of LPSS. We propose geospatially enhanced NDN (GENDN), an adapted NDN framework for supporting LPSS. GENDN incorporates Geohash encoding in content names, allowing flexible use of geospatial characteristics in data subscription. GENDN enhances request aggregation, enabling a single Interest packet (I-pkt) to subscribe to all data in adjacent areas without sequential matching and retrieval. Simulation experiments demonstrate that, compared to traditional NDN, GENDN: 1) effectively leverages geospatial data characteristics, increasing the hit rate of I-pkts in LPSS; 2) reduces the PIT size and network communication overhead, enhancing real-time performance and energy efficiency; and 3) shows potential for large-scale deployment in NTN-enabled IoT environments. Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002, Gaofeng Lv |
IEEE Internet Things J. | 4 |
| 2025 | FooDog: Empower TSN for Efficient PolicingabstractTime-Sensitive Networking (TSN) is an emerging real-time Ethernet technology that provides deterministic communication for time-sensitive (TS) traffic. At its core, TSN utilizes Per-Stream Filtering and Policing (PSFP) gates to mitigate the disruption of unavoidable frame drift. However, as first identified in this work, the naive PSFP gate design results in heavy memory usage, which hinders normal switching functions. This work proposes an efficient PSFP gate design called FooDog. FooDog employs a two-stage structure and a dual-engine policing mechanism to realize memory-efficient, logic-compact, and fast policing while maintaining minimal latency and jitter for TS traffic. Results on FPGA prototypes show that FooDog consumes only hundreds of kilobits of memory, reducing on-chip memory overheads by more than 90% compared to the unoptimized PSFP gate design. Additionally, it maintains end-to-end latency in the microsecond range and jitter below 150 nanoseconds under abnormal traffic conditions, comparable to typical TSN performance without anomalies. Xuyan Jiang, Xiangrui Yang 0002, Tongqing Zhou, Wenfei Wu, Wenwen Fu, Wei Quan 0004, Yingwen Chen 0001, Yihao Jiao, Zhigang Sun 0002 |
IEEE Trans. Netw. | 2 |
| 2025 | Magneto: Load-Balanced Key-Value Service for Write-Intensive WorkloadsabstractHigh-performance key-value (KV) storage is critical for the cloud, providing KV services to various cloud applications. A key challenge for KV services is that workloads of cloud applications are often write-intensive and exhibit highly-skewed characteristics, which result in load imbalance among storage servers thus lowering system performance. To address this problem, prior works set up an in-switch write-back caching mechanism, which adopts a centralized controller to balance writes. However, the limited bandwidth between the switches and the controller is a severe bottleneck for achieving high performance. In this paper, we present Magneto, a novel key-value service architecture for the cloud. At the core of Magneto is a delayed-write mechanism in the switch data plane to absorb frequent write queries for hot items. It effectively balances the load under write-intensive workloads without involving the controller. Magneto also designs a reliability mechanism to ensure system reliability during switch state transitions and failures. We implement a prototype using an FPGA-integrated switch, which has high packet-processing performance and contains enough memory to provide both the in-switch cache and the write buffer. Extensive evaluation shows that Magneto can achieve 8.4x system throughput gains compared to baseline systems when handling a skewed workload consisting of 70% reads and 30% writes. Moreover, it can reduce the load on back-end servers up to 56% in total. Yuanhang Gao, Yingwen Chen 0001, Xiangrui Yang 0002, Huan Zhou 0006, Shihua Tang |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | A Geohash-based Naming Scheme for NDN Optimization
Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002 |
APNet | 4 |
| 2024 | WeMu: A design of wireless network emulator
Mingtai Lv, Xiangrui Yang 0002, Huan Zhou 0006, Wenfei Wu, Yusheng Xia, Jinshu Su |
APNet | 2 |
| 2024 | DeInfer: A GPU resource allocation algorithm with spatial sharing for near-deterministic inferring tasksabstractFor the applications of artificial intelligence, training models with GPUs are widely noted, while inferring requirements are somehow neglected. In some scenario, it is quite important to finish the deep learning inference (DLI) task and get the response in time, e.g., anomaly detection in AIOps or QoE (Quality of Experience) assurance for customers. However, the challenges of GPU inferring are as follows: 1) the interference among inferring tasks on the shared GPU are not well studied and even not considered, which may cause a surge in the inference latency due to the hardware contention; 2) the deadline miss rate caused by the arrival rate is not clearly considered, which often exhibits significant fluctuations in real-world cases. Therefore, the interference among tasks and arrival rates of tasks should be well designed to decrease the deadline miss rate, when sharing GPU resources. To tackle the issue, we propose the algorithm, DeInfer, through following manners: 1) we identify the key factors that lead to interference, conduct a systematic study and develop a highly accurate interference prediction algorithm based on the random forest algorithm, achieving a four times improvement compared to the state-of-the-art interference prediction algorithms; 2) we utilize the queue theory to model the randomness of the arrival process and put forward a GPU resource allocation algorithm, which reduces the deadline miss rate by an average of over 30%. Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002, Yanfei Yin |
ICPP | 4 |
| 2024 | GeoNDN: Naming the localized data with the Geohash-based Scheme for NDN optimizationabstractNamed Data Networking (NDN) is an emerging paradigm for future networks that facilitates content retrieval by managing data directly through their names. However, in Non-Terrestrial Networks (NTN), including satellites and UAVs, data often have strong geospatial characteristics due to dynamic topologies and geographic dependencies. The current NDN protocol struggles to leverage these geospatial characteristics due to the mismatch between two-dimensional geographical coordinates and one-dimensional data names, resulting in inefficient data retrieval. We propose GeoNDN, a Geohash-based naming scheme that effectively utilizes geospatial characteristics during data retrieval. GeoNDN enhances the aggregation of Interest packets when consumers retrieve data from nearby areas, allowing all data from a target area to be obtained with a single Interest packet instead of sequentially matching each piece of data. Simulation experiments tailored to NTN wireless scenarios demonstrate that GeoNDN reduces the Pending Interest Table (PIT) scale by approximately 25%, decreases network communication overhead by 21%, and increases the Interest packet hit rate by about 12% compared to the traditional NDN naming. The experimental results also show that GeoNDN’s optimization effects are proportional to request density, highlighting its potential for large-scale deployment. Yingwen Chen 0001, Huan Zhou 0006, Xiangrui Yang 0002, Gaofeng Lv |
ISPA | 4 |
| 2024 | Optimizing In-network Caching for Key-Value Stores under Write-intensive WorkloadsabstractKey-value stores are critical for modern data centers, yet they struggle with performance degradation under write-intensive workloads due to frequent cache invalidations. Traditional read-optimized caching mechanisms fail to maintain efficiency as frequent updates lead to obsolete cached data. As a result, queries sent by clients will suffer from long queuing delays or even be dropped and overall system performance will be degraded severely. This paper presents a novel key-value stores architecture, named Freq-absorb, which addresses this issue through a delayed-write mechanism delaying writes commitment to backend servers. By buffering write queries within the switch for a specific period, Freq-absorb reduces the impact of cache invalidations, improves switch hit rates, and enhances overall system performance. We implement a prototype using an FPGA-based switch that acts as the ToR (Top of Rack) switch to achieve cache and the appended write buffer. Our experimental results show that compared to the read cache mechanism, our approach can reduce switch miss from 99% to 48.4% when handling write-intensive workloads, significantly enhancing performance. Yuanhang Gao, Yingwen Chen 0001, Xiangrui Yang 0002, Huan Zhou 0006, Shihua Tang, Yuanfeng Chen |
ISPA | 3 |
| 2024 | Hebo: FPGA-based Transfer Time Planning for Volatile Traffic in TSNabstractTime-Sensitive Networking (TSN) is an advanced technology designed for real-time Ethernet communications, providing extremely low latency, minimal jitter, and lossless data transfer for time-sensitive critical traffic. Despite its benefits, TSN faces challenges with volatile traffic, where the time between frames constantly changes, leading to potential network performance issues. To tackle this issue, this paper introduces Hebo, a novel solution designed for zero frame loss and minimal latency of volatile traffic. Hebo employs a centralized controller, built with field-programmable gate array (FPGA), to dynamically plan the timing of volatile traffic in real-time. This approach ensures real-time data transmission across the network by efficiently allocating network resources. Our real-world and simulation experiments show that Hebo could significantly improve network performance, achieving less than 100 microseconds in end-to-end delay and eliminating nearly all frame loss (reducing it from approximately 80% to zero) in industrial automation scenarios. Xuyan Jiang, Zitong Wang 0002, Xiangrui Yang 0002, Yihao Jiao, Tianci Yu, Wenwen Fu, Yinhan Sun, Zhigang Sun 0002 |
IWQoS | 3 |
| 2022 | QUIC Cryption Offloading Based on NanoBPFabstractQUIC is a new transmission protocol parallel with TCP. Compared to TCP, QUIC has advantages, but still, apparent bottlenecks need to be optimized. The optimization method follows the TCP research route. The mainstream is the hardware offloading technology, which offloads the computing-intensive functional modules to the network equipment, and the hardware processing replaces the host CPU for computing. However, the performance of hardware offloading is high, but the versatility and programmability are not guaranteed. To overcome the limitation above, we proposed an offloading model named NanoBPF, based on the RISC multicore DPU. The model modified the boot code of the Bootloader, guided and activated the BPF code as a runtime environment, and offloaded the QUIC's cryption module, which is high CPU occupancy. The model prototype is verified by dual host interconnection and Docker-based simulation topology. Experimental results showed that the offloading of en/decryption improved the throughput by nearly 13% and guaranteed fairness with TCP under certain conditions. Jichang Wang, Gaofeng Lv, Zhongpei Liu, Xiangrui Yang 0002 |
APNOMS | 4 |
| 2022 | Memory-efficient RMT Matching Optimization Based on MBitTreeabstractReconfigurable match tables (RMT) is a pro-grammable pipeline architecture for packet processing. The ar-chitecture searches for action instructions by matching keywords in the packet header vector to modify the packet header. Among them, exact matching uses hash matching, while mask matching is currently more widely implemented using the Ternary Content Addressable Memory (TCAM). TCAM has high classification performance, but its high cost and power consumption make it difficult to scale to large-scale rule sets. MBitTree, a decision tree based on multi-bit cutting implemented on FPGA, is considered to be one of the most scalable packet classification algorithms due to its fast classification speed and low memory footprint. Therefore, MBitTree is applied in the matching action stage of RMT to improve the mask matching and reduce the memory overhead of RMT. According to the characteristics of RMT pipeline, MBitTree is mapped and optimized to improve pipeline efficiency and make full use of hardware resources. In addition, for the first time, we propose to move the key extractor in each stage of RMT to the action engine of the previous stage to save the memory overhead and processing time caused by the key extractor in each stage. We implement a prototype RMT based on MBitTree matching on FPGA, and the implementation results show that our method can achieve a throughput of over 200 Gbps for 10K rule sets and greatly reduce the memory overhead. Zhongpei Liu, Gaofeng Lv, Jichang Wang, Xiangrui Yang 0002 |
FPT | 4 |
| 2022 | TASP: Enabling Time-Triggered Task Scheduling in TSN-Based Mixed-Criticality SystemsabstractDistributed mixed-criticality system (DMCS) has been widely used in various critical domains such as self-driving cars and space crafts. To guarantee the end-to-end QoS (i.e., deadline/jitter requirements) of sensing-controlling-actuating control loops (CL) applications, DMCS adopts Time-Sensitive Networking (TSN), an emerging real-time Ethernet technology, for communication between end systems (ES). TSN provides a synchronized global clock and guarantees bounded delay for time-critical traffic in CLs, making it possible for DMCS to collaboratively schedule the computation (on ES) and communication (in TSN) to meet the Quality of Service (QoS) requirement. However, as modern DMCS tends to use fully-fledged Linux distributions (rather than a custom real-time OS) on ES to enjoy Linux’s mature ecosystem, it is challenging for DMCS to realize TSN-based QoS guarantees because the event-triggered scheduling of Linux on ES is incompatible with TSN.This paper proposes TAsk Scheduling Puppeteer (TASP), a mechanism that schedules CL tasks based on TSN without modifying the Linux OS. The key idea of TASP is to manipulate task scheduling by controlling the timing of CL packet submissions at the interface between TSN and ES. Specifically, TASP extracts two parameters: Fore Guardband (ForeGB) and Back Guardband (BackGB). During the ForeGB period before a CL packet’s submission, TASP forbids any packets’ submission; while during the BackGB period after a CL packet’s submission, TASP forbids any other CL packets’ submission. ForeGB and BackGB can ensure that there is at most one schedulable task on the ES at any time, and thus Linux has no choice but to schedule the only task, making the Linux scheduler a puppet. We have deployed TASP and evaluated it in real-world TSN switches based on an open-source TSN project, OpenTSN. The TASP-enabled ES can achieve task scheduling precisely based on TSN’s global clock, which outperforms the original ES by reducing end-to-end jitter from milliseconds to microseconds. Xuyan Jiang, Yiming Zhang 0003, Wenwen Fu, Xiangrui Yang 0002, Yinhan Sun, Zhigang Sun 0002 |
IWQoS | 4 |
| 2022 | Isolation Mechanisms for High-Speed Packet-Processing Pipelines
Tao Wang 0088, Xiangrui Yang 0002, Gianni Antichi, Anirudh Sivaraman, Aurojit Panda |
NSDI | 2 |
| 2020 | TSN-Builder: Enabling Rapid Customization of Resource-Efficient Switches for Time-Sensitive NetworkingabstractTime-Sensitive Networking (TSN) emerges as a promising technique empowering deterministic forwarding on standard Ethernet without sacrificing compatibility. There are some commercial off-the-shelf (COTS) switches that support TSN recently. However, the resource partitioning on these switches is normally inefficient for the on-chip memory resource in many specific application scenarios. We observe that the critical requirements (e.g., topology, flow features) of these scenarios are pre-determined. Thus, developing a TSN switch in a Top-down approach is feasible and urgently needed.In this paper, we propose TSN-Builder, a template-based developing model for customizing resource-efficient TSN switches rapidly with targeted application-dependent requirements. TSN-Builder decomposes the integrated TSN switching function into multiple function templates. With a fine-grained resource abstraction, TSN-Builder provides platform-independent customization interfaces for developers to customize the resource parameters. We prototype TSN switches on FPGA to evaluate the resource consumption and performance under different application scenarios. Experimental results show that TSN-Builder reduces the on-chip memory by up to 80.53% under the same Quality-of-Service, compared to the resource configuration in the COTS switch. Jinli Yan, Wei Quan 0004, Xiangrui Yang 0002, Wenwen Fu, Zhigang Sun 0002 |
DAC | 3 |
| 2019 | FAST: enabling fast software/hardware prototype for network experimentationabstractThe evolution of new technologies in network community is getting ever faster. Yet it remains the case that prototyping those novel mechanisms on a real-world system (i.e. CPU-FPGA platforms) is both time and labor consuming, which has a serious impact on the research timeliness. In order to bring researchers out of trivial process in prototype development, this paper proposed FAST, a software hardware co-design framework for fast network prototyping. With the programming abstraction of FAST, researchers are able to prototype (using C, verilog or both) a wide spectrum of network boxes rapidly based on all kinds of CPU-FPGA platforms. FAST framework takes care of managing DMA, PCIe and Linux Kernel while providing a unified API for researchers so they can focus only on the packet processing functions. We demonstrate FAST framework's easy to use features with a number of prototypes and show we can get over 10x gains in performance or 1000x better accuracy in clock synchronization compared with their software versions. Xiangrui Yang 0002, Zhigang Sun 0002, Junnan Li 0002, Jinli Yan, Tao Li 0008, Wei Quan 0004, Donglai Xu, Gianni Antichi |
IWQoS | 1 |
| 2018 | OverWatch: A Cross-Plane DDoS Attack Defense Framework with Collaborative Intelligence in SDNabstractDistributed Denial of Service (DDoS) attacks are one of the biggest concerns for security professionals. Traditional middle-box based DDoS attack defense is lack of network-wide monitoring flexibility. With the development of software-defined networking (SDN), it becomes prevalent to exploit centralized controllers to defend against DDoS attacks. However, current solutions suffer with serious southbound communication overhead and detection delay. In this paper, we propose a cross-plane DDoS attack defense framework in SDN, called OverWatch, which exploits collaborative intelligence between data plane and control plane with high defense efficiency. Attack detection and reaction are two key procedures of the proposed framework. We develop a collaborative DDoS attack detection mechanism, which consists of a coarse-grained flow monitoring algorithm on the data plane and a fine-grained machine learning based attack classification algorithm on the control plane. We propose a novel defense strategy offloading mechanism to dynamically deploy defense applications across the controller and switches, by which rapid attack reaction and accurate botnet location can be achieved. We conduct extensive experiments on a real-world SDN network. Experimental results validate the efficiency of our proposed OverWatch framework with high detection accuracy and real-time DDoS attack reaction, as well as reduced communication overhead on SDN southbound interface. Biao Han 0003, Xiangrui Yang 0002, Zhigang Sun 0002, Jinshu Su |
Secur. Commun. Networks | 2 |
| 2017 | SDN-Based DDoS Attack Detection with Cross-Plane Collaboration and Lightweight Flow MonitoringabstractDistributed Denial of Service (DDoS) attacks are one of the biggest concerns for security professionals. Traditional DDoS attack detection mechanisms are based on middle-box devices or SDN controllers, which either lack network-wide monitoring information or suffer with serious southbound communication overhead and detection delay. In this paper, we propose a SDN-based DDoS attack detection framework with cross-plane collaboration called OverWatch, which performs a two-stage granularity filtering procedure between coarse-grained detection data plane and fine- grained detection control plane for abnormal flows. It leverages computational capabilities that currently underutilized on OpenFlow switches to shrink the detection range for fine-grained DDoS attack detections. In OverWatch, we propose a lightweight flow monitoring algorithm to capture the key features of DDoS attack traffics on the data plane by polling the values of counters in OpenFlow switches. Experiments are conducted in an evaluating network with a FPGA-based OpenFlow switch prototype and the Ryu controller, which reveal that our proposed OverWatch framework and flow monitoring algorithm can greatly improve the detection efficiency, as well as reduce the detection delay and southbound communication overhead. Xiangrui Yang 0002, Biao Han 0003, Zhigang Sun 0002 |
GLOBECOM | 1 |