Mimi Qian

dblp:328/1122 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0001-8454-6702ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 dVRM: Cross-Switch Memory Sharing and Self-Adaptive Allocation in Distributed Data Plane
abstract
Programmable switches have revolutionized networking by enabling a new spectrum of applications, such as network telemetry, in-network computation, and machine learning. These applications heavily utilize register memory but their performance is significantly constrained by the scarcity of on-chip resources, such as the 15 MB of SRAM available on a Tofino switch. To effectively accommodate increasingly memorydemanding applications, we aim to pool register resources across multiple switches, creating a larger unified register memory space. This resource pooling approach addresses the limitations of existing single-switch Virtual Register Memory (VRM) solutions, which cannot meet the demands of these applications in distributed environments. To achieve this, we propose dVRM, a distributed VRM deployment framework that enables crossswitch memory sharing and self-dynamic memory allocation on the data plane. dVRM introduces three innovations: (1) Grouped Multi-Switch Registers (GMRs), virtualizing distributed pipeline stages into a unified memory pool; (2) a self-adaptive, bit-width allocation mechanism driven by real-time data-plane feedback; and (3) lightweight heuristics for concurrent application deployment with distributed VRM, formulated as a mixed-integer linear programming (MILP) problem. We have implemented dVRM on both P4 hardware switches (with Intel Tofino ASIC) and BMv2. Experimental results show that dVRM significantly reduces hash unit consumption by up to 26% and achieves an improvement in accuracy (ARE) of up to 57.3% across various workloads.
Mimi Qian, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Zhen Zhang 0017, Weijia Jia 0001
IEEE Trans. Computers1
2026 sketchPro: Identifying Top-k Items Based on Probabilistic Update on Programmable Data Plane
abstract
Detecting the top-k heaviest items in network traffic is fundamental to traffic engineering, congestion control, and security analytics. Controller-side solutions suffer from high communication latency and heavy resource overhead, motivating the migration of this task to programmable data planes (PDP). However, PDP hardware (e.g., Tofino ASIC) offers only a few megabytes of on-chip SRAM per pipeline stage and supports neither loops nor complex arithmetic, making accurate top-k detection highly challenging. This paper proposes sketchPro, a novel sketch-based solution that employs a probabilistic update scheme to retain large items, enabling accurate top-k identification on PDP with minimal memory. sketchPro dynamically adjusts the probability of updates based on the current statistical size of the items and the frequency of hash collisions, thus allowing sketchPro to effectively detect top-k items. We have implemented sketchPro on PDP, including P4 software switch (i.e., BMv2) and hardware switch (Intel Tofino ASIC). Extensive evaluation results demonstrate that sketchPro can achieve more than 95% precision with only 10KB of memory.
Keke Zheng, Mai Zhang, Mimi Qian, Waiming Lau, Lin Cui 0001
IEEE Trans. Netw. Serv. Manag.3
2025 DisPLOY: Target-Constrained Distributed Deployment for Network Measurement Tasks on Data Plane
abstract
In programmable networks, measurement tasks are placed on programmable switches to monitor network traffic at line rate. These tasks typically require substantial resources (e.g., significant SRAM), while programmable switches are constrained by limited resources due to their hardware design (e.g., Tofino ASIC), making distributed deployment essentially. Measurement tasks must monitor specific network locations or traffic flows, introducing significant complexity in deployment optimization. This target-constrained nature makes task optimization on switches (e.g., task merging) become device-dependent and order-dependent, which can lead to deployment failures or performance degradation if ignored. In this paper, we introduceDisPLOY, a novel target-constrained distributed deployment framework specifically designed for network measurement tasks on the data plane.DisPLOYenables operators to specify monitoring targets—network traffic or device/link—across multiple switches. Given the monitoring targets,DisPLOYeffectively minimizes redundant operations and optimizes deployment to achieve both resource efficiency (e.g., minimizing stage consumption) and high-performance monitoring (e.g., high accuracy). We implement and evaluateDisPLOYthrough deployment on both P4 hardware switches (Intel Tofino ASIC) and BMv2. Experimental results show thatDisPLOYsignificantly reduces stage consumption by up to 66% and improves ARE by up to 78.4% in flow size estimation while maintaining end-to-end performance.
Mimi Qian, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001, Zhetao Li, Weijia Jia 0001
IEEE Trans. Parallel Distributed Syst.1
2025 FlxVRM: Enabling Online Configuring Memory via Virtualization on Programmable Data Plane
abstract
Programmable data plane (PDP) has emerged as a powerful platform for line-rate packet processing, utilizing on-chip register memory to execute stateful applications. Yet most existing efforts concentrate on static approaches for allocating register memory, necessitating switch restarting and service interruption. Despite the availability of research on sharing memory for concurrent applications, the rigid requirement of limiting memory sharing to the same pipeline stages hampers application flexibility and poses scalability challenges. To address this limitation, we presentFlxVRM,a flexible register memory virtualization layerfor data plane P4 programs which supports high-flexibility sharing of register memory for concurrent applications on PDP.FlxVRMenables memory allocation at any stage and location of the pipeline on PDP for each application at run time. To reduce resource usage during virtualization in the data plane pipeline,FlxVRMfurther merges different tables and actions with similar structures within P4 programs. Additionally,FlxVRMprovides a compiler to generate data plane programs for virtualization as well as the control plane API configuration. A prototype ofFlxVRMis implemented based on P4 hardware switches with Intel Tofino ASIC. Our experiment results show thatFlxVRMsignificantly improves the allocatable memory space for applications by up to 50%, while reducing the resource of the table up to 68%.
Mimi Qian, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Zhen Zhang 0017, Weijia Jia 0001
IEEE Trans. Serv. Comput.1
2024 OffsetINT: Achieving High Accuracy and Low Bandwidth for In-Band Network Telemetry
abstract
Network measurement is essential for efficient network management and operations. In-band network telemetry (INT) offers fine-grained per-device per-packet information which could provide full-visibility for networks. However, the existing solutions fall short in achieving high accuracy, generality, and low overhead simultaneously. To address this limitation, we introduceOffsetINTto meet these three criteria. The key idea ofOffsetINTis to use minimal bits to carry collected states during monitoring, which is based on our observation that the value of telemetry states are usually very close (e.g., the time of adjacent arrival packets) or small (e.g., only a few tens of microseconds for processing latency) for most of the time in real networks. Instead of embedding complete values of state in packets,OffsetINToptimizes bit usage by encoding an offset (using fewer bits), which is carried in-band by passing packets to the end-hosts for recovery and analysis. We theoretically derive the bounds of bandwidth mitigation forOffsetINT. We have implementedOffsetINTin both P4 hardware switches (with Intel Tofino ASIC) and BMv2. Expensive evaluation results show thatOffsetINTcan achieve an accuracy of up to 100% compared to the original INT while reducing INT bandwidth by up to 48%.
Mimi Qian, Lin Cui 0001, Fung Po Tso 0001, Yuhui Deng 0001, Weijia Jia 0001
IEEE Trans. Serv. Comput.1
2023 A survey on sliding window sketch for network measurement
Zijie Zeng, Lin Cui 0001, Mimi Qian, Zhen Zhang 0017, Kaimin Wei
Comput. Networks3
2022 dDrops: Detecting silent packet drops on programmable data plane
Mimi Qian, Lin Cui 0001, Xiaoquan Zhang, Fung Po Tso 0001, Yuhui Deng 0001
Comput. Networks1