VLDB 2026 Research / reviewers in the wild / expert
Wenfei Wu
dblp:55/8398
· DBLP profile ↗
80ranked-venue papers
7as first author
57since 2021 · last 2027
0000-0002-1357-3137ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 54 · 5 first-author · 34 since 2021Systems, architecture and hardware · 16 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | PCGDNet: Reliability-aware glass detection with physics-inspired transparent-surface consistency from monocular RGB images
Jinsheng Xiao, Wenfei Wu |
Expert Syst. Appl. | 5 |
| 2026 | Multipath Collective Communication Beyond Scale-up Networks in GPU Clouds
Yuchen Xu 0003, Jianglong Nie, Baojia Li 0002, Mingzhuo Chen, Guanyu Qu, Zhenchuan Liu, Shuangshuang Yin, Chunzhi He, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Congcong Miao, Wenfei Wu |
EuroSys | 18 |
| 2026 | LCMP: Distributed Long-Haul Cost-Aware Multi-Path Routing for Inter-Datacenter RDMA NetworksabstractRDMA-empowered cloud services are gradually deployed across datacenters (DCs) with multiple paths, which exhibit new properties of path asymmetry, delayed congestion signals, and simultaneous flow routing collisions, and further fail existing routing methods. Dong-Yang Yu 0001, Yuchao Zhang 0004, Jun Wang 0178, Wenfei Wu, Haipeng Yao, Wendong Wang 0003, Ke Xu 0002 |
EuroSys | 5 |
| 2026 | TurboTSS: A Packet Classifier with Fast Rule Lookup and Update for the Cloud
Shaoke Fang, Yuchen Xu 0003, Weize Gao, Jianglong Nie, Wenfei Wu |
INFOCOM | 6 |
| 2026 | Breaking Bucket Effect in In-Network Aggregation via Memory-Bandwidth Coordination
Junxu Xia, Geyao Cheng, Deke Guo, Lailong Luo, Wenfei Wu |
INFOCOM | 5 |
| 2026 | SwitchNN: In-Network CNN Inference for Edge-Assisted Smart Roadside Networks
Jianqiang Zhong, Jingpu Duan, Wenfei Wu, Deke Guo, Bingyang Liu, Xu Chen 0004 |
IWQoS | 6 |
| 2026 | Enabling Flexible and Efficient Collective Communication Scheduling in Distributed AI
Yuchen Xu 0003, Xiting Ju, Wenfei Wu |
LANMAN | 3 |
| 2026 | Turbo: Efficiently Serving Long-Context Large Language Models with In-Network AggregationabstractLLM supporting long contexts faces a critical memory bottleneck due to the linear growth of KV cache. Distributing the storage across multiple GPUs alleviates this burden but introduces significant communication overhead or traffic incast, especially during the decoding phase. We propose Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches. We address three key challenges to map complex attention mechanisms onto restricted switch hardware: (i) To bypass the switch's inability to buffer global states or perform complex operations, we devise online table-based aggregation, which decomposes global reduction into pairwise operations and approximates nonlinear functions via lookup tables. (ii) To circumvent the restriction on retroactive state access in RMT pipelines, we introduce a rolling forward scheme that propagates states to enable cross-stage updates. (iii) To mitigate aggregation stragglers caused by topology-induced load imbalance, we construct a load-aware aggregation tree that optimizes workload distribution. Evaluations on a Tofino2-based testbed show that Turbo reduces end-to-end inference latency by up to 37%. Large-scale simulations on NS-3 demonstrate that Turbo significantly outperforms state-of-the-art baselines in both inference latency and network traffic reduction with negligible accuracy loss. Ying Wan 0001, Yuchen Xu 0003, Chuwen Zhang, Yingsheng Huang, Wenquan Xu, Jialin Li 0001, Mingwei Xu 0001, Wenfei Wu, Congcong Miao |
SIGCOMM | 9 |
| 2026 | EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
Yitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu 0003, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Zhan Wang 0003, Yuchao Zhang 0004, Yang Liu 0038, Xiangrui Yang 0002, Xiaohe Hu, Limin Xiao 0001, Weifeng Zhang 0003, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu |
SIGCOMM | 31 |
| 2026 | Intelligent assembly of shield tunnel lining segments: A vision-guided integrated approach
Yeting Zhu, Wenfei Wu, Peixin Chen, Weili Fang |
Adv. Eng. Informatics | 4 |
| 2026 | Enhanced depth completion for transparent objects via failure-aware RGB-D pretraining
Wenfei Wu, Jinsheng Xiao, Danqi Meng |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | Virtual Slicing: Achieving Control Plane Availability and Traffic Engineering Efficiency in Data CentersabstractMany proposals have demonstrated the efficiency advantages of software-defined networking (SDN) in managing data center networks. Common practices employ centralized traffic engineering (TE) in the SDN control plane to optimize load balancing and throughput. Meanwhile, for high availability purposes, the control plane is partitioned to ensure the impact of a single faulty controller is contained. However, the interaction between these two aspects is often overlooked. In particular, we show that the current control plane partitioning approach leads to imbalanced link loads and degraded application performance. To address this issue, we proposevirtual slicing, a new control plane partitioning scheme. Virtual slicing achieves desirable traffic engineering performance while retaining the availability guarantees from the current approach. Virtual slicing is implemented and evaluated with real-world and synthetic traffic traces on production spine-free data center networks. Results show that virtual slicing reduces tail link utilizations by up to 28.4%, and improves flow completion times by up to 36%. Brian Chang, Keqiang He, Shawn Shuoshuo Chen, Mingyang Zhang 0005, Wenfei Wu, Fan Wu 0006, Chen Tian 0001, Aditya Akella |
IEEE Trans. Netw. | 6 |
| 2026 | ICH: In-Network Cache Hierarchy for Dynamic WorkloadsabstractIn key-value storage systems, a lot of traffic is concentrated on some of the hotspots. The servers that keep hotspots will become bottlenecks in the whole system because of this skewed workload. Recent studies have demonstrated that load imbalance can be effectively mitigated by deploying a small cache node in front of the back-end servers—commonly referred to as an in-network cache. With the advent of programmable switches, it has become feasible to position this cache node directly on the switch, a critical point through which all traffic naturally flows. A class of in-network cache solutions is proposed to balance workloads among backend servers, but they fall short in supporting dynamic workloads where hotspots vary quickly with time. We show that the bottleneck is caused by the slow rule update speed on switch hardware and programming abstraction — Match-Action Table (MAT). To achieve faster cache update, we designed a new cache calledRAM cachebased on register arrays that support the write-back method. The update of theRAM cacheis done in a data-flow-driven manner, and we ensure its correctness. However, due to the hardware limitations of programmable switches, the RAM cache cannot implement complex cache update policies and has low space utilization. To address workload dynamics without sacrificing system performance, we then design acache hierarchynamed ICH, which combines the advantages of both MAT and register. We carefully devise thecache admission and evictionto keep the cache coherence, continuously available, and the server workload light. Our ICH prototype and extensive experiments demonstrate ICH improves the key-value store throughput for various access patterns (read/write intensive), especially significantly for highly dynamic workloads (up to 155%). Jiangyuan Chen, Xiaohua Xu 0002, Wenfei Wu |
IEEE Trans. Netw. | 3 |
| 2026 | INARouting: Efficient Multi-Job Routing Optimization for Hierarchical In-Network AggregationabstractIn-network aggregation (INA) has emerged as a key technology to alleviate communication bottlenecks in large-scale distributed training, but its performance is often hindered by suboptimal routing. Existing INA-aware routing algorithms suffer from certain limitations: they either lack a global, multi-job coordination mechanism, or operate on incomplete network models that ignore key hardware constraints such as switch processing capacity. These deficiencies lead to network congestion and inefficient resource utilization, ultimately undermining the full potential of INA. To address these challenges, we present INARouting, a novel framework that holistically solves the multi-job hierarchical aggregation routing problem. We propose TINA, a hierarchical aggregation protocol that supports multi-job in-network aggregation. To address different deployment scenarios, we develop two variants: INARouting-Opt that provides optimal solutions for moderate-scale networks, and INARouting-Relax, a fast and effective heuristic using LP-relaxation and a greedy score-based rounding algorithm for large-scale deployments. Through extensive experiments on various scales of Fat-Tree and Spine-Leaf topologies, we demonstrate that INARouting significantly outperforms state-of-the-art methods. INARouting- Opt achieves provably optimal solutions, reducing average job completion time by up to 56% compared to existing methods. Meanwhile, INARouting-Relax outperforms existing algorithms while being 5× faster in solving time, enabling efficient routing in large-scale, dynamic environments. Jianglong Nie, Yidan Yuan, Yuchen Xu 0003, Yitao Yuan, Kehan Yao, Lu Lu 0016, Xiaodong Duan, Wenfei Wu |
IEEE Trans. Netw. | 8 |
| 2026 | Network Functions With Dynamic State Management on Programmable SwitchesabstractRecent popular programmable switches provide programmability in packet processing at the line rate, promising to host various Network Functions (NFs). However, the limitations on the switch programmability make it unable to manage runtime dynamic NF states. In the status quo, on-switch NFs rely on the network controller for dynamic state management, suffering from high first-packet latency. This paper proposes DNF, a hardware-software co-design system for dynamic state management of network functions on programmable switches. DNF designs a dynamic State Table on the switch, which providesopportunisticservice to NF flows and sets upfallbackNFs on the server to handle flows that fail on the switch. DNF proposes a state pre-computation andruntime-bindingmethod for per-flow state initialization, which avoids a flow’s runtime interaction with the controller and the consequent latency. DNF devises astate replacementmethod, which periodically swaps hot flows into the switch. It could maximize the switch resource utilization and the throughput. We prototype DNF on P4 switches, implement six representative NFs, and compare DNF with state-of-the-art solutions. Our experiments show DNF reduces the first-packet latency by up to 98% compared with the controller-assisted architecture and improves the system throughput by up to 4.87 times compared with the software solution. Yitao Yuan, Wenfei Wu, Xin Yao 0008, Renhai Chen, Gong Zhang 0001 |
IEEE Trans. Netw. | 2 |
| 2026 | HarmonyCache: Scalable In-Network Cache With Read-Write SeparationabstractIn a key-value storage system, a small amount of hot item will account for most of the traffic. Skewed workloads can lead to load imbalance between servers, and some servers that keep hotspots will become system bottlenecks and impact the performance of the entire system. The recent studies have shown that the load imbalance can be eliminated by placing a small, fast cache node in front of the back-end servers, i.e., in-network cache. And the appearance of programmable switches made it possible to place this node on the switch, the place through which the traffic must pass. While existing in-network cache effectively balances loads between servers in large-scale storage systems, they struggle in write-intensive workloads and lose scalability with increasing client numbers due to imbalanced cache nodes. This paper introduces HarmonyCache, a scalable and high-performance in-network cache system that supports write-back methods. HarmonyCache use cache replication and read-write separation mode, only one node in all the cache is responsible for handling write requests, the other node can only handle read requests. To achieve system scalability and minimize cache coherence overhead, HarmonyCache proposes an adaptive cache replication scheme to decide where the cache should be replicated and how many replications should be made. Additionally, according to the requirements of different cache nodes, we design different types of in-network caches using different switch resources and propose a hybrid cache scheme. Our HarmonyCache prototype and extensive experiments demonstrate substantial improvements in key-value store throughput across various access patterns (read/write intensive), achieving a throughput gain of up to 7.6× over state-of-the-art solutions under skewed write-intensive workloads. Jiangyuan Chen, Xiaohua Xu 0002, Wenfei Wu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Temporal Quality as a Metric: The MORS Routing Protocol for Model Training
Chenyue Zheng, Yuchao Zhang 0004, Wenfei Wu, Zhuo Jiang, Jianglong Nie, Wendong Wang 0003 |
APNet | 3 |
| 2025 | Megabits Down to Kilobits: Memory-Efficient Time-Aware Shaping for TSNabstractTime-Sensitive Networking (TSN) provides bounded latency and low jitter for cyber-physical systems, such as industrial control. As a key component of TSN, the Time-Aware Shaper (TAS) applies gate control rules to control the transmission time of frames in critical flows. TAS stores the gate control rules for each frame in the gate control table. However, in typical industrial setups, the memory usage of the table could reach over tens of megabits and even exceed the total memory capacity of TSN switches.To address this issue, we propose a memory-efficient TAS design named METAS. It transitions from a per-frame to a per-flow approach. METAS stores one persistent rule for a flow and dynamically generates a temporary rule for a frame only when the frame arrives. We prototyped METAS on an FPGA, and experimental results show that METAS reduces memory usage from 14.34 Mbits to 288 Kbits when supporting 1,024 flows, using just 1.56% of the FPGA’s logic resources while maintaining microsecondlevel latency and nanosecond-level jitter for critical flows. Xuyan Jiang, Wenwen Fu, Xiangrui Yang 0002, Wenfei Wu |
DAC | 5 |
| 2025 | MORS: Traffic-Aware Routing based on Temporal Attributes for Model Training ClustersabstractTo train large AI models, clusters are constructed with abundant connectivity and bandwidth; but the commodity protocol ECMP and recent proposals fail to fully utilize the network bandwidth for AI traffic pattern. As model training jobs and AI clusters exhibit a predictable and periodic traffic pattern, so in this paper, we propose a MOdel training Routing System — MORS — for traffic routing in AI clusters. MORS defines temporal attributes to characterize the periodic traffic pattern of flows and network links, and temporal quality to quantify whether a path could deliver a flow quickly in the near future. MORS runs In-band Network Telemetry (INT) to collect temporal attributes of the network, and periodic analysis to extend the collected attributes in the time domain. Based on the time series of link utilization and latency, MORS computes the temporal quality of candidate paths. It enforces high-quality path selection while maintaining compatibility with commodity ECMP by manipulating the source UDP port to ensure the flow complies with the target path in the ECMP protocol. MORS is light-weight and readily deployable in the RDMA commodity cluster. Our prototype and experiments demonstrate that MORS achieves performance comparable to adaptive routing and delivers up to 14% and 50% better FCT than PLB and ECMP, respectively. Yuchao Zhang 0004, Chenyue Zheng, Wenfei Wu, Zhuo Jiang, Huichen Dai, Jianglong Nie, Wendong Wang 0003 |
ICNP | 3 |
| 2025 | FlowGram: Resource Allocation for Network Flow Measurement in CloudsabstractNetwork measurement is essential for cloud tenants in their virtual network management. However, existing measurement systems exhibit suboptimal resource efficiency while accommodating an increasing number of tenants with limited switch resources. To address this issue, this paper proposes FlowGram, a measurement framework to provide flow frequency estimation services for cloud tenants. FlowGram provides a user-friendly interface for tenants and automatically adjusts the memory allocation to meet the tenants' demands. In the data plane, FlowGram employs sketches on programmable switches to perform the measurement. In the control plane, FlowGram utilizes a precise error model of sketches and an efficient memory allocation algorithm that ensures the measurement errors are within the specified error bounds. We prototype FlowGram on programmable switches and commodity servers and conduct comprehensive experiments. Experiment results demonstrate that FlowGram outperforms the state-of-the-art system, achieving a 14 % higher task satisfaction rate while utilizing the same amount of resources. Chenqi Zhao, Wenfei Wu, Qun Huang 0001, Keqiang He |
IWQoS | 2 |
| 2025 | SGLB: Scalable and Robust Global Load Balancing in Commodity AI ClustersabstractInternet companies are constructing large-scale AI clusters with commodity Ethernet switches for AI model training to support their businesses. AI training workloads impose stringent network requirements, mandating that cluster networks deliver high peak throughput while maintaining robustness and resilience in the face of link failures. We present SGLB, a distributed, global congestion-aware load balancing system for AI clusters. SGLB operates a control-plane protocol, SyncMesh, to enable a new load balancing abstraction in modern commodity switches—Global Load Balancing (GLB) engine—which utilizes global congestion information to distribute traffic across all available paths. We address three key challenges in designing SGLB: fast routing convergence to minimize downtime in the event of link failures, scalable maintenance of congestion profiles within the constraints of limited switch hardware resources, and preventing GLB throughput suppression in scenarios where path bandwidths are asymmetric. We prototype SGLB and conduct extensive experiments to evaluate SGLB. SGLB ensures rapid routing convergence in the event of link failures, recovering in as little as 45 μs to guarantee network robustness for long-term, stable model training. Additionally, SGLB effectively load-balances traffic across paths, avoiding those with global congestion, which accelerates All-to-All collective communication by up to 60%. Chenchen Qi, Wenfei Wu, Yongcan Wang, Keqiang He, Yu-Hsiang Kao, Zongying He, Chen-Yu Yen, Zhuo Jiang, Feng Luo 0006, Surendra Anubolu, Yanjin Gao, Bingfeng Lin, Wenda Ni, Donglin Wei, Shan Ding |
SIGCOMM | 2 |
| 2025 | MimoSketch: A Framework for Frequency-Based Mining Tasks on Multiple Nodes With SketchesabstractIn distributed data stream mining, we abstract a MIMO scenario where a stream ofmultipleitems is mined bymultiple nodes. We design a framework named MimoSketch for the MIMO-specific scenario, which improves the fundamental mining tasks of item frequency estimation, item size distribution estimation, heavy hitter detection, heavy change detection, and entropy estimation. MimoSketch consists of an algorithm design and a policy to schedule items to nodes. MimoSketch's algorithm applies random counting to preserve a mathematically provenunbiasednessproperty, which makes it friendly to the aggregate query on multiple nodes; its memory layout isdynamicallyadaptive to the runtime item size distribution, which maximizes the estimation accuracy by storing more items. MimoSketch's scheduling policy balances items among nodes, avoiding nodes being overloaded or underloaded, which improves the overall mining accuracy. Our prototype and evaluation show that our algorithm can improve the accuracy of five typical mining tasks by an order of magnitude compared with the state-of-the-art solutions, and the scheduling policy further promotes the performance in MIMO scenarios. Wenfei Wu, Yuchen Xu 0003 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Exploring Communication-Efficient Federated Learning via Stateless in-Network AggregationabstractAs an ambitious training paradigm, federated learning has garnered increasing attention in recent years, which enables collaborative training of a global model without accessing users’ private data. However, due to the simultaneous and constant model updates gathering from massive distributed clients, the central server generally becomes a performance bottleneck. Additionally, the stateful aggregation (retaining all the updates from each client) conducted by the central server further poses potential threats to privacy, since it may recover the raw data based on such model updates inversely. The state-of-the-art methodologies, however, fail to address these two problems concurrently and efficiently. To this end, we propose GAIN, a secure aggregation acceleration service for federated learning. At its core, GAIN leverages programmable switches deployed at the edge network to aggregate model updates in a stateless manner before transmitting them to the central server. Consequently, GAIN can accelerate the transmission and aggregation of model updates while eliminating the chance of recovering private data. We evaluate the performance of GAIN through FPGA-based experiments and large-scale simulations. The results show that GAIN can effectively reduce bandwidth overhead and achieve up to 4.11× training throughput acceleration while prioritizing privacy protection. Junxu Xia, Geyao Cheng, Wenfei Wu, Lailong Luo, Deke Guo |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | FooDog: Empower TSN for Efficient PolicingabstractTime-Sensitive Networking (TSN) is an emerging real-time Ethernet technology that provides deterministic communication for time-sensitive (TS) traffic. At its core, TSN utilizes Per-Stream Filtering and Policing (PSFP) gates to mitigate the disruption of unavoidable frame drift. However, as first identified in this work, the naive PSFP gate design results in heavy memory usage, which hinders normal switching functions. This work proposes an efficient PSFP gate design called FooDog. FooDog employs a two-stage structure and a dual-engine policing mechanism to realize memory-efficient, logic-compact, and fast policing while maintaining minimal latency and jitter for TS traffic. Results on FPGA prototypes show that FooDog consumes only hundreds of kilobits of memory, reducing on-chip memory overheads by more than 90% compared to the unoptimized PSFP gate design. Additionally, it maintains end-to-end latency in the microsecond range and jitter below 150 nanoseconds under abnormal traffic conditions, comparable to typical TSN performance without anomalies. Xuyan Jiang, Xiangrui Yang 0002, Tongqing Zhou, Wenfei Wu, Wenwen Fu, Wei Quan 0004, Yingwen Chen 0001, Yihao Jiao, Zhigang Sun 0002 |
IEEE Trans. Netw. | 4 |
| 2025 | In-Network Aggregation as a Generic Service for Distributed ApplicationsabstractThe performance of distributed applications has long been hindered by network communication, which has emerged as a significant bottleneck. At the core of this issue, the many-to-one incast transfer stands out as one of the primary culprits. Existing works typically decompose the transmission into multiple concurrent sub-processes and utilize servers to aggregate relevant traffic, thus avoiding the incast transfer. However, limited by their theoretical bounds, these methods can only obtain limited performance improvement. In this paper, we discover that leveraging network devices for aggregating incast traffic proves highly effective in surpassing such limitations, while the advent of programmable switches further makes this envision practical. Based on this, we propose GISA as a solution for providing network acceleration across diverse distributed applications. GISA offers generic and uniform interfaces to various applications along with a switch resource sharing mechanism and policy for concurrent tasks. It also ensures correct and reliable transport while minimizing overhead through a low-overhead routing mechanism. Our FPGA-based prototype demonstrates that GISA can achieve line-rate processing when performing data aggregation with minor traffic overhead. Additionally, it supports a wide range of concurrent applications with little development effort. Junxu Xia, Wenfei Wu, Lailong Luo, Deke Guo, Geyao Cheng |
IEEE Trans. Netw. | 2 |
| 2024 | WeMu: A design of wireless network emulator
Mingtai Lv, Xiangrui Yang 0002, Huan Zhou 0006, Wenfei Wu, Yusheng Xia, Jinshu Su |
APNet | 4 |
| 2024 | Training Job Placement in Clusters with Statistical In-Network AggregationabstractIn-Network Aggregation (INA) offloads the gradient aggregation in distributed training (DT) onto programmable switches, where the switch memory could be allocated to jobs in either synchronous or statistical multiplexing mode. Statistical INA has advantages in switch memory utilization, control-plane simplicity, and management safety, but it faces the problem of cross-layer resource efficiency in job placement. This paper presents a job placement system NetPack for clusters with statistical INA, which aims to maximize the utilization of both computation and network resources. NetPack periodically batches and places jobs into the cluster. When placing a job, NetPack runs a steady state estimation algorithm to acquire the available resources in the cluster, heuristically values each server according to its available resources (GPU and bandwidth), and runs a dynamic programming algorithm to efficiently search for servers with the highest value for the job. Our prototype of NetPack and the experiments demonstrate that NetPack outperforms prior job placement methods by 45% in terms of average job completion time on production traces. Bohan Zhao, Wei Xu 0005, Shuo Liu 0002, Yang Tian 0012, Qiaoling Wang, Wenfei Wu |
ASPLOS (1) | 6 |
| 2024 | Accelerating and Securing Federated Learning with Stateless In-Network Aggregation at the EdgeabstractIn federated learning, sending the trained models (instead of raw data) from clients to the central server can surely decrease the volume of exchanged data and preserve data privacy to some extent. However, the central server can still be a system bottleneck due to the simultaneous and constant model gathering from massive distributed clients. Besides, the central server conducts stateful aggregation (retaining all the updates from each client), making it a potential threat to privacy, since it may recover the raw data based on such model updates inversely. The state-of-the-art methodologies, however, fail to address these two problems concurrently. To this end, we propose GAIN, a secure aggregation acceleration service for federated learning. At its core, GAIN aggregates the model updates at the programmable ingress switches in a stateless manner (storing the aggregated model parameters from the clients temporarily rather than permanently) before proceeding to the central server. Consequently, GAIN can accelerate the transmission and aggregation of model parameters while eliminating the chance of data recovery. We implemented a prototype of GAIN on an FPGA-based testbed to validate its performance. The results demonstrate that GAIN can achieve up to 4.11x speedup in training throughput and reduce up to 86.5% of traffic overhead. Furthermore, through theoretical analysis, we illustrate that GAIN can achieve even more substantial performance gains with a larger number of clients while guaranteeing privacy protection. Junxu Xia, Wenfei Wu, Lailong Luo, Geyao Cheng, Deke Guo, Qifeng Nian |
ICDCS | 2 |
| 2024 | Balancing Sdn Control Plane Availability and Traffic Engineering Efficiency in Data CentersabstractMany proposals have demonstrated the efficiency advantages of software-defined networking (SDN) in managing data center networks. Common practices employ centralized traffic engineering (TE) in the SDN control plane to optimize load balancing and throughput. Meanwhile, for high availability purposes, the control plane is partitioned to ensure the impact of a single faulty controller is contained. However, the interaction between these two aspects is often overlooked. In particular, we show that the current control plane partitioning approach leads to imbalanced link loads and degraded application performance. To address this issue, we propose virtual slicing, a new control plane partitioning scheme. Virtual slicing achieves desirable traffic engineering performance while retaining the availability guarantees from the current approach. Virtual slicing is implemented and evaluated with real-world and synthetic traffic traces on production spine-free data center networks. Results show that virtual slicing reduces tail link utilizations by up to 28.4 %, and improves flow completion times by up to 36 %. Brian Chang, Keqiang He, Shawn Shuoshuo Chen, Mingyang Zhang 0005, Wenfei Wu, Aditya Akella |
ICNP | 6 |
| 2023 | A Generic Service to Provide In-Network Aggregation for Key-Value StreamsabstractKey-value stream aggregation is a common operation in distributed systems, which requires intensive computation and network resources. We propose a generic in-network aggregation service for key-value streams, ASK, to accelerate the aggregation operations in diverse distributed applications. ASK is a switch-host co-designed system, where the programmable switch provides a best-effort aggregation service, and the host runs a daemon to interact with applications. ASK makes in-depth optimization tailored to traffic characteristics, hardware restrictions, and network unreliable natures: it vectorizes multiple key-value tuples’ aggregation of one packet in one switch pipeline pass, which improves the per-host’s goodput; it develops a lightweight reliability mechanism for key-value stream’s asynchronous aggregation, which guarantees computation correctness; it designs a hot-key agnostic prioritization for key-skewed workloads, which improves the switch memory utilization. We prototype ASK and use it to support Spark and BytePS. The evaluation shows that ASK could accelerate pure key-value aggregation tasks by up to 155 times and big data jobs by 3-5 times, and be backward compatible with existing INA-empowered distributed training solutions with the same speedup. Yongchao He, Wenfei Wu, Yanfang Le, Ming Liu 0027, ChonLam Lao |
ASPLOS (2) | 2 |
| 2023 | In-Network Aggregation with Transport Transparency for Distributed TrainingabstractRecent In-Network Aggregation (INA) solutions offload the all-reduce operation onto network switches to accelerate and scale distributed training (DT). On end hosts, these solutions build custom network stacks to replace the transport layer. The INA-oriented network stack cannot take advantage of the state-of-the-art performant transport layer implementation, and also causes complexity in system development and operation. Shuo Liu 0002, Qiaoling Wang, Junyi Zhang 0005, Wenfei Wu, Qinliang Lin, Yao Liu 0006, Marco Canini, Ray C. C. Cheung, Jianfei He |
ASPLOS (3) | 4 |
| 2023 | Learning To Regularized Resource Allocation with Budget ConstraintsabstractOnline resource allocation problem with budget constraints has a wide range of applications in network science and operation research. In this problem, the decision maker needs to make actions that consume resources to accumulate rewards. Contrary to prior work, we introduce a non-linear and non-separable regularizer to this problem that acts on the total resource consumption. The motivation for introducing the regularizer is allowing the decision maker to tradeoff the reward maximization and metrics optimization such as load-balancing and fairness that appears frequently in practical needs. Our goal is to simultaneously maximize additively separable rewards and the value of a non-separable regularizer without violating resource budget constraints. We develop a primal-dual-type online algorithm for this problem in the online learning setting and confirm its no-regret guarantee and zero constraint violations for stochastic i.i.d. input models. Furthermore, the general convex resource consumption functions allow our model to be more applicable. Numerical experiments are conducted to demonstrate the theoretical guarantee of our algorithm. Shaoke Fang, Qingsong Liu 0001, Wenfei Wu |
ICASSP | 4 |
| 2023 | Preemptive Switch Memory Usage to Accelerate Training Jobs with Shared In-Network AggregationabstractRecent works introduce In-Network Aggregation (INA) for distributed training (DT), which moves the gradient summation into network programmable switches. INA can reduce the traffic volume and accelerate communication in DT jobs. However, switch memory is a scarce resource, unable to support massive DT jobs in data centers, and existing INA solutions have not utilized switch memory to the best extent. We propose DSA, an Efficient Data-Plane switch memory Scheduler for in-network Aggregation. DSA introduces preemption to the switch memory management for INA jobs. In the data plane, DSA allows gradient tensors with high priority to preempt the switch aggregators (basic computation unit in INA) from tensors with low priority, which avoids an aggregator wasting time in idle. In the control plane, DSA devises a priority policy which assigns high priority to gradient tensors that benefit overall job efficiency more, e.g., communication-intensive jobs. We prototype DSA and experiments show that DSA can improve the average JCT by up to 1.35x compared with baseline solutions. Yuxuan Qin, ChonLam Lao, Yanfang Le, Wenfei Wu |
ICNP | 5 |
| 2023 | Identifying Performance Bottleneck in Shared In-Network Aggregation during Distributed TrainingabstractAs the emergence of recently popular large language model, distributed training (DT) optimizes the performance via using different parallelization strategies, resource schedulers and advanced compression techniques. Meanwhile, a promising acceleration primitive, In-Network Aggregation (INA), offloads the gradient aggregation to programmable switches to further reduce the communication overhead using the switch memory. However, to the best of our knowledge, how to identify the performance bottleneck in real time remains challenging. In this paper, we build Argus, a performance bottleneck monitoring framework for INA. Argus implements an aggregation digest extracting mechanism for real-time monitoring of DT jobs at the multi-tenant, multi-rack clusters. Argus models aggregation to identify performance bottlenecks in aggregation, which assists the scheduler in deciding the resource allocation. Extensive evaluation and prototype implementation show that Argus provides real-time and packet-level aggregation monitoring for identifying bottlenecks in INA with minimal performance overhead. Chang Liu 0001, Jiaqi Zheng 0001, Wenfei Wu, Bohan Zhao, Guihai Chen |
ICPADS | 3 |
| 2023 | In-Network Key-Value Cache with LinearizabilityabstractRecently, In-Network Cache (INC) systems have been proposed to promote the performance of remote storage systems. INC offloads cache onto programmable switches between the clients and the servers, responding to clients’ data queries within a sub-RTT time. However, most existing INC solutions do not take applications’ linearizability requirement into consideration, which could lead to query errors and storage state errors in the runtime. We propose a new INC system — NetKV-L, which preserves the high-performance I/O without compromising the linearizability. NetKV-L devises a sequentiality enforcement mechanism, a PSN correction mechanism, and a response memorization mechanism to guarantee linearizability under possible unreliable network conditions. Our prototype and experiments show that NetKV-L could achieve almost the same performance as the state-of-the-art systems while additionally guaranteeing linearizability. Yuxuan Qin, Weize Gao, ChonLam Lao, Wenfei Wu |
ICPADS | 4 |
| 2023 | Enabling Switch Memory Management for Distributed Training with In-Network Aggregation
Bohan Zhao, Jianbo Dong, Wenfei Wu |
INFOCOM | 6 |
| 2023 | AggTree: A Routing Tree With In-Network Aggregation for Distributed TrainingabstractFor distributed training (DT) based on the parameter servers (PS) architecture, the communication overhead is huge in the network for servers synchronizing parameters. In the PS architecture, the workers send gradients over the network to PS for aggregation. With the development of programmable switches, in-network aggregation (INA) is proposed to accelerate distributed training by utilizing the programmable switches in the network to implement gradients aggregation, not only at PS. However, the existing routing methods can not fully utilize the capability of INA, resulting in load imbalance and long communication time. This paper analyzes and models the routing problem in INA under the constraint of network resources. And we propose a routing algorithm named AggTree to solve this problem by searching the high-rate routing path. The result of simulations shows that AggTree can reduce communication time by 4.1%-37.9% for a single DT job and 12.7%-74.0% for multiple DT jobs compared with state-of-the-art solutions. Jianglong Nie, Wenfei Wu |
IPCCC | 2 |
| 2023 | RateSheriff: Multipath Flow-aware and Resource Efficient Rate Limiter Placement for Data Center NetworksabstractEmerging cloud services and applications request different Quality of Service (QoS) in Data Center Networks (DCNs). To meet these various requirements, programmable switch-based rate limiters are introduced to provide performance isolation and benefit from easy control and fast deployment. However, existing programmable switch-based rate limiters have two limitations: (1) multipath flows (i.e., MultiPath TCP) cannot be precisely limited, and (2) rate limiter placement solutions in DCNs are missing. These limitations could lead to poor rate limiting performance and low bandwidth utilization. In this paper, we propose RateSheriff to improve rate limiting performance by providing multipath flow-aware and resource efficient rate limiter placement for programmable switch-enabled DCNs. We identify and associate subflows to a multipath flow by extracting and comparing specific packets and header fields. By solving the formulated resource efficient rate limiter placement problem, we can improve rate limiting performance and balance memory utilization among programmable switches in DCNs. Simulation results show that RateSheriff can correctly limit the rate of multipath flows, improve rate limiting performance by up to 46%, and improve memory balancing performance by up to 79% with low computation time, compared with baselines. Songshi Dou, Yongchao He, Sen Liu 0002, Wenfei Wu, Zehua Guo 0001 |
IWQoS | 4 |
| 2023 | MimoSketch: A Framework to Mine Item Frequency on Multiple Nodes with SketchesabstractWe abstract a MIMO scenario in distributed data stream mining, where a stream of multiple items is mined by multiple nodes. We design a framework named MimoSketch for the MIMO-specific scenario, which improves the fundamental mining task of item frequency estimation. MimoSketch consists of an algorithm design and a policy to schedule items to nodes. MimoSketch's algorithm applies random counting to preserve a mathematically proven unbiasedness property, which makes it friendly to the aggregate query on multiple nodes; its memory layout is dynamically adaptive to the runtime item size distribution, which maximizes the estimation accuracy by storing more items. MimoSketch's scheduling policy balances items among nodes, avoiding nodes being overloaded or underloaded, which improves the overall mining accuracy. Our prototype and evaluation show that our algorithm can improve the item frequency estimation accuracy by an order of magnitude compared with the state-of-the-art solutions, and the scheduling policy further promotes the performance in MIMO scenarios. Yuchen Xu 0003, Wenfei Wu, Bohan Zhao, Tong Yang 0003, Yikai Zhao 0001 |
KDD | 2 |
| 2023 | NetRPC: Enabling In-Network Computation in Remote Procedure Calls
Bohan Zhao, Wenfei Wu |
NSDI | 2 |
| 2023 | ClickINC: In-network Computing as a Service in Heterogeneous Programmable Data-center NetworksabstractIn-Network Computing (INC) has found many applications for performance boosts or cost reduction. However, given heterogeneous devices, diverse applications, and multi-path network typologies, it is cumbersome and error-prone for application developers to effectively utilize the available network resources and gain predictable benefits without impeding normal network functions. Previous work is oriented to network operators more than application developers. We develop ClickINC to streamline the INC programming and deployment using a unified and automated workflow. Click-INC provides INC developers a modular programming abstractions, without concerning to the states of the devices and the network topology. We describe the ClickINC framework, model, language, workflow, and corresponding algorithms. Experiments on both an emulator and a prototype system demonstrate its feasibility and benefits. Wenquan Xu, Haoyu Song 0001, Zhikang Chen, Wenfei Wu, Guyue Liu, Yinchao Zhang, Zerui Tian, Bin Liu 0001 |
SIGCOMM | 6 |
| 2023 | Toward Flexible and Predictable Path Programmability Recovery Under Multiple Controller Failures in Software-Defined WANsabstractSoftware-Defined Networking (SDN) promises good network performance in Wide Area Networks (WANs) with the logically centralized control using physically distributed controllers. In Software-Defined WANs (SD-WANs), maintaining path programmability, which enables flexible path change on flows, is crucial for maintaining network performance under traffic variation. However, when controllers fail, existing solutions are essentially coarse-grained switch-controller mapping solutions and only recover the path programmability of a limited number of offline flows, which traverse offline switches controlled by failed controllers. In this paper, we propose FlexibleProgrammabilityMedic (FlexPM) to provide predictable path programmability recovery under multiple controller failures in SD-WANs. The key idea of FlexPM is to approximately realize flow-controller mappings using hybrid SDN/legacy routing supported by high-end commercial SDN switches. Using the hybrid routing, we can recover programmability by selecting a routing mode for each offline flow at each offline switch in a fine-grained way to fit the given control resource from active controllers and release a few control resource of active controllers by reasonably configuring some normal flows under legacy routing mode. Thus, FlexPM can promise ample control resource to improve the recovery efficiency and further effectively map offline switches to active controllers. Simulation results show that FlexPM outperforms existing switch-level solutions by maintaining balanced programmability and increasing the total programmability of recovered offline flows up to 660% under AT&T topology and 590% under Belnet topology. Zehua Guo 0001, Songshi Dou, Wenfei Wu, Yuanqing Xia |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | SFP: Service Function Chain Provision on Programmable Switches for Cloud TenantsabstractRecent progress in programmable switches provides opportunities for service function chains (SFCs) provision to cloud tenants, which has the advantage of flexible deployment and high performance. We devise SFP for such SFC provision in the cloud. SFP's data plane installs physical NFs and is virtualized to host logical SFCs from multiple tenants. SFP's control plane uses a relaxed integer programming model to jointly optimize the placement of physical and logical NFs, which can achieve resource efficiency and high tenant traffic processing throughput within efficient execution time. Our prototype and evaluation shows that SFP can significantly offload NFV computation from server to the switch and maximize the switch resource utilization. Hongyi Huang, Wenfei Wu, Yongchao He, Zehua Guo 0001 |
IPDPS | 2 |
| 2022 | NQ/ATP: Architectural Support for Massive Aggregate Queries in Data Center NetworksabstractNetwork queries become increasingly challenging for online service providers with massive network devices and massive network queries due to the tradeoff between system scale and query granularity. We re-architect the traditional three-tier architecture, i.e., data collection, data storage, and data query, for aggregate queries, and build a system named NQ/ATP. NQ/ATP offloads the aggregation operation in network queries onto network switches, which accelerates the query execution and frees up network resources. NQ/ATP further devises a route learning mechanism, query hierarchy load balancing policy, and hierarchy clustering mechanism to save forwarding table entries on switches, which better supports massive queries. The evaluation shows that NQ/ATP can support network aggregate queries with higher capacity, less traffic volume, finer granularity, and better scalability than traditional three-tier polling architectures. The three optimizations can effectively reduce the forwarding table usage by up to 97.55%. Wenfei Wu, Shan-Hsiang Shen, Ying Zhang 0022 |
IWQoS | 2 |
| 2022 | Consistent and Fine-Grained Rule Update with In-Network Control for Distributed Rate LimitingabstractCoexisting applications contend for limited WAN bandwidth when communicating over distributed data centers in private clouds. Online service providers deploy distributed rate limiting systems to dynamically estimate each application instance’s bandwidth demand, and update rate limiting rules to provide performance isolation and bandwidth guarantees for applications with different priorities. However, the isolation violation caused by the inconsistent update of rate limiting rules among servers would eventually violate the Service Level Agreement requirements. Motivated by the observation that InNetwork Control can update rate limiting rules consistently with ultra-low latency for hundreds of thousands of end-hosts through in-band control messages, this paper presents DistRL, a dynamic distributed rate limiting system that can achieve fine-grained consistent updates. The main idea of DistRL is to replace the traditional controller with programmable switches to improve communication efficiency, thereby achieving finer-grained consistent updates. The evaluation shows that DistRL can support sub-second distributed updates of rate limiting rules without isolation violation for O(105) servers. Yongchao He, Wenfei Wu |
IWQoS | 2 |
| 2022 | HyperSFP: Fault-Tolerant Service Function Chain Provision on Programmable Switches in Data CentersabstractWith the cloud networks being equipped with programmable switch, Service Function Chain (SFC) provision has started to be migrated to the switches for better performance and manageability. In this paper, we design HyperSFP which places multiple SFCs to a data center network (DCN). In the placement algorithm, HyperSFP builds an integer programming (IP) model to achieve functionality, fault tolerance, and load balance. To support large-scale networks, HyperSFP IP model is relaxed to two approximate approaches: Stage-Separated IP model and linear programming (LP) model. Both approaches can improve the algorithm efficiency. HyperSFP’s data plane is designed to deploy the active and backup NFs in the control-plane plan, and migrate traffic from failed active NFs to its backup NFs. Our prototype and evaluation shows that HyperSFP achieves performance gain by implementing NFs on programmable switches, its control plane achieves fault tolerance, load balance, and scalability, and its data plane can handle network failures promptly. Hongyi Huang, Wenfei Wu |
NOMS | 2 |
| 2022 | WRS: Workflow Retrieval System for Cloud Automatic RemediationabstractRemediation systems in modern information infrastructure gradually take over the increasing workload of unexpected events from operators. The core logic in such systems — remediation rule— still needs operators to fill in manually, which is a tedious and error-prone process. We propose Workflow Retrieval System (a.k.a. WRS), which helps the operator to build remediation rules.WRS is a recommendation system that recommends structures from existing rules to operators when they are creating new ones. WRS formalizes the workflows in remediation rules as trees and extracts representative atomic structures from them. Then WRS organizes atomic structures in two ways to accelerate the later retrieval — a two-level hierarchy which helps searching similar structures, and a keyword indexing structure which helps searching by words. With the two retrieval structures, we build two applications: one is real-time workflow auto completion application which recommends remaining structures based on the existing partial structure, and another is workflow recommendation where WRS recommends atomic structures based on the description of the remediation symptoms.We prototype WRS and evaluate it with legacy device vendors’ operation manual. Our evaluation shows that WRS reasonably extracts and organizes atomic structures in remediation workflows, and the two applications atop it show fast execution time and high accuracy. Hongyi Huang, Wenfei Wu, Shimin Tao |
NOMS | 2 |
| 2022 | Detecting Outlier Machine Instances Through Gaussian Mixture Variational Autoencoder With One Dimensional CNNabstractToday's large datacenters house a massive number of machines, each of which is being closely monitored with multivariate time series (e.g., CPU idle, memory utilization) to ensure service quality. Detecting outlier machine instances with multivariate time series is crucial for service management. However, it is a challenging task due to the multiple classes and various shapes, high dimensionality, and lack of labels of multivariate time series. In this article, we propose DOMI, a novel unsupervised model that combines Gaussian mixture VAE with 1D-CNN, todetectoutliermachineinstances. Its core idea is to capture the normal patterns of machine instances by learning their latent representations that consider the shape characteristics, reconstruct input data by the learned representations, and apply reconstruction probabilities to determine outliers. Moreover, DOMI interprets the detected outlier instance based on the reconstruction probability changes of univariate time series. Extensive experiments have been conducted on the dataset collected from 1821 machines with a 1.5-month-period, which are deployed in ByteDance, a top global content service provider. DOMI achieves the best F1-Score of 0.94 and AUC score of 0.99, significantly outperforming the best performing baseline method by 0.08 and 0.03, respectively. Moreover, its interpretation accuracy is up to 0.93. Ya Su, Youjian Zhao, Shenglin Zhang, Xidao Wen, Yongsu Zhang, Junliang Tang, Wenfei Wu, Dan Pei |
IEEE Trans. Computers | 10 |
| 2021 | HyperNAT: Scaling Up Network Address Translation with SmartNICs for CloudsabstractNetwork address translation (NAT) is a basic functionality in cloud gateways. With the increasing traffic volume and number of flows introduced by the cloud tenants, the NAT gateway needs to be implemented on a cluster of servers. We propose to scale up the gateway servers, which could reduce the number of servers so as to reduce the capital expense and operation expense. We design HyperNAT, which leverages smartNICs to improve the server's processing capacity. In HyperNAT, the NAT functionality is distributed on multiple NICs, and the flow space is divided and assigned accordingly. HyperNAT overcomes the challenge that the packets in two directions of one connection need to be processed by the same NAT rule (named two-direction consistency, TDC) by cloning the rule to both data paths of the two directions. Our implementation and evaluation of HyperNAT show that HyperNAT could scale up cloud gateway effectively with low overhead. Shaoke Fang, Qingsong Liu 0001, Wenfei Wu |
GLOBECOM | 3 |
| 2021 | NFReducer: Redundant Logic Elimination for Network Functions with Runtime ConfigurationsabstractNetwork functions (NFs) are critical components in the network data plane. Their efficiency is important to the whole network's end-to-end performance. We identify three types of runtime redundant logic in individual NF and NF chains when they are deployed with concrete configured rules. We use program analysis techniques to optimize away the redundancy where we also overcome the NF specific challenges - we combine symbolic execution and dead code elimination to eliminate unused logic, we customize the common sub-expression elimination to eliminate duplicated logic, and we add network semantics to the dead code elimination to eliminate overwritten logic. We implement a prototype named NFReducer using LLVM. Our evaluation on both legacy and platform NFs shows that after eliminating the redundant logic, the packet processing rate of the NFs can be significantly improved and the operational overhead is small. Bangwen Deng, Wenfei Wu |
INFOCOM | 2 |
| 2021 | Scalable On-Switch Rate Limiters for the CloudabstractWhile most clouds use on-server rate limiters for bandwidth allocation, we propose to implement them on switches. On-switch rate limiters can simplify network management and promote the performance of control-plane rate limiting applications. We leverage the recent progress of programmable switches to implement on-switch rate limiters, named SwRL. In the design of SwRL, we make design choices according to the programmable hardware characteristics, we deeply optimize the memory usage of the algorithm so as to fit a cloud-scale (one million) rate limiters in a single switch, and we complement the missing computation primitives of the hardware using a pre-computed approximate table. We further developed three control-plane applications and integrate them with SwRL, showing the control-plane interoperability of SwRL. We prototype and evaluate SwRL in both testbed and production environments, demonstrating its good properties of precision rate control, scalability, interoperability, and manageability (execution environmental isolation). Yongchao He, Wenfei Wu, Xuemin Wen, Yongqiang Yang |
INFOCOM | 2 |
| 2021 | NFD: Using Behavior Models to Develop Cross-Platform Network FunctionsabstractNFV ecosystem is flourishing and more and more NF platforms appear, but this makes NF vendors difficult to deliver NFs rapidly to diverse platforms. We propose an NF development framework named NFD for cross-platform NF development. NFD's main idea is to decouple the functional logic from the platform logic -it provides a platform-independent language to program NFs' behavior models, and a compiler with interfaces to develop platform-specific plugins. By enabling a plugin on the compiler, various NF models would be compiled to executables integrated with the target platform. We prototype NFD, build 14 NFs, and support 6 platforms (standard Linux, OpenNetVM, GPU, SGX, DPDK, OpenNF). Our evaluation shows that NFD can save development workload for cross-platform NFs and output valid and performant NFs. Hongyi Huang, Wenfei Wu, Yongchao He, Bangwen Deng, Ying Zhang 0022, Yongqiang Xiong, Guo Chen 0001, Yong Cui 0001, Peng Cheng 0005 |
INFOCOM | 2 |
| 2021 | CTF: Anomaly Detection in High-Dimensional Time Series with Coarse-to-Fine Model TransferabstractAnomaly detection is indispensable in modern IT infrastructure management. However, the dimension explosion problem of the monitoring data (large-scale machines, many key performance indicators, and frequent monitoring queries) causes a scalability issue to the existing algorithms. We propose a coarse-to-fine model transfer based framework CTF to achieve a scalable and accurate data-center-scale anomaly detection. CTF pre-trains a coarse-grained model, uses the model to extract and compress per-machine features to a distribution, clusters machines according to the distribution, and conducts model transfer to fine-tune per-cluster models for high accuracy. The framework takes advantage of clustering on the per-machine latent representation distribution, reusing the pre-trained model, and partial-layer model fine-tuning to boost the whole training efficiency. We also justify design choices such as the clustering algorithm and distance algorithm to achieve the best accuracy. We prototype CTF and experiment on production data to show its scalability and accuracy. We also release a labeling tool for multivariate time series and a labeled dataset to the research community. Ya Su, Shenglin Zhang, Yuanpu Cao, Dan Pei, Wenfei Wu, Yongsu Zhang, Junliang Tang |
INFOCOM | 7 |
| 2021 | DHS: Adaptive Memory Layout Organization of Sketch Slots for Fast and Accurate Data Stream ProcessingabstractData stream processing is a crucial computation task in data mining applications. The rigid and fixed data structures in existing solutions limit their accuracy, throughput, and generality in measurement tasks. We propose Dynamic Hierarchical Sketch (DHS), a sketch-based hybrid solution targeting these properties. During the online stream processing, DHS hashes items to buckets and organizes cells in each bucket dynamically; the size of all cells in a bucket is adjusted adaptively to the actual size and distribution of flows. Thus, memory is efficiently used to precisely record elephant flows and cover more mice flows. Implementation and evaluation show that DHS achieves high accuracy, high throughput, and high generality on five measurement tasks: flow size estimation, flow size distribution estimation, heavy hitter detection, heavy changer detection, and entropy estimation. Bohan Zhao, Xiang Li 0156, Boyu Tian, Zhiyu Mei, Wenfei Wu |
KDD | 5 |
| 2021 | ATP: In-network Aggregation for Multi-tenant Learning
ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Wenfei Wu, Aditya Akella, Michael M. Swift |
NSDI | 5 |
| 2021 | Sphinx: A transport protocol for high-speed and lossy mobile networks
Dan Li 0001, Wenfei Wu, K. K. Ramakrishnan, Jinkun Geng, Fanzhao Wang, Kai Zheng 0003 |
Comput. Networks | 3 |
| 2021 | Simultaneously achieving sublinear regret and constraint violations for online convex optimization with time-varying constraints
Qingsong Liu 0001, Wenfei Wu, Longbo Huang, Zhixuan Fang |
Perform. Evaluation | 2 |
| 2020 | T2DNS: A Third-Party DNS Service with Privacy Preservation and TrustworthinessabstractWe design a third-party DNS service named T2DNS. T2DNS serves client DNS queries with the following features: protecting clients from channel and server attackers, providing trustworthiness proof to clients, being compatible with the existing Internet infrastructure, and introducing bounded overhead. T2DNS’s privacy preservation is achieved by a hybrid protocol of encryption and obfuscation, and its service proxy is implemented on Intel SGX. We overcome the challenges of scaling the initialization process, bounding the obfuscation overhead, and tuning practical system parameters. We prototype T2DNS, and experiment results show that T2DNS is fully functional, has acceptable overhead in comparison with other solutions, and is scalable to the number of clients. Qingxiu Liu, Wenfei Wu, Qun Huang 0001 |
ICCCN | 2 |
| 2020 | Symbolic Execution for Network Functions with Time-Driven LogicabstractSymbolic Execution is a commonly used technique in network function (NF) verification, and it helps network operators to find implementation or configuration bugs before the deployment. By studying most existing symbolic execution engine, we realize that they only focus on packet arrival based event logic; we propose that NF modeling language should include time-driven logic to describe the actual NF implementations more accurately and performing complete verification. Thus, we define primitives to express time-driven logic in NF modeling language and develop a symbolic execution engine NF-SE that can verify such logic for NFs for multiple packets. Our prototype of NF-SE and evaluation on multiple example NFs demonstrate its usefulness and correctness. Harsha Sharma, Wenfei Wu, Bangwen Deng |
MASCOTS | 2 |
| 2019 | SpeedyBox: Low-Latency NFV Service Chains with Cross-NF Runtime ConsolidationabstractSoftware-based service chains in Network Function Virtualization (NFV) typically suffers high processing latency. This latency grows as chain lengths increase and possibly violates application requirements. Previous efforts focus on reducing latency while maintaining the perspective of each NF being an independent, isolated module. This results in processing redundancy that could eventually become the performance bottleneck. In this paper, we propose a low-latency NFV framework called SpeedyBox, that innovatively enables cross-NF runtime optimizations in a service chain to eliminate processing redundancy. SpeedyBox builds a fast data path for flows at runtime by consolidating the aggregate actions across diverse network functions (NFs) in a service chain. In SpeedyBox, each NF is instrumented with a stateful Local Match-Action Table (MAT), and leverages our easy-to-use APIs to record its per-flow behavior in the Local MAT. Next, SpeedyBox uses a Global MAT to build the fast data path by consolidating actions from each Local MAT, while providing the ability to express the stateful NF behaviors with an Event Table. We have implemented a prototype of SpeedyBox on the BESS and OpenNetVM NFV platforms. Our trace-driven evaluation on common NFs shows that SpeedyBox achieves significant latency reduction under real world scenarios. Yong Cui 0001, Wenfei Wu, Jiahan Gu, K. K. Ramakrishnan, Yongchao He, Xuehai Qian |
ICDCS | 3 |
| 2019 | Sphinx: A Transport Protocol for High-Speed and Lossy Mobile NetworksabstractModern mobile wireless networks have been demonstrated to be high-speed but lossy, while mobile applications have more strict requirements including reliability, goodput guarantee, bandwidth efficiency, and computation efficiency. Such a complicated combination of requirements and conditions in networks pushes the pressure to transport layer protocol design. We analyze and argue that few of existing network transport layer solutions are able to handle all these requirements. We design and implement Sphinx to satisfy the four requirements in high-speed and lossy networks. Sphinx has (1) a proactive coding-based method named semi-random LT codes for loss recovery, which estimates packet loss rate and adjusts the redundancy level accordingly, (2) a reactive retransmission method named Instantaneous Compensation Mechanism (ICM) for loss retransmission, which compensates the lost packets once actual loss exceeds the estimation, and (3) a parallel coding architecture, which leverages multi-core, shared memory and kernel-bypass DPDK. Prototype and evaluation show that Sphinx outperforms TCP and other coding solutions significantly in microbenchmarks across all four requirements, and improves the performance of applications such as video streaming and block data transfer. Dan Li 0001, Wenfei Wu, K. K. Ramakrishnan, Jinkun Geng, Fei Gui, Fanzhao Wang, Kai Zheng 0003 |
IPCCC | 3 |
| 2019 | Alembic: Automated Model Inference for Stateful Network Functions
Soo-Jin Moon, Jeffrey Helt, Yves Bieri, Sujata Banerjee, Vyas Sekar, Wenfei Wu, Mihalis Yannakakis, Ying Zhang 0022 |
NSDI | 7 |
| 2018 | Improving Quality of Experience for Mobile Broadcasters in Personalized Live Video StreamingabstractEnsuring high video quality of experience (QoE) on the broadcaster side is critical for interactive live streaming. However, measurements on multiple live streaming platforms show that they all suffer from broadcaster-side video quality degradation in the presence of transient bandwidth fluctuations. This paper presents Greedy Variable Bitrate (GVBR), a suite of solutions that optimizes the QoE through an approriate keyframe interval that trades cross-frame compression for lowered inter-frame interdependency, a simple-yet-efficient frame dropping strategy to prevent excessive frame drops, and a bitrate adaptation strategy customized for broadcasters who have shallow buffer. We compare GVBR with state-of-art algorithms in different network conditions, and find that GVBR can cut video interruption incidents by 90%, while achieving comparable bitrate. Qingmei Ren, Yong Cui 0001, Wenfei Wu, Changfeng Chen, Yuchi Chen, Jiangchuan Liu, Hongyi Huang |
IWQoS | 3 |
| 2018 | NetCP: Consistent, Non-Interruptive and Efficient Checkpointing and Rollback of SDNabstractNetwork failures are inevitable due to its increasing complexity, which significantly hampers system availability and performance. While adopting checkpointing and rollback recovery protocols (C/R for abbreviation) from distributed systems into computer networks is promising, several specific challenges appear as we design a C/R system for Software-Defined Networks (SDN). The C/R should be coordinated with other applications in the SDN controller, each individual switch C/R should not interrupt traffic traversing it, and SDN controller C/R faces the challenge of time and space overhead. We propose a C/R framework for SDN, named NetCP. NetCP coordinates C/R and other applications to get consistent global checkpoints, it leverages redundant forwarding tables in SDN switches for C/R so as to avoid interrupting traversing traffic, and it analyzes the dependencies between controller applications to make minimal C/R decision. We have implemented NetCP in a prototype system using the current standard SDN tools and demonstrate that it achieves consistency, non-interruption, and efficiency with negligible overhead. Ye Yu 0001, Chen Qian 0001, Wenfei Wu, Ying Zhang 0022 |
IWQoS | 3 |
| 2017 | Low Latency Software Rate Limiters for Cloud NetworksabstractA lot of recent work has focused on reducing in network queueing latency in datacenter networks. In this paper, we focus on a less explored topic --- latency increases caused by queueing in rate limiters on the end-host. First, we show that latency can be increased by an order of magnitude by rate limiters in cloud networks. To solve this problem, we extend ECN marking into rate limiters and use a datacenter congestion control algorithm --- DCTCP. Unfortunately, while this reduces latency, it also leads to throughput oscillation. Thus, this solution is not sufficient. In this paper, we also analyze the specific reasons that ECN marking in software rate limiters leads to the throughput oscillation problem. Finally, we propose two potential solutions to design software rate limiters that can achieve stable high throughput and low latency. Keqiang He, Weite Qin, Wenfei Wu, Tian Pan 0001, Chengchen Hu, Jiao Zhang 0002, Brent E. Stephens, Aditya Akella, Ying Zhang 0022 |
APNet | 4 |
| 2017 | Supporting Diverse Dynamic Intent-based Policies using JanusabstractExisting network policy abstractions handle basic group based reachability and access control list based security policies. However, QoS policies as well as dynamic policies are also important and not representing them in the high level policy abstraction poses serious limitations. At the same time, efficiently configuring and composing group based QoS and dynamic policies present significant technical challenges, such as (a) maintaining group granularity during configuration, (b) dealing with network-bandwidth contention among policies from distinct writers and (c) dealing with multiple path changes corresponding to dynamically changing policies, group membership and end-point mobility. In this paper we propose Janus, a system which makes two major contributions. First, we extend the prior policy graph abstraction model to represent complex QoS and dynamic tateful/temporal policies. Second, we convert the policy configuration problem into an optimization problem with the goal of maximizing the number of satisfied and configured policies, and minimizing the number of path changes under dynamic environments. To solve this, Janus presents several novel heuristic algorithms. We evaluate our system using a diverse set of bandwidth policies and network topologies. Our experiments demonstrate that Janus can achieve near-optimal solutions in a reasonable amount of time. Anubhavnidhi Abhashkumar, Joon-Myung Kang, Sujata Banerjee, Aditya Akella, Ying Zhang 0022, Wenfei Wu |
CoNEXT | 6 |
| 2017 | SLA-verifier: Stateful and quantitative verification for service chainingabstractNetwork verification has been recently proposed to detect network misconfigurations. Existing work focuses on the reachability. This paper proposes a framework that verifies the Service Level Agreement (SLA) compliance of the network using static verification. This work proposes a quantitative model and a set of algorithms for verifying performance properties of a network with switches and middleboxes, i.e., service chains. We develop SLA-Verifier and evaluate its efficiency using simulation on real-world data and testbed experiments. To improve the SLA violation detection accuracy, our system uses verification results to optimize online monitoring. Ying Zhang 0022, Wenfei Wu, Sujata Banerjee, Joon-Myung Kang, Mario A. Sánchez |
INFOCOM | 2 |
| 2016 | Automatic Synthesis of NF Models by Program AnalysisabstractNetwork functions (NFs), like firewall, NAT, IDS, have been widely deployed in today’s modern networks. However, currently there is no standard specification or modeling language that can accurately describe the complexity and diversity of different NFs. Recently there have been research efforts to propose NF models. However, they are often generated manually and thus error-prone. This paper proposes a method to automatically synthesize NF models via program analysis. We develop a tool called NFactor, which conducts code refactoring and program slicing on NF source code, in order to generate its forwarding model. We demonstrate its usefulness on two NFs and evaluate its correctness. A few applications of NFactor are described, including network verification. Wenfei Wu, Ying Zhang 0022, Sujata Banerjee |
HotNets | 1 |
| 2015 | Management Plane AnalyticsabstractWhile it is generally held that network management is tedious and error-prone, it is not well understood which specific management practices increase the risk of failures. Indeed, our survey of 51 network operators reveals a significant diversity of opinions, and our characterization of the management practices in the 850+ networks of a large online service provider shows significant diversity in prevalent practices. Motivated by these observations, we develop a management plane analytics (MPA) framework that an organization can use to: (i) infer which management practices impact network health, and (ii) develop a predictive model of health, based on observed practices, to improve network management. We overcome the challenges of sparse and skewed data by aggregating data from many networks, reducing data dimensionality, and oversampling minority cases. Our learned models predict network health with an accuracy of 76-89%, and our causal analysis uncovers some high impact practices that operators thought had a low impact on network health. Our tool is publicly available, so organizations can analyze their own management practices. Aaron Gember, Wenfei Wu, Xiujun Li, Aditya Akella, Ratul Mahajan |
Internet Measurement Conference | 2 |
| 2015 | PerfSight: Performance Diagnosis for Software DataplanesabstractThe advent of network functions virtualization (NFV) means that data planes are no longer simply composed of routers and switches. Instead they are very complex and involve a variety of sophisticated packet processing elements that reside on the OSes and software running on compute servers where network functions (NFs) are hosted. In this paper, we argue that these new "software data planes" are susceptible to at least three new classes of performance problems. To diagnose such problems, we design, implement and evaluate, PerfSight, a ground-up system that works by extracting comprehensive low-level information regarding packet processing and I/O performance of the various elements in the software data plane. Name then analyzes the information gathered in various dimensions (e.g., across all VMs on a machine, or all VMs deployed by a tenant). By looking across aggregates, we show that it becomes possible to detect and diagnose key performance problems. Experimental results show that our framework can result in accurate detection of the root causes of key performance problems in software data planes, and it imposes very little overhead. Wenfei Wu, Keqiang He, Aditya Akella |
Internet Measurement Conference | 1 |
| 2014 | SoftMoW: Recursive and Reconfigurable Cellular WAN ArchitectureabstractThe current LTE network architecture is organized into very large regions, each having a core network and a radio access network. The core network contains an Internet edge comprised of packet data network gateways (PGWs). The radio network consists of only base stations. There are minimal interactions among regions other than interference management at the edge. The current architecture has several problems. First, mobile application performance is seriously impacted by the lack of Internet egress points per region. Second, the continued exponential growth of mobile traffic puts tremendous pressure on the scalability of PGWs. Third, the fast growth of signaling traffic known as the signaling storm problem poses a major challenge to the scalability of the control plane. To address these problems, we present SoftMoW, a recursive and reconfigurable cellular WAN architecture that supports seamlessly inter-connected core networks, reconfigurable control plane, and global optimization. Mehrdad Moradi, Wenfei Wu, Li Erran Li, Z. Morley Mao |
CoNEXT | 2 |
| 2014 | PRAN: Programmable Radio Access NetworksabstractWith the continued exponential growth of mobile traffic and the rise of diverse applications, the current LTE radio access network (RAN) architecture of cellular operators face mounting challenges. Current RAN suffers from insufficient radio resource coordination, inefficient infrastructure utilization, and inflexible data paths. We present the high level design of PRAN, which centralizes base stations' L1/L2 processing into a cluster of commodity servers. PRAN uses a flexible data path model to support new protocols; multiple base stations' L1/L2 processing tasks are scheduled on servers with performance guarantees; and a RAN scheduler coordinates the allocation of shared radio resources between operators and base stations. Our evaluation shows the feasibility of fast data path control and efficiency of resource pooling (a potential for a 30× reduction on resources). Wenfei Wu, Li Erran Li, Aurojit Panda, Scott Shenker |
HotNets | 1 |
| 2013 | Virtual network diagnosis as a serviceabstractToday's cloud network platforms allow tenants to construct sophisticated virtual network topologies among their VMs on a shared physical network infrastructure. However, these platforms provide little support for tenants to diagnose problems in their virtual networks. Network virtualization hides the underlying infrastructure from tenants as well as prevents deploying existing network diagnosis tools. This paper makes a case for providing virtual network diagnosis as a service in the cloud. We identify a set of technical challenges in providing such a service and propose a Virtual Network Diagnosis (VND) framework. VND exposes abstract configuration and query interfaces for cloud tenants to troubleshoot their virtual networks. It controls software switches to collect flow traces, distributes traces storage, and executes distributed queries for different tenants for network diagnosis. It reduces the data collection and processing overhead by performing local flow capture and on-demand query execution. Our experiments validate VND's functionality and shows its feasibility in terms of quick service response and acceptable overhead; our simulation proves the VND architecture scales to the size of a real data center network. Wenfei Wu, Aditya Akella, Anees Shaikh |
SoCC | 1 |
| 2013 | Adaptive data transmission in the cloudabstractData centers provide resources for a broad range of services, such as web search, email, web sites, etc., each with different delay requirements. For example, web search should cater to users' requests quickly, while data backup has no special requirement on completion time. Different applications also introduce flows with very different properties (e.g., size and duration). The default method of transport in data centers, namely TCP, treats flows equally, forcing equal share of the bottleneck network bandwidth. This fairness property leads to poor outcomes for time-sensitive applications. A better solution is to allocate more bandwidth to time-sensitive applications. However, the state-of-the-art approaches that do this all require forklift changes to data center networking gear. In some cases, substantial changes need to be made to end-system stacks and applications as well. In this paper, we argue that a simple modification to TCP can help better meet the requirements of latency-sensitive applications in the data center. No modification to end-systems, applications or networking gear is necessary. We motivate our Adaptive TCP (ATCP) design using measurements of real data center traffic. We analytically derive the parameters to use in our proposed modification to TCP. Finally, we use extensive simulations in NS2 to show the benefits of ATCP. Wenfei Wu, Yizheng Chen 0005, Ramakrishnan Durairajan, Ashok Anand, Aditya Akella |
IWQoS | 1 |
| 2012 | XIA: Efficient Support for Evolvable Internetworking
Dongsu Han, Ashok Anand, Fahad R. Dogar, Hyeontaek Lim, Michel Machado, Arvind Mukundan, Wenfei Wu, Aditya Akella, David G. Andersen, John W. Byers, Srinivasan Seshan, Peter Steenkiste |
NSDI | 8 |
| 2012 | DAC: Generic and Automatic Address Configuration for Data Center NetworksabstractData center networks encode locality and topology information into their server and switch addresses for performance and routing purposes. For this reason, the traditional address configuration protocols such as DHCP require a huge amount of manual input, leaving them error-prone. In this paper, we present DAC, a generic and automatic Data center Address Configuration system. With an automatically generated blueprint that defines the connections of servers and switches labeled by logical IDs, e.g., IP addresses, DAC first learns the physical topology labeled by device IDs, e.g., MAC addresses. Then, at the core of DAC is its device-to-logical ID mapping and malfunction detection. DAC makes an innovation in abstracting the device-to-logical ID mapping to the graph isomorphism problem and solves it with low time complexity by leveraging the attributes of data center network topologies. Its malfunction detection scheme detects errors such as device and link failures and miswirings, including the most difficult case where miswirings do not cause any node degree change. We have evaluated DAC via simulation, implementation, and experiments. Our simulation results show that DAC can accurately find all the hardest-to-detect malfunctions and can autoconfigure a large data center with 3.8 million devices in 46 s. In our implementation, we successfully autoconfigure a small 64-server BCube network within 300 ms and show that DAC is a viable solution for data center autoconfiguration. Kai Chen 0005, Chuanxiong Guo, Zhenqian Feng, Yan Chen 0004, Songwu Lu, Wenfei Wu |
IEEE/ACM Trans. Netw. | 8 |
| 2011 | Routing Optimization for Ensemble RoutingabstractThe Ensemble Routing architecture (presented at ANCS 2010) implements multipath routing for data center networks. Rather than managing individual flows, ensemble routing manages flows in groups or ensembles to provide scalable responsive management using simple hardware. Ensemble Routing combines: routing VLANs that define a set of diverse paths through complex networks, and load balancing algorithms to split traffic among those VLANS and optimize traffic flow. This extended abstract describes improved algorithms for the formation of routing VLANs and traffic load-balancing for Ensemble routing. The VLAN formation algorithms are improved by incorporating traffic flow estimates into the VLAN formation heuristics. The previous load balancing algorithm using a greedy heuristic is replaced by linear programming that determines optimal traffic splitting among VLANs. Simulations show that these mechanisms significantly enhance performance. Wenfei Wu, Yoshio Turner, Mike Schlansker |
ANCS | 1 |
| 2011 | XIA: an architecture for an evolvable and trustworthy internetabstractMotivated by limitations in today's host-based IP network architecture, recent studies have proposed clean-slate network architectures centered around alternative first-class principals, such as content, services, or users. However, much like the host-centric IP design, elevating one principal type above others hinders communication between other principals and inhibits the network's capability to evolve. Our work presents the eXpressive Internet Architecture (XIA), an architecture with native support for multiple principals and the ability to evolve its functionality to accommodate new, as yet unforeseen, principals over time. XIA also provides intrinsic security: communicating entities validate that their underlying intent was satisfied correctly without relying on external databases or configuration. Ashok Anand, Fahad R. Dogar, Dongsu Han, Hyeontaek Lim, Michel Machado, Wenfei Wu, Aditya Akella, David G. Andersen, John W. Byers, Srinivasan Seshan, Peter Steenkiste |
HotNets | 7 |
| 2010 | SecondNet: a data center network virtualization architecture with bandwidth guaranteesabstractIn this paper, we propose virtual data center (VDC) as the unit of resource allocation for multiple tenants in the cloud. VDCs are more desirable than physical data centers because the resources allocated to VDCs can be rapidly adjusted as tenants' needs change. To enable the VDC abstraction, we design a data center network virtualization architecture called SecondNet. SecondNet achieves scalability by distributing all the virtual-to-physical mapping, routing, and bandwidth reservation state in server hypervisors. Its port-switching based source routing (PSSR) further makes SecondNet applicable to arbitrary network topologies using commodity servers and switches. SecondNet introduces a centralized VDC allocation algorithm for bandwidth guaranteed virtual to physical mapping. Simulations demonstrate that our VDC allocation achieves high network utilization and low time complexity. Our implementation and experiments show that we can build SecondNet on top of various network topologies, and SecondNet provides bandwidth guarantee and elasticity, as designed. Chuanxiong Guo, Guohan Lu, Helen J. Wang, Chao Kong, Wenfei Wu, Yongguang Zhang |
CoNEXT | 7 |
| 2010 | Generic and automatic address configuration for data center networksabstractData center networks encode locality and topology information into their server and switch addresses for performance and routing purposes. For this reason, the traditional address configuration protocols such as DHCP require huge amount of manual input, leaving them error-prone.In this paper, we present DAC, a generic and automatic Data center Address Configuration system. With an automatically generated blueprint which defines the connections of servers and switches labeled by logical IDs, e.g., IP addresses, DAC first learns the physical topology labeled by device IDs, e.g., MAC addresses. Then at the core of DAC is its device-to-logical ID mapping and malfunction detection. DAC makes an innovation in abstracting the device-to-logical ID mapping to the graph isomorphism problem, and solves it with low time-complexity by leveraging the attributes of data center network topologies. Its malfunction detection scheme detects errors such as device and link failures and miswirings, including the most difficult case where miswirings do not cause any node degree change.We have evaluated DAC via simulation, implementation and experiments. Our simulation results show that DAC can accurately find all the hardest-to-detect malfunctions and can autoconfigure a large data center with 3.8 million devices in 46 seconds. In our implementation, we successfully autoconfigure a small 64-server BCube network within 300 milliseconds and show that DAC is a viable solution for data center autoconfiguration. Kai Chen 0005, Chuanxiong Guo, Zhenqian Feng, Yan Chen 0004, Songwu Lu, Wenfei Wu |
SIGCOMM | 8 |