Yanfang Le

dblp:148/1943 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
4since 2021 · last 2023
0000-0002-5680-9015ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 A Generic Service to Provide In-Network Aggregation for Key-Value Streams
abstract
Key-value stream aggregation is a common operation in distributed systems, which requires intensive computation and network resources. We propose a generic in-network aggregation service for key-value streams, ASK, to accelerate the aggregation operations in diverse distributed applications. ASK is a switch-host co-designed system, where the programmable switch provides a best-effort aggregation service, and the host runs a daemon to interact with applications. ASK makes in-depth optimization tailored to traffic characteristics, hardware restrictions, and network unreliable natures: it vectorizes multiple key-value tuples’ aggregation of one packet in one switch pipeline pass, which improves the per-host’s goodput; it develops a lightweight reliability mechanism for key-value stream’s asynchronous aggregation, which guarantees computation correctness; it designs a hot-key agnostic prioritization for key-skewed workloads, which improves the switch memory utilization. We prototype ASK and use it to support Spark and BytePS. The evaluation shows that ASK could accelerate pure key-value aggregation tasks by up to 155 times and big data jobs by 3-5 times, and be backward compatible with existing INA-empowered distributed training solutions with the same speedup.
Yongchao He, Wenfei Wu, Yanfang Le, Ming Liu 0027, ChonLam Lao
ASPLOS (2)3
2023 Preemptive Switch Memory Usage to Accelerate Training Jobs with Shared In-Network Aggregation
abstract
Recent works introduce In-Network Aggregation (INA) for distributed training (DT), which moves the gradient summation into network programmable switches. INA can reduce the traffic volume and accelerate communication in DT jobs. However, switch memory is a scarce resource, unable to support massive DT jobs in data centers, and existing INA solutions have not utilized switch memory to the best extent. We propose DSA, an Efficient Data-Plane switch memory Scheduler for in-network Aggregation. DSA introduces preemption to the switch memory management for INA jobs. In the data plane, DSA allows gradient tensors with high priority to preempt the switch aggregators (basic computation unit in INA) from tensors with low priority, which avoids an aggregator wasting time in idle. In the control plane, DSA devises a priority policy which assigns high priority to gradient tensors that benefit overall job efficiency more, e.g., communication-intensive jobs. We prototype DSA and experiments show that DSA can improve the average JCT by up to 1.35x compared with baseline solutions.
Yuxuan Qin, ChonLam Lao, Yanfang Le, Wenfei Wu
ICNP4
2023 Flor: An Open High Performance RDMA Framework Over Heterogeneous RNICs
Qiang Li 0045, Yixiao Gao, Xiaoliang Wang 0001, Haonan Qiu, Yanfang Le, Derui Liu, Qiao Xiang, Bo Li 0061, Jianbo Dong, Lingbo Tang, Hongqiang Harry Liu, Shaozong Liu, Rui Miao 0001, Yaohui Wu, Zhiwu Wu, Zheng Cao 0003, Zhongjie Wu, Chen Tian 0001, Guihai Chen, Dennis Cai, Jiaji Zhu, Jiesheng Wu, Jiwu Shu
OSDI5
2021 ATP: In-network Aggregation for Multi-tenant Learning
ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Wenfei Wu, Aditya Akella, Michael M. Swift
NSDI2
2019 On the Impact of Cluster Configuration on RoCE Application Design
abstract
RDMA over Converged Ethernet (RoCE) allows RDMA-enabled NICs to operate in datacenter networks. This study focuses on identifying how different aspects of datacenter cluster configuration impact the latency, and throughput, and CPU utilization of different ways of transferring data in RoCE (RDMA verbs). We look into the impact of colocated applications competing for both the CPU and access to the NIC as well as the impact of the network MTU. We find that RDMA applications do not fairly share the NIC, large frames should not be used, and that correct verb choice is dependent on many variables, including application access patterns, object size, and the load of both the local and remote CPU.
Yanfang Le, Mojtaba MalekpourShahraki, Brent E. Stephens, Aditya Akella, Michael M. Swift
APNet1
2018 RoGUE: RDMA over Generic Unconverged Ethernet
abstract
RDMA over Converged Ethernet (RoCE) promises low latency and low CPU utilization over commodity networks, and is attractive for cloud infrastructure services. Current implementations require Priority Flow Control (PFC) that uses backpressure-based congestion control to provide lossless networking to RDMA. Unfortunately, PFC compromises network stability. As a result, RoCE's adoption has been slow and requires complex network management. Recent efforts, such as DCQCN, reduce the risk to the network, but do not completely solve the problem.
Yanfang Le, Brent E. Stephens, Arjun Singhvi, Aditya Akella, Michael M. Swift
SoCC1
2017 UNO: uniflying host and smart NIC offload for flexible packet processing
abstract
Increasingly, smart Network Interface Cards (sNICs) are being used in data centers to offload networking functions (NFs) from host processors thereby making these processors available for tenant applications. Modern sNICs have fully programmable, energy-efficient multi-core processors on which many packet processing functions, including a full-blown programmable switch, can run. However, having multiple switch instances deployed across the host hypervisor and the attached sNICs makes controlling them difficult and data plane operations more complex.
Yanfang Le, Hyunseok Chang, Sarit Mukherjee, Limin Wang 0010, Aditya Akella, Michael M. Swift, T. V. Lakshman
SoCC1
2015 On Datacenter-Network-Aware Load Balancing in MapReduce
abstract
MapReduce has emerged as a powerful tool for distributed and scalable processing of voluminous data. For skewed data input, load balancing is necessary among the MapReduce worker nodes to minimize the overall finishing time, which however can incur massive data movement in a data center network. In this paper, we for the first time examine this problem of data center-network-aware load balancing in the shuffle sub phase in MapReduce. Different from earlier studies that generally assume the network inside a data center has negligible delay and infinite capacity, we consider the traffic and bottlenecks in real data center networks by introducing the constraints on available network bandwidth, and demonstrate that the corresponding problem can be decomposed into two sub problems for network flow and load balancing, respectively. We show effective solutions to both of them, which together yield a complete solution towards near optimal data center-network-aware load balancing. A much simpler yet performance-wise comparable greedy algorithm is also developed for fast implementation in practice. The effectiveness of our solution has been demonstrated on synthetic and real public datasets.
Yanfang Le, Feng Wang 0001, Jiangchuan Liu, Funda Ergün
CLOUD1
2014 Online load balancing for MapReduce with skewed data input
abstract
MapReduce has emerged as a powerful tool for distributed and scalable processing of voluminous data. In this paper, we, for the first time, examine the problem of accommodating data skew in MapReduce with online operations. Different from earlier heuristics in the very late reduce stage or after seeing all the data, we address the skew from the beginning of data input, and make no assumption about a priori knowledge of the data distribution nor require synchronized operations. We examine the input in a continuous fashion and adaptively assign tasks with a load-balanced strategy. We show that the optimal strategy is a constrained version of online minimum makespan and, in the MapReduce context where pairs with identical keys must be scheduled to the same machine, there is an online algorithm with a provable 2-competitive ratio. We further suggest a sample-based enhancement, which, probabilistically, achieves a 3/2-competitive ratio with a bounded error.
Yanfang Le, Jiangchuan Liu, Funda Ergün, Dan Wang 0002
INFOCOM1