EDBT 2026 Demo / reviewers in the wild / expert
Xiaoxiao Ma 0004
dblp:32/8037-4
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0002-1973-5248ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards connection-scalable RNIC architecture
Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan |
J. Supercomput. | 4 |
| 2023 | Understanding the Scalability Problem of RNIC Cache at the Micro-architecture LevelabstractHigh-performance Remote Direct Memory Access (RDMA) has been applied in High-Performance Computing (HPC) clusters and large-scale data centers as InfiniBand(IB) or RDMA over Converged Ethernet (RoCE). As the number of connections and the data space accessed by the network increases, the scalability problem of the RDMA Network Interface Card (RNIC) is getting worse because of cache misses on the RNIC. However, the RNIC is a black box for users. To understand the scalability problem, we give a detailed analysis of RNIC communication behavior based on the commercial RNIC driver and PCIe logic analyzer. We design a set of Cache Tests (CacheT) and a universal test methodology to profile the micro-architecture of the cache on the RNIC. We get statistical data by software tests and PCIe logic analyzer and obtain the cache micro-architecture parameters of commercial RNIC. Finally, we present an accurate performance model with 98% accuracy. Xiaoxiao Ma 0004, Fan Yang 0096, Zhan Wang 0003, Ning Kang 0007, Xunjun An |
ICC | 1 |
| 2023 | A Scalable RDMA Network Interface Card with Efficient Cache ManagementabstractRemote Direct Memory Access (RDMA) has been applied in large-scale clusters due to its high bandwidth and low latency features in recent decades. However, the scalability issue is a known intractable problem that the current RDMA Network Interface Card (RNIC) cannot overcome as the number of connections increases to thousands. The key to the scalability bottleneck is managing the connection information cached on the RNIC. To solve the scalability problem, firstly, we test and analyze the behavior of the commercial RNIC when the network performance declines in the case of large-scale connections. Further, we give the performance model of RNIC and point out that the cache design is the key to scalability issues. Then, we propose a Scalable RDMA NIC (ScalaRNIC) architecture with a non-blocking and priority-programmable cache design. Besides, our cache design supports parametric configuration by the extended API. ScalaRNIC can maintain$\mathrm{a}\approx 100\%$cache hit ratio for higher priority connections and keep message rates nearly equal to peak performance. In contrast, commercial RNIC's performance drops by 47% when there are a few thousand connections. Xiaoxiao Ma 0004, Fan Yang 0096, Zhan Wang 0003, Ning Kang 0007, Guojun Yuan, Xunjun An |
ISCAS | 1 |
| 2022 | csRNA: Connection-Scalable RDMA NIC Architecture in Datacenter EnvironmentabstractRDMA has been widely deployed in datacenter networking as an ideal optimization strategy in recent years. Due to its mechanisms such as kernel bypass and hardware offloading, RDMA is expected to offer better performance than traditional kernel-based TCP/IP networking. However, the hardware offloading in RDMA requires the RDMA Network Interface Card (RNIC) to manage the connection metadata, and the limited on-chip memory size in RNIC leads to its limited connection scalability. When the RNIC maintains a large number of connections, its performance drops dramatically.This paper first finds that the head-of-line blocking in connection metadata management is a major factor affecting RNIC scalability. Based on the findings, we propose csRNA, a connection-scalable RNIC architecture that maintains near-peak performance when connection scales. To achieve the non-blocking RNIC processing path, csRNA utilizes a non-blocking connection scheduler to schedule different connections when blocking. Furthermore, using a non-blocking connection management model, csRNA departs from the conventional RNIC design by returning the prepared connections first. csRNA effectively avoids the performance degradation caused by the head-of-line blocking of connection metadata management when the number of connections increases. We implement and evaluate csRNA and demonstrate that with less on-chip memory occupancy, csRNA could still maintain near-peak performance when scaling up to more than 15,000 connections. Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan |
ICCD | 4 |
| 2019 | SwitchAgg: A Further Step Towards In-Network ComputationabstractMany distributed applications adopt a partition/aggregation pattern to achieve high performance and scalability. The aggregation process, which usually takes a large portion of the overall execution time, incurs large amount of network traffic and bottlenecks the system performance. To reduce network traffic, some researches take advantage of network devices to commit innetwork aggregation. However, these approaches use either special topology or middle-boxes, which cannot be easily deployed in current datacenters. The emerging programmable RMT switch brings us new opportunities to implement in-network computation task. However, we argue that the architecture of RMT switch is not suitable for in-network aggregation since it is designed primarily for implementing traditional network functions. In this paper, we first give a detailed analysis of in-network aggregation, and point out the key factor that affects the data reduction ratio. We then propose SwitchAgg, which is an innetwork aggregation system that is compatible with current datacenter infrastructures. We also evaluate the performance improvement we have gained from SwitchAgg. Our results show that, SwitchAgg can process data aggregation tasks at line rate and gives a high data reduction rate, which helps us to cut down network traffic and alleviate pressure on server CPU. In the system performance test, the job-completion-time can be reduced as much as 50%. Fan Yang 0096, Zhan Wang 0003, Xiaoxiao Ma 0004, Guojun Yuan, Xuejun An |
FPGA | 3 |