EDBT 2026 Demo / reviewers in the wild / expert
Ning Kang 0007
dblp:47/2109-7
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0003-7383-9018ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Credit-Guided Congestion Control on Wafer-Scale On-Chip Networks for Molecular DynamicsabstractMolecular dynamics (MD) is a cornerstone of scientific computing, but strong scaling often collapses at high parallelism because communication is bursty and highly sensitive to tail latency. MD advances by repeating a fixed timestep loop (one iteration of force computation and state update), and performance is largely determined by how quickly timesteps complete. A key reason is that each timestep contains short, synchronized communication phases, followed by a global dependency before the next timestep. Wafer-scale chips (WSCs) offer cycle-level latency and high on-chip bandwidth, yet their 2D mesh fabrics can still suffer burst-induced queue buildup; existing wavelet scheduling relies on a static stride that either over-injects (triggering credit backpressure) or over-throttles (wasting bandwidth) as conditions evolve. Shixiong Qi, Zhan Wang 0003, Ning Kang 0007, Fan Yang 0096, Yuanzhe Wang, Guanglei Chen, Guangming Tan, Guojun Yuan |
SIGCOMM | 4 |
| 2025 | MD-pipe: A Strong Scaling Enhanced Pipeline Architecture for Ab Initio Accuracy Molecular DynamicsabstractMolecular Dynamics (MD) simulations with first-principles accuracy are widely applied in various fields, including materials science and molecular pharmacology.Current research focus on reducing the solution time of ab initio molecular dynamics (AIMD) from both Ning Kang 0007, Guojun Yuan, Beining Zhang, Guanglei Chen, Jiayi Rao, Zhan Wang 0003, Weile Jia, Ninghui Sun, Guangming Tan |
ISCA | 1 |
| 2024 | Palos: Fair and Flexible Flow Scheduling on RNICabstractIn recent years, Remote Direct Memory Access (RDMA) has gained significant attraction within modern hyperscale data centers. However, RNIC fails to provide fine-grained performance isolation among network flows with different traffic patterns which co-exist in multi-tenant data centers and typically have various bandwidth, throughput and latency requirements.In this paper, we reveal that the drawbacks on isolation root in the packet-level flow scheduling mechanism implemented in the RNIC hardware. To solve this problem, we introduce Palos, a fair and flexible flow-scheduling mechanism. In the hardware layer, Palos adopts a data chunk based scheduling mechanism by reconstructing communication descriptors. The data chunk based scheduling diminishes the performance interference between large flows and small flows. Palos configures the scheduler in the software layer using a hierarchical weight setting to enable customized performance policy while preventing the configuration of users from interfering each other. Our experiments demonstrate that Palos provides better performance isolation and performance control flexibility compared with the commodity RDMA NIC and existing optimization framework. Zhenlong Ma, Fan Yang 0096, Ning Kang 0007, Guojun Yuan, Zhan Wang 0003, Ninghui Sun |
HPCC | 3 |
| 2024 | FNCC: Fast Notification Congestion Control in Data Center NetworksabstractCongestion control plays a pivotal role in large-scale data centers, facilitating ultra-low latency, high bandwidth, and optimal utilization. Even with the deployment of data center congestion control mechanisms such as DCQCN and HPCC, these algorithms often respond to congestion sluggishly. This sluggishness is primarily due to the slow notification of congestion. It takes almost one round-trip time (RTT) for the congestion information to reach the sender. In this paper, we introduce the Fast Notification Congestion Control (FNCC) mechanism, which achieves sub-RTT notification. FNCC leverages the acknowledgment packet (ACK) from the return path to carry in-network telemetry (INT) information of the request path, offering the sender more timely and accurate INT. To further accelerate the responsiveness of last-hop congestion control, we propose that the receiver notifies the sender of the number of concurrent congested flows, which can be used to adjust the congested flows to a fair rate quickly. Our experimental results demonstrate that FNCC reduces flow completion time by 27.4% and 88.9% compared to HPCC and DCQCN, respectively. Moreover, FNCC triggers minimal pause frames and maintains high utilization even at 400Gbps. Zhan Wang 0003, Fan Yang 0096, Ning Kang 0007, Zhenlong Ma, Guojun Yuan, Guangming Tan, Ninghui Sun |
ICPP | 4 |
| 2024 | Towards connection-scalable RNIC architecture
Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan |
J. Supercomput. | 1 |
| 2023 | Understanding the Scalability Problem of RNIC Cache at the Micro-architecture LevelabstractHigh-performance Remote Direct Memory Access (RDMA) has been applied in High-Performance Computing (HPC) clusters and large-scale data centers as InfiniBand(IB) or RDMA over Converged Ethernet (RoCE). As the number of connections and the data space accessed by the network increases, the scalability problem of the RDMA Network Interface Card (RNIC) is getting worse because of cache misses on the RNIC. However, the RNIC is a black box for users. To understand the scalability problem, we give a detailed analysis of RNIC communication behavior based on the commercial RNIC driver and PCIe logic analyzer. We design a set of Cache Tests (CacheT) and a universal test methodology to profile the micro-architecture of the cache on the RNIC. We get statistical data by software tests and PCIe logic analyzer and obtain the cache micro-architecture parameters of commercial RNIC. Finally, we present an accurate performance model with 98% accuracy. Xiaoxiao Ma 0004, Fan Yang 0096, Zhan Wang 0003, Ning Kang 0007, Xunjun An |
ICC | 4 |
| 2023 | A Scalable RDMA Network Interface Card with Efficient Cache ManagementabstractRemote Direct Memory Access (RDMA) has been applied in large-scale clusters due to its high bandwidth and low latency features in recent decades. However, the scalability issue is a known intractable problem that the current RDMA Network Interface Card (RNIC) cannot overcome as the number of connections increases to thousands. The key to the scalability bottleneck is managing the connection information cached on the RNIC. To solve the scalability problem, firstly, we test and analyze the behavior of the commercial RNIC when the network performance declines in the case of large-scale connections. Further, we give the performance model of RNIC and point out that the cache design is the key to scalability issues. Then, we propose a Scalable RDMA NIC (ScalaRNIC) architecture with a non-blocking and priority-programmable cache design. Besides, our cache design supports parametric configuration by the extended API. ScalaRNIC can maintain$\mathrm{a}\approx 100\%$cache hit ratio for higher priority connections and keep message rates nearly equal to peak performance. In contrast, commercial RNIC's performance drops by 47% when there are a few thousand connections. Xiaoxiao Ma 0004, Fan Yang 0096, Zhan Wang 0003, Ning Kang 0007, Guojun Yuan, Xunjun An |
ISCAS | 4 |
| 2022 | csRNA: Connection-Scalable RDMA NIC Architecture in Datacenter EnvironmentabstractRDMA has been widely deployed in datacenter networking as an ideal optimization strategy in recent years. Due to its mechanisms such as kernel bypass and hardware offloading, RDMA is expected to offer better performance than traditional kernel-based TCP/IP networking. However, the hardware offloading in RDMA requires the RDMA Network Interface Card (RNIC) to manage the connection metadata, and the limited on-chip memory size in RNIC leads to its limited connection scalability. When the RNIC maintains a large number of connections, its performance drops dramatically.This paper first finds that the head-of-line blocking in connection metadata management is a major factor affecting RNIC scalability. Based on the findings, we propose csRNA, a connection-scalable RNIC architecture that maintains near-peak performance when connection scales. To achieve the non-blocking RNIC processing path, csRNA utilizes a non-blocking connection scheduler to schedule different connections when blocking. Furthermore, using a non-blocking connection management model, csRNA departs from the conventional RNIC design by returning the prepared connections first. csRNA effectively avoids the performance degradation caused by the head-of-line blocking of connection metadata management when the number of connections increases. We implement and evaluate csRNA and demonstrate that with less on-chip memory occupancy, csRNA could still maintain near-peak performance when scaling up to more than 15,000 connections. Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan |
ICCD | 1 |