EDBT 2026 Demo / reviewers in the wild / expert
Fan Yang 0096
dblp:29/3081-96
· DBLP profile ↗
13ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0004-8087-5903ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 8 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoCoTree: A Computation-Capable Architecture for Collective Communication in Scalable PIMabstractThe growing demand for high-bandwidth and largecapacity memory access in data-intensive workloads has driven the development and deployment of Processing-in-Memory (PIM) architectures. However, existing DIMM-based PIM systems suffer from the severe communication bottleneck between the processing elements (PEs) near the PIM banks due to their requirement on host CPU forwarding. This bottleneck limits the efficiency of collective operations and degrades scalability and performance for workloads that require inter-PE communication. To address the communication limitation, we propose CoCoTree, a computation-capable architecture for collective communication in scalable DIMM-based PIM. CoCoTree supports direct and high-throughput inter-PE communication without host intervention. CoCoTree accelerates key collective communication using novel hierarchical binary tree topology and lightweight in-network computation support. We design and implement microarchitectures for the main building blocks: Co-Leaf and Co-Node, to efficiently handle the data packing, routing, and processing in CoCoTree. Furthermore, we also introduce a packet-based communication protocol tailored to the CoCoTree architecture, which decouples control and data through a twophase configuration-computation communication mechanism to efficiently support a wide range of collective communication operations. CoCoTree effectively mitigates inter-PE communication bottlenecks, enabling scalable PIM systems capable of meeting the demands of growing data size. Experimental results show that CoCoTree achieves up to$95.6 \times$improvement for collective operations and improves end-to-end application performance by up to$10.5 \times$across various workloads over the baseline PIM, while outperforming state-of-the-art PIM communication architectures in both performance and scalability. Shunchen Shi, Qijia Yang, Fan Yang 0096, Yu Huang 0013, Youwei Zhuo, Zhichun Li, Ninghui Sun, Xueqi Li 0001 |
HPCA | 3 |
| 2026 | HOPESim: A Lightweight and Modern C++ based Accelerator Simulation Approach
Xueqi Li 0001, Ruihao Gao, Xiaoyu Zhang 0009, Xiaoming Chen 0003, Shunchen Shi, Fan Yang 0096, Ninghui Sun |
ISCAS | 6 |
| 2026 | Credit-Guided Congestion Control on Wafer-Scale On-Chip Networks for Molecular DynamicsabstractMolecular dynamics (MD) is a cornerstone of scientific computing, but strong scaling often collapses at high parallelism because communication is bursty and highly sensitive to tail latency. MD advances by repeating a fixed timestep loop (one iteration of force computation and state update), and performance is largely determined by how quickly timesteps complete. A key reason is that each timestep contains short, synchronized communication phases, followed by a global dependency before the next timestep. Wafer-scale chips (WSCs) offer cycle-level latency and high on-chip bandwidth, yet their 2D mesh fabrics can still suffer burst-induced queue buildup; existing wavelet scheduling relies on a static stride that either over-injects (triggering credit backpressure) or over-throttles (wasting bandwidth) as conditions evolve. Shixiong Qi, Zhan Wang 0003, Ning Kang 0007, Fan Yang 0096, Yuanzhe Wang, Guanglei Chen, Guangming Tan, Guojun Yuan |
SIGCOMM | 5 |
| 2025 | upTSA: A DIMM-Based Near Data Processing Accelerator for Time Series Analysis
Shunchen Shi, Fan Yang 0096, Qijia Yang, Xiaohui Peng 0002, Xueqi Li 0001, Ninghui Sun |
NPC (1) | 2 |
| 2024 | Palos: Fair and Flexible Flow Scheduling on RNICabstractIn recent years, Remote Direct Memory Access (RDMA) has gained significant attraction within modern hyperscale data centers. However, RNIC fails to provide fine-grained performance isolation among network flows with different traffic patterns which co-exist in multi-tenant data centers and typically have various bandwidth, throughput and latency requirements.In this paper, we reveal that the drawbacks on isolation root in the packet-level flow scheduling mechanism implemented in the RNIC hardware. To solve this problem, we introduce Palos, a fair and flexible flow-scheduling mechanism. In the hardware layer, Palos adopts a data chunk based scheduling mechanism by reconstructing communication descriptors. The data chunk based scheduling diminishes the performance interference between large flows and small flows. Palos configures the scheduler in the software layer using a hierarchical weight setting to enable customized performance policy while preventing the configuration of users from interfering each other. Our experiments demonstrate that Palos provides better performance isolation and performance control flexibility compared with the commodity RDMA NIC and existing optimization framework. Zhenlong Ma, Fan Yang 0096, Ning Kang 0007, Guojun Yuan, Zhan Wang 0003, Ninghui Sun |
HPCC | 2 |
| 2024 | FNCC: Fast Notification Congestion Control in Data Center NetworksabstractCongestion control plays a pivotal role in large-scale data centers, facilitating ultra-low latency, high bandwidth, and optimal utilization. Even with the deployment of data center congestion control mechanisms such as DCQCN and HPCC, these algorithms often respond to congestion sluggishly. This sluggishness is primarily due to the slow notification of congestion. It takes almost one round-trip time (RTT) for the congestion information to reach the sender. In this paper, we introduce the Fast Notification Congestion Control (FNCC) mechanism, which achieves sub-RTT notification. FNCC leverages the acknowledgment packet (ACK) from the return path to carry in-network telemetry (INT) information of the request path, offering the sender more timely and accurate INT. To further accelerate the responsiveness of last-hop congestion control, we propose that the receiver notifies the sender of the number of concurrent congested flows, which can be used to adjust the congested flows to a fair rate quickly. Our experimental results demonstrate that FNCC reduces flow completion time by 27.4% and 88.9% compared to HPCC and DCQCN, respectively. Moreover, FNCC triggers minimal pause frames and maintains high utilization even at 400Gbps. Zhan Wang 0003, Fan Yang 0096, Ning Kang 0007, Zhenlong Ma, Guojun Yuan, Guangming Tan, Ninghui Sun |
ICPP | 3 |
| 2024 | Towards connection-scalable RNIC architecture
Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan |
J. Supercomput. | 3 |
| 2023 | Understanding the Scalability Problem of RNIC Cache at the Micro-architecture LevelabstractHigh-performance Remote Direct Memory Access (RDMA) has been applied in High-Performance Computing (HPC) clusters and large-scale data centers as InfiniBand(IB) or RDMA over Converged Ethernet (RoCE). As the number of connections and the data space accessed by the network increases, the scalability problem of the RDMA Network Interface Card (RNIC) is getting worse because of cache misses on the RNIC. However, the RNIC is a black box for users. To understand the scalability problem, we give a detailed analysis of RNIC communication behavior based on the commercial RNIC driver and PCIe logic analyzer. We design a set of Cache Tests (CacheT) and a universal test methodology to profile the micro-architecture of the cache on the RNIC. We get statistical data by software tests and PCIe logic analyzer and obtain the cache micro-architecture parameters of commercial RNIC. Finally, we present an accurate performance model with 98% accuracy. Xiaoxiao Ma 0004, Fan Yang 0096, Zhan Wang 0003, Ning Kang 0007, Xunjun An |
ICC | 2 |
| 2023 | A Scalable RDMA Network Interface Card with Efficient Cache ManagementabstractRemote Direct Memory Access (RDMA) has been applied in large-scale clusters due to its high bandwidth and low latency features in recent decades. However, the scalability issue is a known intractable problem that the current RDMA Network Interface Card (RNIC) cannot overcome as the number of connections increases to thousands. The key to the scalability bottleneck is managing the connection information cached on the RNIC. To solve the scalability problem, firstly, we test and analyze the behavior of the commercial RNIC when the network performance declines in the case of large-scale connections. Further, we give the performance model of RNIC and point out that the cache design is the key to scalability issues. Then, we propose a Scalable RDMA NIC (ScalaRNIC) architecture with a non-blocking and priority-programmable cache design. Besides, our cache design supports parametric configuration by the extended API. ScalaRNIC can maintain$\mathrm{a}\approx 100\%$cache hit ratio for higher priority connections and keep message rates nearly equal to peak performance. In contrast, commercial RNIC's performance drops by 47% when there are a few thousand connections. Xiaoxiao Ma 0004, Fan Yang 0096, Zhan Wang 0003, Ning Kang 0007, Guojun Yuan, Xunjun An |
ISCAS | 2 |
| 2022 | csRNA: Connection-Scalable RDMA NIC Architecture in Datacenter EnvironmentabstractRDMA has been widely deployed in datacenter networking as an ideal optimization strategy in recent years. Due to its mechanisms such as kernel bypass and hardware offloading, RDMA is expected to offer better performance than traditional kernel-based TCP/IP networking. However, the hardware offloading in RDMA requires the RDMA Network Interface Card (RNIC) to manage the connection metadata, and the limited on-chip memory size in RNIC leads to its limited connection scalability. When the RNIC maintains a large number of connections, its performance drops dramatically.This paper first finds that the head-of-line blocking in connection metadata management is a major factor affecting RNIC scalability. Based on the findings, we propose csRNA, a connection-scalable RNIC architecture that maintains near-peak performance when connection scales. To achieve the non-blocking RNIC processing path, csRNA utilizes a non-blocking connection scheduler to schedule different connections when blocking. Furthermore, using a non-blocking connection management model, csRNA departs from the conventional RNIC design by returning the prepared connections first. csRNA effectively avoids the performance degradation caused by the head-of-line blocking of connection metadata management when the number of connections increases. We implement and evaluate csRNA and demonstrate that with less on-chip memory occupancy, csRNA could still maintain near-peak performance when scaling up to more than 15,000 connections. Ning Kang 0007, Zhan Wang 0003, Fan Yang 0096, Xiaoxiao Ma 0004, Zhenlong Ma, Guojun Yuan, Guangming Tan |
ICCD | 3 |
| 2019 | SwitchAgg: A Further Step Towards In-Network ComputationabstractMany distributed applications adopt a partition/aggregation pattern to achieve high performance and scalability. The aggregation process, which usually takes a large portion of the overall execution time, incurs large amount of network traffic and bottlenecks the system performance. To reduce network traffic, some researches take advantage of network devices to commit innetwork aggregation. However, these approaches use either special topology or middle-boxes, which cannot be easily deployed in current datacenters. The emerging programmable RMT switch brings us new opportunities to implement in-network computation task. However, we argue that the architecture of RMT switch is not suitable for in-network aggregation since it is designed primarily for implementing traditional network functions. In this paper, we first give a detailed analysis of in-network aggregation, and point out the key factor that affects the data reduction ratio. We then propose SwitchAgg, which is an innetwork aggregation system that is compatible with current datacenter infrastructures. We also evaluate the performance improvement we have gained from SwitchAgg. Our results show that, SwitchAgg can process data aggregation tasks at line rate and gives a high data reduction rate, which helps us to cut down network traffic and alleviate pressure on server CPU. In the system performance test, the job-completion-time can be reduced as much as 50%. Fan Yang 0096, Zhan Wang 0003, Xiaoxiao Ma 0004, Guojun Yuan, Xuejun An |
FPGA | 1 |
| 2017 | Regional Congestion Control in Datacenter NetworksabstractThe rapid deployment of cloud computing and online services poses great challenges for data center networks, and congestion control is one of the top concerns. Although numbers of proposals in different network layers have been put forward to alleviate the negative impact of congestion, the short-lived flows, which are latency-sensitive and constitute the majority of total traffic in data centers, still suffer severe performance degradation. Since the existing congestion control methods all rely on end hosts to perceive congestion and then adjust their network sending rate, the response time is relatively long when compared with the duration of short-lived flows, which increases latency significantly. In this paper, we propose RCC, a regional congestion control mechanism, which aims to respond to congestion more quickly and eliminate the mismatch mentioned above. Different from host-based mechanisms, RCC is implemented in the switch, which detects the congestion state and schedule the traffic around the congestion point locally, without sending feedback to the distal host. Evaluation has shown that, compared with the host-based mechanism, our method achieves better performance for short-lived flows and maintains stable buffer occupancy of the switch. In addition, mixed long- and short-lived flows which contend for the same bottleneck link can share the bandwidth more fairly. Fan Yang 0096, Zhan Wang 0003, Xiaoli Liu 0002, Zheng Cao 0003, Guojun Yuan, Xuejun An |
ICPADS | 1 |
| 2017 | Regional Congestion Mitigation in Lossless Datacenter Networks
Xiaoli Liu 0002, Fan Yang 0096, Yanan Jin, Zhan Wang 0003, Zheng Cao 0003, Ninghui Sun |
NPC | 2 |