Shanwei Ye

dblp:406/8797 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0006-7437-8737ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Interconnection networks and networks-on-chip · 87% Reconfigurable computing and FPGAs · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Interconnection networks and networks-on-chip
network interface
0.912025
A High-Performance RDMA NIC With Ultrahighly Scalable Connections · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Interconnection networks and networks-on-chip
remote direct memory access
0.912025
A High-Performance RDMA NIC With Ultrahighly Scalable Connections · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Reconfigurable computing and FPGAs › FPGA-based network processing
FPGA-based network interface
0.312025
A High-Performance RDMA NIC With Ultrahighly Scalable Connections · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025

Methods — techniques the papers use, named apart from their topics

multitiered cache structure · 0.9chain prefetching · 0.9
YearPublicationVenuePosition
2025 A High-Performance RDMA NIC With Ultrahighly Scalable Connections
abstract
Remote direct memory access (RDMA) technology has significantly enhanced network bandwidth and decreased transmission latency through kernel bypass and protocol offloading, overcoming obstacles in distributed computing systems. However, with the deployment of more intricate services in RDMA networks, current RDMA network interface cards (RNICs) have experienced a notable performance decline as the number of queue pair (QP) connections increases, substantially constraining the broad acceptance of RDMA networks. To address this challenge, this article proposes a novel RNIC architecture with high connection scalability. This architecture incorporates a multitiered cache structure to handle diverse communication contexts, enabling RNIC to support ultrahigh QP connection numbers while minimizing on-chip memory usage. In addition, the architecture facilitates chain prefetching, allowing on-chip caches to manage multiple concurrent requests; thus, averting latency resulting from cache misses and access conflicts during communication under concurrent multiple QP scenarios. This ensures transmission performance in multi-QPs connection scenarios. This article implements and validates the performance of a 100G RNIC based on this architecture on Xilinx’s U280 FPGA. With approximately 1 M memory usage on-chip for context, it can support 64 K performant QP connections ($25\times $than CX-6) and can be extended if necessary. Experimental results confirm the high connection scalability of the RNIC, achieving approximately 92 Gb/s network throughput for data packet transmission with concurrent execution of 1–64 K QPs.
Zhenlong Wan, Pingjing Liu, Qilin Dai, Shanwei Ye, Yingcheng Lin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8