Qingsong Ning

dblp:36/11358 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0001-4951-3099ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Reconfigurable computing and FPGAs · 32% Hardware accelerators and domain-specific architectures · 26% Cloud and datacenter computing · 20%
Computer networks
2 papers
Transport protocols and congestion control · 79% Datacenter networks · 21%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs › cloud FPGA
cloud FPGA acceleration
0.912025
Harmonia: A Unified Framework for Heterogeneous FPGA Acceleration in the Cloud · ASPLOS (2) 2025
Interconnection networks and networks-on-chip
in-network computing
0.912025
A Generic and Efficient Communication Framework for Message-Level In-Network Computing · INFOCOM 2025
Transport protocols and congestion control › rate control
rate limiting
0.812024
Fast, Scalable, and Accurate Rate Limiter for RDMA NICs · SIGCOMM 2024
Reconfigurable computing and FPGAs
FPGA accelerator
0.612022
FAERY: An FPGA-accelerated Embedding-based Retrieval System · OSDI 2022
Reconfigurable computing and FPGAs › FPGA architecture
heterogeneous FPGA
0.312025
Harmonia: A Unified Framework for Heterogeneous FPGA Acceleration in the Cloud · ASPLOS (2) 2025
Cloud and datacenter computing
traffic isolation
0.212024
Fast, Scalable, and Accurate Rate Limiter for RDMA NICs · SIGCOMM 2024
Datacenter networks
RDMA
0.212023
SRNIC: A Scalable Architecture for RDMA NICs · NSDI 2023
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval
0.212022
FAERY: An FPGA-accelerated Embedding-based Retrieval System · OSDI 2022
Information retrieval
retrieval models
0.212022
FAERY: An FPGA-accelerated Embedding-based Retrieval System · OSDI 2022

Methods — techniques the papers use, named apart from their topics

packet filtering · 1.5hierarchical rate limiting · 1.5adaptive batching · 1.5FPGA · 1.1shell-role architecture · 0.9in-network computing · 0.9
YearPublicationVenuePosition
2025 Harmonia: A Unified Framework for Heterogeneous FPGA Acceleration in the Cloud
abstract
FPGAs are gaining popularity in the cloud as accelerators for various applications. To make FPGAs more accessible for users and streamline system management, cloud providers have widely adopted the shell-role architecture on their homogeneous FPGA servers. However, the increasing heterogeneity of cloud FPGAs poses new challenges for this architecture. Previous studies either focus on homogeneous FPGAs or only partially address the portability issues for roles, while still requiring laborious shell development for providers and ad-hoc software modifications for users.
Xinchen Wan, Zilong Wang 0007, Qian Zhao 0001, Feng Ning, Qingsong Ning, Shideng Zhang, Zhenyu Li 0001, Layong Luo, Gaogang Xie
ASPLOS (2)8
2025 A Generic and Efficient Communication Framework for Message-Level In-Network Computing
Xinchen Wan, Han Tian, Xudong Liao, Chaoliang Zeng, Zilong Wang 0007, Qingsong Ning, Guyue Liu, Layong Luo, Kai Chen 0005
INFOCOM10
2024 Fast, Scalable, and Accurate Rate Limiter for RDMA NICs
abstract
RDMA NICs desire a rate limiter that is accurate, scalable, and fast: to precisely enforce the policies such as congestion control and traffic isolation, to support a large number of flows, and to sustain high packet rates. Prior works such as SENIC and PIEO can achieve accuracy and scalability, but they are not fast enough, thus fail to fulfill the performance requirement of RNICs, due primarily to their monolithic design and one-packet-per-sorting transmission. We present Tassel, a hierarchical rate limiter for RDMA NICs that can deliver high packet rates by enabling multiple-packet-per-sorting transmission, while preserving accuracy and scalability. At its heart, Tassel renovates the workflow of the rate limiter hierarchically: by first applying scalable rate limiting to the flows to be scheduled, followed by accurate rate limiting to the packets to be transmitted, while leveraging adaptive batching and packet filtering to improve the performance of these two steps. We integrate Tassel into the RNIC architecture by replacing the original QP scheduler module and implement the prototype of Tassel using FPGA. Experimental results show that Tassel delivers 125 Mpps packet rate, outperforming SENIC and PIEO by 3.6×, while supporting 16 K flows with low resource usage, 7.5% - 25.6% as compared to SENIC and PIEO, and preserving high accuracy, precisely enforcing rate limits from 100 Kbps to 100 Gbps.
Zilong Wang 0007, Xinchen Wan, Yijun Sun, Qingsong Ning, Junxue Zhang 0001, Kai Chen 0005
SIGCOMM7
2023 SRNIC: A Scalable Architecture for RDMA NICs
Zilong Wang 0007, Layong Luo, Qingsong Ning, Chaoliang Zeng, Wenxue Li 0004, Xinchen Wan, Xiongfei Geng, Tianhao Wang 0025, Weicheng Ling, Kejia Huo, Pingbo An, Kui Ji, Shideng Zhang, Ruiqing Feng, Kai Chen 0005, Chuanxiong Guo
NSDI3
2022 FAERY: An FPGA-accelerated Embedding-based Retrieval System
Chaoliang Zeng, Layong Luo, Qingsong Ning, Yaodong Han, Ding Tang, Zilong Wang 0007, Kai Chen 0005, Chuanxiong Guo
OSDI3
2012 ITester: A FPGA based high performance traffic replay tool
abstract
Packet replay is an important way to reproduce real traffic for network test. Many works focus on the performance and accuracy of packet generation based on hardware, such as NP and FPGA, in order to replace inefficient software based packet replay tools. However, limited onboard storage space constrains the size of trace file. In this paper, we design a novel FPGA based packet replay tool called iTester which supports to replay large trace file without impairing the performance and accuracy. It combines high-performance feature of FPGA with vast storage space in host PC. In RAM disk mode, file of Giga Bytes, depending on the memory size in host PC, can be replayed at 10Gbps while the average error of packet gap is kept within 10ns. Additionally, compared to copy with packet, it has a 69% performance improvement in RAM disk mode by adjusting the block size for memory copy.
Fuxing Zhang, Yingke Xie, Layong Luo, Qingsong Ning
FPL5
2012 Building a Flexible and Scalable Virtual Hardware Data Plane
Yingke Xie, Gaogang Xie, Layong Luo, Fuxing Zhang, Qingsong Ning, Hongtao Guan
Networking (1)7