EDBT 2026 Demo / reviewers in the wild / expert
Qianfeng Shen
dblp:238/9832 · also Qianfeng Clark Shen
· DBLP profile ↗
6ranked-venue papers
2as first author
3since 2021 · last 2022
0000-0001-7525-0248ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Parallel CRC On An FPGA At Terabit SpeedsabstractThe Cyclic Redundancy Check Algorithm (CRC) is critical for ensuring high data reliability in serial communication such as Ethernet networks, allowing for the detection of corrupted packets with a programmable and arbitrarily small probability of failure. The baseline algorithm, however, is highly serialized due to read after write (RAW) dependencies, preventing efficient parallelization of the algorithm for use in hardware. We built a fully parameterizable open-source IP core that has no such dependencies to produce the equivalent result as the baseline CRC algorithm but in a form that can be fully parallelized, with fully automated pipelining, which works for any CRC polynomial, and with a low-resource end-of-packet alignment. This allows for up to 64-bit CRC to be computed in an FPGA at 4 Tbps. Qianfeng Shen, Juan Camilo Vega, Paul Chow |
FPT | 1 |
| 2022 | AIgean: An Open Framework for Deploying Machine Learning on Heterogeneous ClustersabstractAIgean , pronounced like the sea, is an open framework to build and deploy machine learning (ML) algorithms on a heterogeneous cluster of devices (CPUs and FPGAs). We leverage two open source projects: Galapagos , for multi-FPGA deployment, and hls4ml , for generating ML kernels synthesizable using Vivado HLS. AIgean provides a full end-to-end multi-FPGA/CPU implementation of a neural network. The user supplies a high-level neural network description, and our tool flow is responsible for the synthesizing of the individual layers, partitioning layers across different nodes, as well as the bridging and routing required for these layers to communicate. If the user is an expert in a particular domain and would like to tinker with the implementation details of the neural network, we define a flexible implementation stack for ML that includes the layers of Algorithms, Cluster Deployment & Communication, and Hardware. This allows the user to modify specific layers of abstraction without having to worry about components outside of their area of expertise, highlighting the modularity of AIgean . We demonstrate the effectiveness of AIgean with two use cases: an autoencoder, and ResNet-50 running across 10 and 12 FPGAs. AIgean leverages the FPGA’s strength in low-latency computing, as our implementations target batch-1 implementations. Naif Tarafdar, Giuseppe Di Guglielmo, Philip C. Harris, Jeffrey D. Krupa, Vladimir Loncar, Dylan S. Rankin, Zhenbin Wu, Qianfeng Shen, Paul Chow |
ACM Trans. Reconfigurable Technol. Syst. | 9 |
| 2021 | RIFL: A Reliable Link Layer Network Protocol for FPGA-to-FPGA CommunicationabstractMore and more latency-sensitive applications are being introduced into the data center. Performance of such applications can be limited by the high latency of the network interconnect. Because the conventional network stack is designed not only for LAN, but also for WAN, it carries a great amount of redundancy that is not required in a data center network. This paper introduces the concept of a three-layer protocol stack that can replace the conventional network stack and fulfill the exact demands of data center network communications. The detailed design and implementation of the first layer of the stack, which we call RIFL, is presented. A novel low latency in-band hop-by-hop re-transmission protocol is proposed and adopted in RIFL, which guarantees lossless transmission for links whose longest wire segment is no more than 150 meters. Experimental results show that RIFL achieves 218 nanoseconds round-trip latency on 3 meter zero-hop links, at a throughput of 104.7 Gbps. RIFL is a multi-lane protocol with scalable throughput from 500 Mbps to above 200 Gbps. It is portable to most of the recent FPGAs. It can be the enabler of low latency, high throughput, flexible, scalable, and lossless data center networks. Qianfeng Shen, Paul Chow |
FPGA | 1 |
| 2020 | AIgean: An Open Framework for Machine Learning on Heterogeneous ClustersabstractMachine learning (ML) in the past decade has been one of the most popular topics of research within the computing community. Interest within the computing field ranges across all levels of the computation stack. We show this stack in Figure 1. This work introduces an open framework, called AIgean, to build and deploy machine learning (ML) algorithms on a heterogeneous cluster of devices (CPUs and FPGAs). Users can flexibly modify any layer of the machine learning stack in Figure 1 to suit their need. This allows both machine learning domain experts to focus on higher algorithmic layers, and distributed systems experts to create the communication layers below. Naif Tarafdar, Giuseppe Di Guglielmo, Philip C. Harris, Jeffrey D. Krupa, Vladimir Loncar, Dylan S. Rankin, Zhenbin Wu, Qianfeng Shen, Paul Chow |
FCCM | 9 |
| 2020 | SHIP: Storage for Hybrid Interconnected ProcessorsabstractDrivers for accessing storage are complex. In addition to the complexity involved in using the NVMe protocol, navigating filesystems requires multiple serialized storage accesses and data processing between each access. As a result, efforts to create storage drivers for FPGAs, so that FPGAs can directly access storage without help from a CPU, have either failed, require too many resources/time, or remove some of the functionality expected by storage users (such as removing the filesystem) [1], [2]. Juan Camilo Vega, Qianfeng Shen, Paul Chow |
FCCM | 2 |
| 2019 | Introducing ReCPRI: A Field Re-configurable Protocol for Backhaul Communication in a Radio Access Network
Juan Camilo Vega, Qianfeng Shen, Alberto Leon-Garcia, Paul Chow |
IM | 2 |