EDBT 2026 Demo / reviewers in the wild / expert
Pengcheng Xu 0005
dblp:65/4669-5
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-2724-7893ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FPsPIN: An FPGA-Based Open-Hardware Research Platform for Processing in the NetworkabstractNetwork offload offers low-latency communication for use-cases where data-movement needs to be combined with lightweight processing. Processing incoming data on the network hardware instead of the CPU reduces latency (by eliminating the need for the data to be deposited into main memory) and frees CPU cycles. Various methods have been employed to offload protocol- and data-processing onto network interface cards (NICs), from firmware modification to running full Linux on NICs for application execution. The sPIN project enables users to define handlers executed upon packet arrival. While simulations show sPIN's potential across diverse workloads, a full-system evaluation is lacking. This work presents FPsPIN, a full FPGA-based implementation of sPIN. FPsPIN is showcased through of-floaded MPI datatype processing, achieving a 96 % overlap ratio. FPsPIN provides an adaptable open-source research platform for researchers to conduct end-to-end experiments on smart NICs. Timo Schneider, Pengcheng Xu 0005, Torsten Hoefler |
HOTI | 2 |
| 2025 | The NIC should be part of the OSabstractThe network interface adapter (NIC) is a critical component of a cloud server occupying a unique position. Not only is network performance vital to efficient operation of the machine, but unlike compute accelerators like GPUs, the network subsystem must react to unpredictable events like the arrival of a network packet and communicate with the appropriate application end point with minimal latency. Pengcheng Xu 0005, Timothy Roscoe |
HotOS | 1 |
| 2022 | Critique of "MemXCT: Memory-Centric X-Ray CT Reconstruction With Massive Parallelization" by SCC Team From Peking UniversityabstractHidayetoluet al.(2019) proposed a novel memory-centric computation system, MemXCT. As a challenge at SC20, we reproduce the computational efficiency of MemXCT on our Azure cloud cluster. Our experiments evaluate the overall performance and the strong scalability with real datasets and verify part of the conclusions in the original article. Zejia Fan, Zhewen Hao, Yueyang Pan, Pengcheng Xu 0005, Yuxuan Yan, Fangyuan Yang, Zhenxin Fu, Yun Liang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | HASCO: Towards Agile HArdware and Software CO-design for Tensor ComputationabstractTensor computations overwhelm traditional general-purpose computing devices due to the large amounts of data and operations of the computations. They call for a holistic solution composed of both hardware acceleration and software mapping. Hardware/software (HW/SW) co-design optimizes the hardware and software in concert and produces high-quality solutions. There are two main challenges in the co-design flow. First, multiple methods exist to partition tensor computation and have different impacts on performance and energy efficiency. Besides, the hardware part must be implemented by the intrinsic functions of spatial accelerators. It is hard for programmers to identify and analyze the partitioning methods manually. Second, the overall design space composed of HW/SW partitioning, hardware optimization, and software optimization is huge. The design space needs to be efficiently explored. To this end, we propose an agile co-design approach HASCO that provides an efficient HW/SW solution to dense tensor computation. We use tensor syntax trees as the unified IR, based on which we develop a two-step approach to identify partitioning methods. For each method, HASCO explores the hardware and software design spaces. We propose different algorithms for the explorations, as they have distinct objectives and evaluation costs. Concretely, we develop a multi-objective Bayesian optimization algorithm to explore hardware optimization. For software optimization, we use heuristic and Q-learning algorithms. Experiments demonstrate that HASCO achieves a 1.25X to 1.44X latency reduction through HW/SW co-design compared with developing the hardware and software separately. Qingcheng Xiao, Size Zheng 0001, Bingzhe Wu, Pengcheng Xu 0005, Xuehai Qian, Yun Liang 0001 |
ISCA | 4 |
| 2021 | Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Peking UniversityabstractShi et al. (2018) proposed a highly parallel polynomial filtering eigensolver for the computation of planetary normal modes. As a challenge at the Student Cluster Competition in The International Conference for High Performance Computing, Networking, Storage and Analysis (SC19), we reproduce the computational efficiency of the polynomial filtering eigensolver on our Intel Xeon machine. We present the weak scalability, scaling of runtime with model size (in a fixed interval) and the strong scalability results in this report. Yihua Cheng, Zejia Fan, Jing Mai, Yifan Wu 0005, Pengcheng Xu 0005, Yuxuan Yan, Zhenxin Fu, Yun Liang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |