Chenfeng Zhao

dblp:282/8565 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0001-9952-0628ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 HLPerf: Demystifying the Performance of HLS-based Graph Neural Networks with Dataflow Architectures
abstract
The development of FPGA-based applications using HLS is fraught with performance pitfalls and large design space exploration times. These issues are exacerbated when the application is complicated and its performance is dependent on the input dataset, as is often the case with graph neural network approaches to machine learning. Here, we introduce HLPerf, an open-source, simulation-based performance evaluation framework for dataflow architectures that both supports early exploration of the design space and shortens the performance evaluation cycle. We apply the methodology to GNNHLS, an HLS-based graph neural network benchmark containing six commonly used graph neural network models and four datasets with distinct topologies and scales. The results show that HLPerf achieves over 10, 000× average simulation acceleration relative to RTL simulation and over 400× acceleration relative to state-of-the-art cycle-accurate tools at the cost of 7% mean error rate relative to actual FPGA implementation performance. This acceleration positions HLPerf as a viable component in the design cycle.
Chenfeng Zhao, Clayton J. Faber, Roger D. Chamberlain, Xuan Zhang 0001
ACM Trans. Reconfigurable Technol. Syst.1
2024 HLS Taking Flight: Toward Using High-Level Synthesis Techniques in a Space-Borne Instrument
abstract
FPGAs are widely deployed on high-energy astrophysics telescopes to preprocess and reduce sensor data read out by front-end electronics. Across instruments, these computational pipelines have similar semantics, sharing common stages such as pedestal subtraction, signal integration, zero-suppression, island detection, and centroiding. However, diverse telescope designs require unique implementations of these algorithms, and the logic is often rewritten from scratch for a new instrument.
Marion Sudvarg, Chenfeng Zhao, Ye Htet, Meagan Konst, Thomas Lang, Nick Song, Roger D. Chamberlain, Jeremy Buhler, James H. Buckley
CF2
2023 SuperCut: Communication-Aware Partitioning for Near-Memory Graph Processing
abstract
The parallel execution of many graph algorithms is frequently dominated by data communication overheads between compute nodes. This bottleneck becomes even more pronounced in Near-Memory Processing (NMP) architectures with multiple memory cubes as local memory accesses are less expensive. Existing near-memory architectures typically use graph partitioning methods with a fixed vertex assignment, which limits their potential to improve performance and reduce energy consumption. Here, we argue that an NMP-based graph processing system should also consider the distribution of vertices onto memory cubes. We propose SuperCut, a framework for near-memory architectures to effectively reduce communication overheads while maintaining computational balance. We evaluate SuperCut via architectural simulation with 6 real-world datasets and 4 representative applications. The results show that it provides up to 1.8x total energy reduction and 2.6x speedup relative to current state-of-the-art approaches.
Chenfeng Zhao, Roger D. Chamberlain, Xuan Zhang 0001
CF1
2023 GNNHLS: Evaluating Graph Neural Network Inference via High-Level Synthesis
abstract
We present GNNHLS, an open-source framework to comprehensively evaluate GNN inference acceleration on FPGAs via HLS, containing a software stack for data generation and baseline deployment and FPGA implementations of 6 well-tuned GNN HLS kernels. Evaluating on 4 graph datasets with distinct topologies and scales, the results show that GNNHLS achieves up to 50.8× speedup and 423× energy reduction relative to the CPU baselines. Compared with the GPU baselines, GNNHLS achieves up to 5.16× speedup and 74.5× energy reduction.
Chenfeng Zhao, Zehao Dong, Yixin Chen 0001, Xuan Zhang 0001, Roger D. Chamberlain
ICCD1