Chengyue Wang 0002

dblp:277/5345-2 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0006-7481-018XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 HP-FFT: A General High-Performance FFT Generator Using High-Level Synthesis
abstract
The Fast Fourier Transform (FFT) is a foundational algorithm widely used in fields like digital signal processing and machine learning. While High-Level Synthesis (HLS) tools have boosted the development of customized hardware accelerators in these fields, existing FFT HLS IP libraries often suffer from low throughput and poor usability due to inadequate exploitation of potential parallelism. In constract, many high-throughput RTL FFT designs lack portability and flexibility, limiting their practical adoption. To get rid of this predicament, we conducted an in-depth analysis of the FFT algorithm's loop structure, uncovering hierarchical parallelism to optimize performance. Based on these insights, we developed a general FFT HLS generator, HP-FFT, which supports multiple functionalities and a wide range of customizable parallelism settings to meet diverse user requirements. Experimental results demonstrate that our proposed HLS generator matches or outperforms state-of-the-art RTL/HLS IP libraries or generators, while enabling users to easily generate architectures that balance resource efficiency and high throughput to suit various application needs.
Chengyue Wang 0002, Yingquan Wu, Jason Cong
FCCM1
2025 Reconfigurable Stream Network Architecture
abstract
As AI systems grow increasingly specialized and complex, managing hardware heterogeneity becomes a pressing challenge.How can we efficiently coordinate and synchronize heterogeneous hardware resources to achieve high utilization?How can we minimize the friction of transitioning between diverse computation phases, reducing costly stalls from initialization, pipeline setup, or drain?Our insight is that a network abstraction at the ISA level naturally unifies heterogeneous resource orchestration and phase transitions.This paper presents a Reconfigurable Stream Network Architecture (RSN), a novel ISA abstraction designed for the DNN domain.RSN models the datapath as a circuit-switched network with stateful functional units as nodes and data streaming on the edges.Programming a computation corresponds to triggering a path.Software is explicitly exposed to the compute and communication latency of each functional unit, enabling precise control over data movement for optimizations such as compute-communication overlap and layer fusion.As nodes in a network naturally differ, the RSN abstraction can efficiently virtualize heterogeneous hardware resources by separating control from the data plane, enabling low instruction-level intervention.We build a proof-of-concept design RSN-XNN on VCK190, a heterogeneous platform with FPGA fabric and AI engines.Compared to the SOTA solution on this platform, it reduces latency by 6.1x and improves throughput by 2.4x-3.2x.Compared to the T4 GPU with the same FP32 performance, it matches latency with only 18% of the memory bandwidth.Compared to the A100 GPU at the same 7nm process node, it achieves 2.1x higher energy efficiency in FP32.
Chengyue Wang 0002, Xiaofan Zhang 0001, Jason Cong, James C. Hoe
ISCA1
2021 Extending HLS with High-Level Descriptive Language for Configurable Algorithm-Level Spatial Structure Design
abstract
High-level synthesis (HLS) tools have greatly improved the development efficiency of FPGA accelerators in many application areas. With the HLS tools, FPGA designers can focus more on algorithm specifications using software languages such as C/C++, OpenCL, and Python. However, due to the fact that CPU-oriented software languages are designed to describe sequential execution, the repurposing of these languages yields insufficient support for describing parallel data execution and flexible spatial structures on FPGA architecture. To strengthen HLS's ability to describe configurable algorithmlevel spatial structures, we propose fusing hardware-friendly design patterns, namely high-level descriptive language, into imperative programming model on Python.
Chengyue Wang 0002, Sitao Huang, Wen-Mei W. Hwu, Deming Chen
FCCM1
2021 PyLog: An Algorithm-Centric Python-Based FPGA Programming and Synthesis Flow
abstract
The exploding complexity and computation efficiency requirements of applications are stimulating a strong demand for hardware acceleration with heterogeneous platforms such as FPGAs. However, a high-quality FPGA design is very hard to create as it requires FPGA expertise and a long design iteration time. In contrast, software applications are typically developed in a short development cycle, in high-level languages like Python, which is at a much higher level of abstraction than all existing hardware design flows. To close this gap between hardware design flows and software applications, and simplify FPGA programming, we create PyLog, a high-level, algorithm-centric Python-based programming and synthesis flow for FPGA. PyLog is powered by a set of compiler optimization passes and a type inference system to generate high-quality design. It abstracts away the implementation details and allows designers to focus on algorithm specification. PyLog captures more high-level computation patterns for better optimization than traditional HLS systems. PyLog also has a runtime for running PyLog code directly on FPGA platform without any extra code development. Evaluation shows that PyLog significantly improves FPGA design productivity and generates highly efficient FPGA designs that outperform highly optimized CPU and FPGA version by 3.17× and 1.24× on average.
Sitao Huang, Kun Wu 0002, Hyunmin Jeong, Chengyue Wang 0002, Deming Chen, Wen-Mei W. Hwu
FPGA4
2021 PyLog: An Algorithm-Centric Python-Based FPGA Programming and Synthesis Flow
abstract
The exploding complexity and computation efficiency requirements of applications are stimulating a strong demand for hardware acceleration with heterogeneous platforms such as FPGAs. However, a high-quality FPGA design is very hard to create as it requires FPGA expertise and a long design iteration time. In contrast, software applications are typically developed in a short development cycle, with high-level languages like Python, which is at a much higher level of abstraction than all existing hardware design flows. To close this gap and simplify FPGA programming, we create PyLog, a high-level, algorithm-centric programming and synthesis flow for FPGA. PyLog features a set of compiler optimization passes and a type inference system to generate high-quality design. It abstracts away the implementation details, and allows designers to focus on algorithm specification. PyLog takes in Python functions and generates complete optimized FPGA system design. PyLog also has a runtime that allows users to run the PyLog code directly on the target FPGA platform without any extra code development. The whole design flow is automated. The evaluation shows that PyLog significantly improves FPGA design productivity and generates highly efficient FPGA designs that outperform highly optimized CPU and FPGA versions by 3.17x and 1.24x on average.
Sitao Huang, Kun Wu 0002, Hyunmin Jeong, Chengyue Wang 0002, Deming Chen, Wen-Mei W. Hwu
IEEE Trans. Computers4