Yiwei Zhang 0009

dblp:86/1695-9 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-2050-572XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core Units
abstract
While Tensor Core Units (TCUs) excel in AI tasks, their application to HPC algorithms like stencil computations faces significant challenges due to sparsity, which leads to underutilization and exacerbates memory-bound limitations. This paper introduces FlashFFTStencil1, a memory-efficient stencil computing system designed to bridge FFT to fully-dense stencil computations on TCUs. Aimed at bound shifting, FlashFFTStencil comprises three key techniques: Kernel Tailoring on HBM fuses distinct kernels to enhance parallelism while reducing memory transfer and footprint; Architecture Aligning on SMEM restructures FFT-based stencil computations into dense matrix multiplications tailored for shared memory architecture; Computation Streamlining on TCU optimizes TCU utilization and thread parallelism by minimizing pipeline stalls and maximizing register reuse. Notably, a distinctive extension is FlashFFTStencil's ability to enable theoretically unrestricted temporal fusion by FFT. Results show that FlashFFTStencil achieves effective sparsity-free bound shifting, with an average speedup of 2.57x over the state-of-the-art. FlashFFTStencil pioneers a new era in unifying computational patterns within the HPC landscape and bridges them with cutting-edge AI-driven hardware innovations like TCUs.
Haozhi Han, Kun Li 0016, Donglin Bai, Yiwei Zhang 0009, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004
PPoPP5
2025 Jigsaw: Toward Conflict-free Vectorized Stencil Computation by Tessellating Swizzled Registers
abstract
Stencil computation plays a pivotal role in numerous scientific and engineering applications. Previous studies have extensively investigated vectorization techniques to enhance in-core parallelism; however, the performance bottleneck caused by data alignment conflicts (DAC) has not been effectively resolved in all dimensions. This paper proposes Jigsaw, a conflict-free vectorization method to reduce DAC across all dimensions by tessellating swizzled finest-grained lanes. Jigsaw comprises three key components: Lane-based Butterfly Vectorization, SVD-based Dimension Flattening, and Iteration-based Temporal Merging. These components effectively address DAC across spatial and temporal dimensions. Experimental results on different machines demonstrate that Jigsaw could achieve a significant improvement compared to the state-of-the-art techniques, with an average speedup of 2.31x on various stencil kernels.
Yiwei Zhang 0009, Kun Li 0016, Haozhi Han, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004
PPoPP1
2024 LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor Cores
abstract
Stencil computations play a pivotal role in numerous scientific and industrial applications, yet their efficient execution on specialized hardware accelerators like Tensor Core Units (TCUs) remains a challenge. This paper introduces LoRAStencil1, a novel stencil computing system designed to mitigate memory access redundancies on TCUs through low-rank adaptation. We first identify a nuanced form of this redundancy, dimension residue, specific to TCUs. Then LoRAStencil leverages orchestrated mathematical transformations to decompose stencil weight matrices into smaller rank-1 matrices, facilitating efficient data gathering along residual dimensions. It comprises three key components: memory-efficient Residual Dimension Gathering to facilitate more data reuse, compute-saving Pyramidal Matrix Adaptation to exploit the inherent low-rank characteristics, and performance-boosting Butterfly Vector Swapping to circumvent all data shuffles. Comprehensive evaluations demonstrate that LoRAStencil address dimension residues effectively, which outperforms state-of-the-arts with up to a 2.16x speedup, offering promising advancements for efficient tensorized stencil computation on TCUs by Low-Rank Adaptation.
Yiwei Zhang 0009, Kun Li 0016, Jiawen Cheng, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004
SC1
2023 Generating Fast FFT Kernels on CPUs via FFT-Specific Intrinsics
abstract
This paper proposes an algorithm-specific instruction (ASI)-based fast Fourier transform (FFT) code generation framework, named FFTASI, to generate unified architecture independent butterfly kernels that can be transformed into architecture-dependent kernels by establishing the mapping between ASIs and architecture-specific instructions for various hardware platforms. FFTASI strikes a good balance between performance and productivity on CPUs.
Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Yuyan Sun, Yiwei Zhang 0009, Tun Chen
PPoPP5
2021 Decision Tree Based Inter Partition Termination For Av1 Encoding
abstract
As a next-generation video coding standard, AV1 introduces numerous new coding tools, leading to high computational complexity and high time cost. To deal with this problem, in this paper, we propose a decision tree based algorithm to early terminate the inter prediction process by predicting splitting decisions at each depth. Motion compensated block is introduced to provide temporal neighborhood information. Nine attributes are selected and analyzed in this paper, and a set of decision trees are generated for different block sizes. According to experimental results, our algorithm can save 23.6% of encoding time on average, with a negligible BD-rate loss of 0.73% under low-delay encoding mode.
Yiwei Zhang 0009, Yanghao Li, Jiangtao Wen
ICASSP2