EDBT 2026 Demo / reviewers in the wild / expert
Yiwei Zhang 0009
dblp:86/1695-9
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-2050-572XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core UnitsabstractWhile Tensor Core Units (TCUs) excel in AI tasks, their application to HPC algorithms like stencil computations faces significant challenges due to sparsity, which leads to underutilization and exacerbates memory-bound limitations. This paper introduces FlashFFTStencil1, a memory-efficient stencil computing system designed to bridge FFT to fully-dense stencil computations on TCUs. Aimed at bound shifting, FlashFFTStencil comprises three key techniques: Kernel Tailoring on HBM fuses distinct kernels to enhance parallelism while reducing memory transfer and footprint; Architecture Aligning on SMEM restructures FFT-based stencil computations into dense matrix multiplications tailored for shared memory architecture; Computation Streamlining on TCU optimizes TCU utilization and thread parallelism by minimizing pipeline stalls and maximizing register reuse. Notably, a distinctive extension is FlashFFTStencil's ability to enable theoretically unrestricted temporal fusion by FFT. Results show that FlashFFTStencil achieves effective sparsity-free bound shifting, with an average speedup of 2.57x over the state-of-the-art. FlashFFTStencil pioneers a new era in unifying computational patterns within the HPC landscape and bridges them with cutting-edge AI-driven hardware innovations like TCUs. Haozhi Han, Kun Li 0016, Donglin Bai, Yiwei Zhang 0009, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
PPoPP | 5 |
| 2025 | Jigsaw: Toward Conflict-free Vectorized Stencil Computation by Tessellating Swizzled RegistersabstractStencil computation plays a pivotal role in numerous scientific and engineering applications. Previous studies have extensively investigated vectorization techniques to enhance in-core parallelism; however, the performance bottleneck caused by data alignment conflicts (DAC) has not been effectively resolved in all dimensions. This paper proposes Jigsaw, a conflict-free vectorization method to reduce DAC across all dimensions by tessellating swizzled finest-grained lanes. Jigsaw comprises three key components: Lane-based Butterfly Vectorization, SVD-based Dimension Flattening, and Iteration-based Temporal Merging. These components effectively address DAC across spatial and temporal dimensions. Experimental results on different machines demonstrate that Jigsaw could achieve a significant improvement compared to the state-of-the-art techniques, with an average speedup of 2.31x on various stencil kernels. Yiwei Zhang 0009, Kun Li 0016, Haozhi Han, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
PPoPP | 1 |
| 2024 | LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor CoresabstractStencil computations play a pivotal role in numerous scientific and industrial applications, yet their efficient execution on specialized hardware accelerators like Tensor Core Units (TCUs) remains a challenge. This paper introduces LoRAStencil1, a novel stencil computing system designed to mitigate memory access redundancies on TCUs through low-rank adaptation. We first identify a nuanced form of this redundancy, dimension residue, specific to TCUs. Then LoRAStencil leverages orchestrated mathematical transformations to decompose stencil weight matrices into smaller rank-1 matrices, facilitating efficient data gathering along residual dimensions. It comprises three key components: memory-efficient Residual Dimension Gathering to facilitate more data reuse, compute-saving Pyramidal Matrix Adaptation to exploit the inherent low-rank characteristics, and performance-boosting Butterfly Vector Swapping to circumvent all data shuffles. Comprehensive evaluations demonstrate that LoRAStencil address dimension residues effectively, which outperforms state-of-the-arts with up to a 2.16x speedup, offering promising advancements for efficient tensorized stencil computation on TCUs by Low-Rank Adaptation. Yiwei Zhang 0009, Kun Li 0016, Jiawen Cheng, Yunquan Zhang, Ting Cao 0003, Mao Yang 0004 |
SC | 1 |
| 2023 | Generating Fast FFT Kernels on CPUs via FFT-Specific IntrinsicsabstractThis paper proposes an algorithm-specific instruction (ASI)-based fast Fourier transform (FFT) code generation framework, named FFTASI, to generate unified architecture independent butterfly kernels that can be transformed into architecture-dependent kernels by establishing the mapping between ASIs and architecture-specific instructions for various hardware platforms. FFTASI strikes a good balance between performance and productivity on CPUs. Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Yuyan Sun, Yiwei Zhang 0009, Tun Chen |
PPoPP | 5 |
| 2021 | Decision Tree Based Inter Partition Termination For Av1 EncodingabstractAs a next-generation video coding standard, AV1 introduces numerous new coding tools, leading to high computational complexity and high time cost. To deal with this problem, in this paper, we propose a decision tree based algorithm to early terminate the inter prediction process by predicting splitting decisions at each depth. Motion compensated block is introduced to provide temporal neighborhood information. Nine attributes are selected and analyzed in this paper, and a set of decision trees are generated for different block sizes. According to experimental results, our algorithm can save 23.6% of encoding time on average, with a negligible BD-rate loss of 0.73% under low-delay encoding mode. Yiwei Zhang 0009, Yanghao Li, Jiangtao Wen |
ICASSP | 2 |