EDBT 2026 Demo / reviewers in the wild / expert
Tun Chen
dblp:178/6501
· DBLP profile ↗
8ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-3459-7960ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Computation-communication overlapping based on vertical layer grouping in the inverse legendre transform stage of YHGSM
Yuntian Zheng, Tun Chen, Zhaokai Song, Fukang Yin |
CCF Trans. High Perform. Comput. | 3 |
| 2026 | Automatic generation of cross-platform vectorization kernels for cloud microphysics parameterization
Tun Chen, Fukang Yin, Xiaoli Ren |
J. Parallel Distributed Comput. | 1 |
| 2025 | Auto-CLOUDSC: An Auto-generation Framework for Vectorization and Optimization of Cloud Microphysics Parameterization on ARM CPUs
Tun Chen, Yuntian Zheng, Fukang Yin, Jinhui Yang, Juan Zhao 0006, Xiaoli Ren |
ICA3PP (2) | 1 |
| 2023 | OpenFFT: An Adaptive Tuning Framework for 3D FFT on ARM Multicore CPUsabstractThe sophisticated hierarchy and shared characteristics of cache in multicore CPU architectures bring challenges to the performance improvement of fundamental algorithms, especially in implementing and optimizing 3D FFT. 3D FFT is a memory-bounded algorithm that contains many highly discretized memory accesses. With the working set scaling, the data locality becomes poor, which is prone to cause serious memory access overhead, especially for high-dimensional data transposition. This paper proposes a 3D FFT optimization framework named OpenFFT. This framework optimizes the memory access of 3D FFT by the following methods, including 1) A novel tiling algorithm, Z-OpenFFT, based on the column-order algorithm for high-dimensional vectorization to improve data locality and eliminate transposition; 2) An efficient search algorithm Section-cache-aware algorithm to optimize the memory access of butterfly network of 1D FFT; 3) A multi-thread allocation model by analyzing the characteristics of cache hierarchy and task size to allocate threads adaptively. Experiments demonstrate that OpenFFT could obtain a more competitive performance than the best configuration of FFTW and ARMPL on ARM CPUs. Tun Chen, Haipeng Jia, Yunquan Zhang, Kun Li 0016, Zhihao Li 0001, Jianyu Yao, Chendi Li |
ICS | 1 |
| 2023 | Generating Fast FFT Kernels on CPUs via FFT-Specific IntrinsicsabstractThis paper proposes an algorithm-specific instruction (ASI)-based fast Fourier transform (FFT) code generation framework, named FFTASI, to generate unified architecture independent butterfly kernels that can be transformed into architecture-dependent kernels by establishing the mapping between ASIs and architecture-specific instructions for various hardware platforms. FFTASI strikes a good balance between performance and productivity on CPUs. Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Yuyan Sun, Yiwei Zhang 0009, Tun Chen |
PPoPP | 6 |
| 2020 | Automatic Generation of High-Performance FFT Kernels on Arm and X86 CPUsabstractThis article presents AutoFFT, a template-based code generation framework that can automatically generate high-performance FFT kernels for all natural-number radices. AutoFFT is based on the Cooley-Tukey FFT algorithm, which exploits the symmetric and periodic properties of the DFT matrix, as the outer parallelization framework. Because butterflies are the core operations of the Cooley-Tukey algorithm, we explore additional symmetric and periodic properties of the DFT matrix and formulate multiple optimized calculation templates to further reduce the number of floating-point operations for butterflies of arbitrary natural numbers. To fully exploit hardware resources, we encapsulate a series of optimizations in an assembly template optimizer. Given any DFT problem, AutoFFT automatically generates C FFT kernels using these calculation templates and converts them into efficient assembly kernels using the template optimizer. Through a series of experiments on Arm, Intel, and AMD processors, we show that AutoFFT-generated kernels can outperform those in Fastest Fourier Transform in the West (FFTW), the Arm Performance Libraries (ARMPL), and the Intel Math Kernel Library (MKL). Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Tun Chen, Richard W. Vuduc |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | AutoFFT: a template-based FFT codes auto-generation framework for ARM and X86 CPUsabstractThe discrete Fourier transform (DFT) is widely used in scientific and engineering computation. This paper proposes a template-based code generation framework named AutoFFT that can automatically generate high-performance fast Fourier transform (FFT) codes. AutoFFT employs the Cooley-Tukey FFT algorithm, which exploits the symmetric and periodic properties of the DFT matrix as the outer parallelization framework. To further reduce the number of floating-point operations of butterflies, we explore more symmetric and periodic properties of the DFT matrix and formulate two optimized calculation templates for prime and power-of-two radices. To fully exploit hardware resources, we encapsulate a series of optimizations in an assembly template optimizer. Given any DFT problem, AutoFFT automatically generates C FFT kernels using these two templates and transfers them to efficient assembly codes using the template optimizer. Experiments show that AutoFFT outperforms FFTW, ARMPL, and Intel MKL on average across all FFT types on ARMv8 and Intel x86-64 processors. Zhihao Li 0001, Haipeng Jia, Yunquan Zhang, Tun Chen, Luning Cao |
SC | 4 |
| 2016 | Node localization algorithm for wireless sensor networks using compressive sensing theory
Yehua Wei, Wenjia Li, Tun Chen |
Pers. Ubiquitous Comput. | 3 |