Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Chia-Lin Yu

dblp:215/5836 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
0since 2021 · last 2020
0000-0003-4530-3665ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 56% High-performance computing · 44%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › performance optimization
auto-tuning
0.412020
Efficient and Portable Workgroup Size Tuning · IEEE Trans. Parallel Distributed Syst. 2020
GPUs and heterogeneous computing › heterogeneous programming models
OpenCL
0.412020
Efficient and Portable Workgroup Size Tuning · IEEE Trans. Parallel Distributed Syst. 2020

Methods — techniques the papers use, named apart from their topics

hardware parameter abstraction · 0.4design space filtering · 0.4
YearPublicationVenuePosition
2020 Efficient and Portable Workgroup Size Tuning
abstract
The performance of an OpenCL program is strongly influenced by both hardware and software attributes. To achieve superior performance, developers may leverage automatic performance tuning techniques to determine the optimal parameters on the target device. Although existing approaches have shown promising tuning results in their target scenarios, other requirements such as efficiency, portability, and usability should also be considered because of the rapid growth of heterogeneous computing applications and platforms. In this paper, we re-examine the workgroup size tuning problem and propose a novel approach to meet the aforementioned requirements. We abstract the architectural details into a set of hardware parameters so that the proposed approach can be applied without the presence of target devices, which makes it more accessible to developers. The proposed approach is evaluated on 20 OpenCL kernels and six devices, including both CPUs and GPUs. Experimental results demonstrate that, with negligible overhead, our approach filters out 88.6 percent of the possible workgroup sizes on average. Among all the workgroup size candidates, the bestand worst-performing candidates can achieve average performance of 95.5 and 92.1 percent, respectively, compared with the optimal workgroup size.
Chia-Lin Yu, Shiao-Li Tsao
IEEE Trans. Parallel Distributed Syst.1