Wenzhi Yin

dblp:267/1248 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Reconfigurable computing and FPGAs · 100%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.412020
Towards Higher Performance and Robust Compilation for CGRA Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2020
Reconfigurable computing and FPGAs › coarse-grained reconfigurable architecture
loop mapping
0.412020
Towards Higher Performance and Robust Compilation for CGRA Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2020
Reconfigurable computing and FPGAs
modulo scheduling
0.412020
Towards Higher Performance and Robust Compilation for CGRA Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2020
Operating systems › resource management › process management
CPU scheduling
0.112020
Towards Higher Performance and Robust Compilation for CGRA Modulo Scheduling · IEEE Trans. Parallel Distributed Syst. 2020

Methods — techniques the papers use, named apart from their topics

reordering · 0.9buffer allocation · 0.9backtracking · 0.9
YearPublicationVenuePosition
2020 Towards Higher Performance and Robust Compilation for CGRA Modulo Scheduling
abstract
Coarse-Grained Reconfigurable Architectures (CGRA) is a promising solution for accelerating computation intensive tasks due to its good trade-off in energy efficiency and flexibility. One of the challenging research topic is how to effectively deploy loops onto CGRAs within acceptable compilation time. Modulo scheduling (MS) has shown to be efficient on deploying loops onto CGRAs. Existing CGRA MS algorithms still suffer from the challenge of mapping loop with higher performance under acceptable compilation time, especially mapping large and irregular loops onto CGRAs with limited computational and routing resources. This is mainly due to the under utilization of the available buffer resources on CGRA, unawareness of critical mapping constraints and time consuming method of solving temporal and spatial mapping. This article focus on improving the performance and compilation robustness of the modulo scheduling mapping algorithm for CGRAs. We decomposes the CGRA MS problem into the temporal and spatial mapping problem and reorganize the processes inside these two problems. For the temporal mapping problem, we provide a comprehensive and systematic mapping flow that includes a powerful buffer allocation algorithm, and efficient interconnection & computational constraints solving algorithms. For the spatial mapping problem, we develop a fast and stable spatial mapping algorithm with backtracking and reordering mechanism. Our MS mapping algorithm is able to map loops onto CGRA with higher performance and faster compilation time. Experiment results show that given the same compilation time budget, our mapping algorithm generates higher compilation success rate. Among the successfully compiled loops, our approach can improve 5.4 to 14.2 percent performance and takes x24 to x1099 less compilation time in average comparing with state-of-the-art CGRA mapping algorithms.
Zhongyuan Zhao 0004, Weiguang Sheng, Qin Wang 0009, Wenzhi Yin, Pengfei Ye, Jinchao Li, Zhigang Mao
IEEE Trans. Parallel Distributed Syst.4