Shunsuke Tatsumi

dblp:191/3590 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 56% GPUs and heterogeneous computing · 44%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing › heterogeneous programming models
OpenCL
0.312017
OpenCL-Based FPGA-Platform for Stencil Computation and Its Optimization Methodology · IEEE Trans. Parallel Distributed Syst. 2017
High-performance computing
stencil computation
0.312017
OpenCL-Based FPGA-Platform for Stencil Computation and Its Optimization Methodology · IEEE Trans. Parallel Distributed Syst. 2017
High-performance computing
scientific computing
0.112017
OpenCL-Based FPGA-Platform for Stencil Computation and Its Optimization Methodology · IEEE Trans. Parallel Distributed Syst. 2017
YearPublicationVenuePosition
2017 OpenCL-Based FPGA-Platform for Stencil Computation and Its Optimization Methodology
abstract
Stencil computation is widely used in scientific computations and many accelerators based on multicore CPUs and GPUs have been proposed. Stencil computation has a small operational intensity so that a large external memory bandwidth is usually required for high performance. FPGAs have the potential to solve this problem by utilizing large internal memory efficiently. However, a very large design, testing and debugging time is required to implement an FPGA architecture successfully. To solve this problem, we propose an FPGA-platform using C-like programming language called open computing language (OpenCL). We also propose an optimization methodology to find the optimal architecture for a given application using the proposed FPFA-platform. According to the experimental results, we achieved 119 - 237 Gflop/s of processing power and higher processing speed compared to conventional GPU and multicore CPU implementations.
Hasitha Muthumala Waidyasooriya, Yasuhiro Takei, Shunsuke Tatsumi, Masanori Hariyama
IEEE Trans. Parallel Distributed Syst.3