EDBT 2026 Demo / reviewers in the wild / expert
Kentaro Sano
dblp:62/699
· DBLP profile ↗
47ranked-venue papers
14as first author
19since 2021 · last 2026
0000-0002-6681-4192ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 12 first-author · 17 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A case study in hardware specialization for Monte Carlo cross-section lookup
Kazutomo Yoshii, John R. Tramm, Bryce Allen, Tomohiro Ueno, Kentaro Sano, Andrew R. Siegel, Pete Beckman |
Parallel Comput. | 5 |
| 2025 | Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing
Franck Cappello, Mario C. Acosta, Emmanuel Agullo, Hartwig Anzt, Jon Calhoun 0001, Sheng Di, Luc Giraud, Thomas Grützmacher, Sian Jin, Kentaro Sano, Kento Sato, Amarjit Singh, Dingwen Tao, Jiannan Tian, Tomohiro Ueno, Robert Underwood, Frédéric Vivien, Xavier Yepes, Kazutomo Yoshii, Boyuan Zhang 0002 |
Future Gener. Comput. Syst. | 10 |
| 2025 | A Scalable Accelerator for Local Score Computation of Structure Learning in Bayesian NetworksabstractA Bayesian network is a powerful tool for representing uncertainty in data, offering transparent and interpretable inference, unlike neural networks’ black-box mechanisms. To fully harness the potential of Bayesian networks, it is essential to learn the graph structure that appropriately represents variable interrelations within data. Score-based structure learning, which involves constructing collections of potentially optimal parent sets for each variable, is computationally intensive, especially when dealing with high-dimensional data in discrete random variables. Our proposed novel acceleration algorithm extracts high levels of parallelism, offering significant advantages even with reduced reusability of computational results. In addition, it employs an elastic data representation tailored for parallel computation, making it FPGA-friendly and optimizing module occupancy while ensuring uniform handling of diverse problem scenarios. Demonstrated on a Xilinx Alveo U50 FPGA, our implementation significantly outperforms optimal CPU algorithms and is several times faster than GPU implementations on an NVIDIA TITAN RTX. Furthermore, the results of performance modeling for the accelerator indicate that, for sufficiently large problem instances, it is weakly scalable, meaning that it effectively utilizes increased computational resources for parallelization. To our knowledge, this is the first study to propose a comprehensive methodology for accelerating score-based structure learning, blending algorithmic and architectural considerations. Ryota Miyagi, Ryota Yasudo, Kentaro Sano, Hideki Takase |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2024 | Flexible Systolic Array Platform on Virtual 2-D Multi-FPGA PlaneabstractSystolic arrays are a promising approach to achieving high-performance processing based on highly parallelized designs in various fields, such as AI and bioinformatics. Many previous studies have devoted considerable effort to exploring efficient circuit designs for specific processing. However, the increasing size of systolic arrays forces us to process increasingly large workloads by dividing them into smaller pieces. Therefore, we propose a systolic array platform based on a two-dimensional FPGA plane in which multiple FPGAs are connected by a virtual network. The systolic array realized by this system can be freely customized in shape and size according to the target. Distributed memory access through off-chip memory on each FPGA board and simple stream processing enable scalable performance. This paper presents a preliminary implementation based on the proposed systolic array platform and its performance evaluation. The evaluation results show that the proposed method improves the processing performance in proportion to the number of FPGAs. The results also show that the proposed platform is highly scalable due to the small circuit area required, and that the processing performance depends on the network bandwidth, which means that recent high-bandwidth FPGA boards can be expected to significantly improve the performance. Tomohiro Ueno, Emanuele Del Sozzo, Kentaro Sano |
HPC Asia | 3 |
| 2024 | Exploration of Trade-offs Between General-Purpose and Specialized Processing Elements in HPC-Oriented CGRAabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable accelerators traditionally used in embedded computing. Recently, CGRA-like devices have gained traction for HPC and AI acceleration; however, typical HPC and AI workloads often require operations that current CGRAs cannot implement, such as complex mathematical calculations. In this work, we present a broad architectural study exploring potential heterogeneous computational resources in CGRA architectures for HPC, which are not commonly considered in typical CGRA architecture research. We first improved the general-purpose Processing Element (PE) of a baseline CGRA to optimize computational resources and then developed a new specialized PE for mathematical functions commonly found in HPC applications. Finally, we evaluated multiple CGRA configurations concerning floorplan, size, general-purpose/specialized PE ratio, and Power, Performance, and Area (PPA) results from hardware synthesis. Emanuele Del Sozzo, Xinyuan Wang 0003, Boma Anantasatya Adhi, Carlos Cortes, Jason Helge Anderson, Kentaro Sano |
IPDPS | 6 |
| 2024 | Automated parallel execution of distributed task graphs with FPGA clustersabstractOver the years, Field Programmable Gate Arrays (FPGA) have been gaining popularity in the High Performance Computing (HPC) field, because their reconfigurability enables very fine-grained optimizations with low energy cost. However, the different characteristics, architectures, and network topologies of the clusters have hindered the use of FPGAs at a large scale. In this work, we present an evolution of OmpSs@FPGA, a high-level task-based programming model and extension to OmpSs-2, that aims at unifying all FPGA clusters by using a message-passing interface that is compatible with FPGA accelerators. These accelerators are programmed with C/C++ pragmas, and synthesized with High-Level Synthesis tools. The new framework includes a custom protocol to exchange messages between FPGAs, agnostic of the architecture and network type. On top of that, we present a new communication paradigm called Implicit Message Passing (IMP), where the user does not need to call any message-passing API. Instead, the runtime automatically infers data movement between nodes. We test classic message passing and IMP with three benchmarks on two different FPGA clusters. One is cloudFPGA, a disaggregated platform with AMD FPGAs that are only connected to the network through UDP/TCP/IP. The other is ESSPER, composed of CPU-attached Intel FPGAs that have a private network at the ethernet level. In both cases, we demonstrate that IMP with OmpSs@FPGA can increase the productivity of FPGA programmers at a large scale thanks to simplifying communication between nodes, without limiting the scalability of applications. We implement the N-body, Heat simulation and Cholesky decomposition benchmarks, and show that FPGA clusters get 2.6x and 2.4x better performance per watt than a CPU-only supercomputer for N-body and Heat. Juan Miguel De Haro Ruiz, Carlos Álvarez 0001, Daniel Jiménez-González, Xavier Martorell, Tomohiro Ueno, Kentaro Sano, Burkhard Ringlein, François Abel, Beat Weiss |
Future Gener. Comput. Syst. | 6 |
| 2024 | Introduction to the Special Issue on FPL 2022abstractThe International Conference on Field-Programmable Logic and Applications (FPL)is widely regarded to be the premier venue for presenting current research on reconfigurable technology in Europe.As the original conference in this domain, FPL continues to be a most comprehensive gathering for experts, researchers, and enthusiasts in this dynamic and evolving area.In 2022, the 32nd event has returned to Belfast, where it already took place in 2001.Despite this long history, FPL's key topics remain as relevant as ever.Field-programmable devices, notably FPGAs (Field-Programmable Gate Arrays), present a unique blend of the benefits of dedicated hardware-such as enhanced performance and power efficiency-combined with a versatility and user-friendly nature akin to software.This duality positions them as an attractive option in scenarios where traditional computing platforms like CPUs or GPUs fall short in performance or flexibility, and where the deployment of fully application-specific integrated circuits (ASICs) is impractical due to their prohibitive nonrecurring costs and the intensive design efforts required for contemporary silicon fabrication technologies.The scope of reconfigurable technology is vast, encompassing a broad spectrum of research areas crucial for its advancement.These areas include the development of innovative tools and design methodologies, the architecture of field-programmable systems, and the exploration of device technology for field-programmable chips.Equally important is the practical application of this technology-understanding how it can be effectively utilized in various domains to translate its potential into tangible benefits for end users.This exploration is fundamental in bridging the gap between theoretical advancements and real-world impacts.Across this wide range of topics, FPL 2022 received an initial 178 abstract submissions and 129 full submissions.After a reviewing process that included a rebuttal phase and at least three reviews for the research papers, 33 contributions could be accepted as full papers and 28 as short papers.In addition, the newly offered FPL journal track, which focused on more mature work profiting from a longer-form presentation, attracted eight submissions.Four of these contributions could be accepted for publication in ACM Transactions on Reconfigurable Technology and Systems (TRETS), and were also presented as regular talks at the event.After the conference, we invited the best 10 papers from the FPL conference to submit extended versions of their work to ACM TRETS for a special issue. Andreas Koch 0001, Kentaro Sano |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2024 | Across Time and Space: Senju's Approach for Scaling Iterative Stencil Loop Accelerators on Single and Multiple FPGAsabstractStencil-based applications play an essential role in high-performance systems as they occur in numerous computational areas, such as partial differential equation solving. In this context, Iterative Stencil Loops (ISLs) represent a prominent and well-known algorithmic class within the stencil domain. Specifically, ISL-based calculations iteratively apply the same stencil to a multi-dimensional point grid multiple times or until convergence. However, due to their iterative and intensive nature, ISLs are highly performance-hungry, demanding specialized solutions. Here, Field Programmable Gate Arrays (FPGAs) represent a valid architectural choice as they enable the design of custom, parallel, and scalable ISL accelerators. Besides, the regular structure of ISLs makes them an ideal candidate for automatic optimization and generation flows. For these reasons, this article introduces Senju , an automation framework for the design of highly parallel ISL accelerators targeting single-/multi-FPGA systems. Given an input description, Senju automates the entire design process and provides accurate performance estimations. The experimental evaluation shows remarkable and scalable results, outperforming single- and multi-FPGA literature approaches under different metrics. Finally, we present a new analysis of temporal and spatial parallelism trade-offs in a real-case scenario and discuss our performance through a single- and novel specialized multi-FPGA formulation of the Roofline Model. Emanuele Del Sozzo, Davide Conficconi, Kentaro Sano |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2023 | Senju: A Framework for the Design of Highly Parallel FPGA-based Iterative Stencil Loop AcceleratorsabstractStencil-based applications play an essential role in high-performance systems as they occur in numerous computational areas, such as partial differential equation solving, seismic simulations, and financial option pricing, to name a few. In this context, Iterative Stencil Loops (ISLs) represent a prominent and well-known algorithmic class within the stencil domain. Specifically, ISL-based calculations iteratively apply the same stencil to a multi-dimensional system of points until it reaches convergence. However, due to their iterative and computationally intensive nature, these workloads are highly performance-hungry, demanding specialized solutions to boost performance and reduce power consumption. Here, FPGAs represent a valid architectural choice as their peculiar features enable the design of custom, parallel, and scalable ISL accelerators. Besides, the regular structure of ISLs makes them an ideal candidate for automatic optimization and generation flows. For these reasons, this paper introduces Senju, an automation framework for FPGA-based ISL accelerators. Starting from an input description, Senju builds highly parallel hardware modules and automatizes all their design phases. The experimental evaluation shows remarkable and scalable results, reaching significant performance and energy efficiency improvements compared to the other single-FPGA literature approaches. Emanuele Del Sozzo, Davide Conficconi, Marco D. Santambrogio, Kentaro Sano |
FPGA | 4 |
| 2023 | ESSPER: Elastic and Scalable FPGA-Cluster System for High-Performance Reconfigurable Computing with Supercomputer FugakuabstractFPGA clusters have yet to be a mainstream of HPC, even for accelerators, and several challenges exist in their architecture and system organization. This work presents ESSPER, a flexible and scalable FPGA cluster prototype system for reconfigurable HPC to meet the concept of customizability, scalability, and interoperability with existing HPC systems. Based on our classification of FPGA cluster architectures, we propose a new category of FPGA clusters with a host-FPGA bridging network using software-bridged APIs for the use of remote FPGAs. We have designed, implemented, verified, and demonstrated a proof-of-concept system of ESSPER, as a functional extension of the supercomputer Fugaku. Kentaro Sano, Atsushi Koshiba, Takaaki Miyajima, Tomohiro Ueno |
HPC Asia | 1 |
| 2023 | Experimental Survey of FPGA-Based Monolithic Switches and a Novel Queue BalancerabstractThis article studies small to medium-sized monolithic switches for FPGA implementation and presents a novel switch design that achieves high algorithmic performance and FPGA implementation efficiency. Crossbar switches based on virtual output queues (VOQs) and variations have been rather popular for implementing switches on FPGAs, with applications in network switches, memory interconnects, network-on-chip (NoC) routers etc. The implementation efficiency of crossbar-based switches is well-documented on ASICs, though we show that their disadvantages can outweigh their advantages on FPGAs. One of the most important challenges in such input-queued switches is the requirement for iterative scheduling algorithms. In contrast to ASICs, this is more harmful on FPGAs, as the reduced operating frequency and narrower packets cannot “hide” multiple iterations of scheduling that are required to achieve a modest scheduling performance. Our proposed design uses an output-queued switch internally for simplifying scheduling, and a queue balancing technique to avoid queue fragmentation and reduce the need for memory-sharing VOQs. Its implementation approaches the scheduling performance of a state-of-the-art FPGA-based switch, while requiring considerably fewer resources. Philippos Papaphilippou, Kentaro Sano, Boma Anantasatya Adhi, Wayne Luk |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | VCSN: Virtual Circuit-Switching Network for Flexible and Simple-to-Operate Communication in HPC FPGA ClusterabstractFPGA clusters promise to play a critical role in high-performance computing (HPC) systems in the near future due to their flexibility and high power efficiency. The operation of large-scale general-purpose FPGA clusters on which multiple users run diverse applications requires flexible network topology to be divided and reconfigured. This paper proposes Virtual Circuit-Switching Network (VCSN) that provides an arbitrarily reconfigurable network topology and simple-to-operate network system among FPGA nodes. With virtualization, user logic on FPGAs can communicate with each other as if a circuit-switching network was available. This paper demonstrates that VCSN with 100 Gbps Ethernet achieves highly-efficient point-to-point communication among FPGAs due to its unique and efficient communication protocol. We compare VCSN with a direct connection network (DCN) that connects FPGAs directly. We also show a concrete procedure to realize collective communication on an FPGA cluster with VCSN. We demonstrate that the flexible virtual topology provided by VCSN can accelerate collective communication with simple operations. Furthermore, based on experimental results, we model and estimate communication performance by DCN and VCSN in a large FPGA cluster. The result shows that VCSN has the potential to accelerate gather communication up to about 1.97 times more than DCN. Tomohiro Ueno, Kentaro Sano |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | The Cost of Flexibility: Embedded versus Discrete Routers in CGRAs for HPCabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a class of reconfigurable architectures that inherit the performance and usability properties of Central Processing Units (CPUs) and the reconfigurability aspects of Field-Programmable Gate Arrays (FPGAs). Historically, CGRAs have been successfully used to accelerate embedded applications and are today also being considered to accelerate High-Performance Computing (HPC) applications in future supercomputers. However, embedded systems and supercomputers are two vastly different domains with different applications and constraints, and it is today not fully understood what CGRA design decisions adequately cater to the HPC market. One such unknown design decision is regarding the interconnect that facilitates intra-CGRA communication. Today, intra-CGRA communication comes in two flavors: using routers closely embedded into the compute units or using discrete routers outside the compute units. The former trades flexibility for a reduction in hardware cost, while the latter has greater flexibility but is more resource hungry. In this paper, we aspire to understand which of both designs best suits the CGRA HPC segment. We extend our previous methodology, which consists of both a parameterized CGRA design and an OpenMPcapable compiler, to accommodate both types of routing designs, including verification tests using RTL simulation. Our results show that the discrete router design can facilitate better use of processing elements (PEs) compared to embedded routers and can achieve up to 79.27% reduction in unnecessary PE occupancy for an aggressively unrolled stencil kernel on a 18 × 16 CGRA at a (estimated) hardware resource overhead cost of 6.3x. This reduction in PE occupancy can be used, for example, to exploit instruction-level parallelism (ILP) through even more aggressive unrolling. Boma Anantasatya Adhi, Carlos Cortes, Yiyu Tan, Takuya Kojima, Artur Podobas, Kentaro Sano |
CLUSTER | 6 |
| 2022 | Exploring Inter-tile Connectivity for HPC-oriented CGRA with Lower Resource UsageabstractThis research aims to explore the tradeoffs between routing flexibility and hardware resource usage, ultimately reducing the resource usage of our CGRA architecture while maintaining compute efficiency. we investigate statistics of connection usages among switch blocks for benchmark DFGs, propose several CGRA architecture with a reduced connection, and evaluate their hardware cost, routability of DFGs, and computational throughput for benchmarks. We found that the topology with horizontal plus diagonal connection saves about 30% of the resource usage while maintaining virtually the same routing flexibility as the full connectivity topology. Boma Anantasatya Adhi, Carlos Cortes, Tomohiro Ueno, Yiyu Tan, Takuya Kojima, Artur Podobas, Kentaro Sano |
FPT | 7 |
| 2022 | Elastic Sample Filter: An FPGA-based Accelerator for Bayesian Network Structure Learningabstractproposed in 1985 by Judea Pearl [1], Ryota Miyagi, Ryota Yasudo, Kentaro Sano, Hideki Takase |
FPT | 3 |
| 2022 | ESSPER: Elastic and Scalable System for High-Performance Reconfigurable Computing with Software-bridged APIsabstractMany-core CPUs and GPUs, present mainstream architectures for HPC, are facing difficulty in maintaining the same performance improvement rate because of the recent slow-down in the semiconductor scaling, the dark silicon problem, and wasteful mechanisms required for accelerating general-purpose computing such as a branch predictor and an out-of-order mechanism. Also, the power efficiency of HPC systems is significantly important to achieve higher performance. Kentaro Sano, Atsushi Koshiba, Takaaki Miyajima, Tomohiro Ueno |
FPT | 1 |
| 2021 | Virtual Circuit-Switching Network with Flexible Topology for High-Performance FPGA ClusterabstractAs the performance of high-end FPGAs has increased in recent years, it’s getting more important to construct an FPGA cluster for both improved processing performance and power efficiency in data centers and supercomputers. For higher utilization of FPGA resources for various applications, we require a flexible inter-FPGA network which provides various topologies appropriately to different applications while a conventional direct-connection network (DCN) provides only a fixed topology, such like a 2D torus. In this paper, we propose a virtual circuit-switching network (VCSN) for a large-scale FPGA cluster to have a flexible inter-FPGA network, where communication links connecting FPGAs are virtualized on the top of Ethernet frames. We can easily configure the VCSN topology optimized for the application by modifying the destination MAC addresses registered in a table of a frame encoder. We present its efficient protocol, hardware implementation, demonstration with 100Gbps Ethernet, and performance comparison with a conventional direct-connection network for FPGAs. We show that VCSN has higher but acceptable latency and slightly higher throughput in comparison with DCN, so that numerical simulation running with a ring of FPGAs achieves comparable performance for DCN. Tomohiro Ueno, Atsushi Koshiba, Kentaro Sano |
ASAP | 3 |
| 2021 | A memory bandwidth improvement with memory space partitioning for single-precision floating-point FFT on Stratix 10 FPGAabstractThe Fast Fourier Transform (FFT) is one of the fundamental computational methods used in the fields of computational science and high-performance computing. Single-precision floating-point complex FFT itself is known as a memory bandwidth bottleneck and often becomes a bottleneck of application acceleration in these fields. We are researching and developing a parallel FFT on FPGA(s) to overcome this problem. In this paper, we discuss the memory bandwidth of the single-precision floating-point complex FFT on an FPGA. Our FFT implementation is based on a state-of-the-art OpenCL implementation provided by Intel. We first show that the computational performance of the FFT on Intel PAC D5005 is proportional to the effective memory bandwidth of the main memory. Then we propose a memory sub-system to improve the effective memory bandwidth. Specifically, a memory space partitioning and the sub-modules that access each memory space individually. In our FPGA design running at 270 MHz, two memory channels of DDR4-2400 memory are used for both reading and writing, respectively. Our proposed memory sub-system achieved an effective memory bandwidth of 22.57 [GB/s] (65.3% of the theoretical peak of this implementation) was achieved when the number of data points for FFT was 16,777,216. Takaaki Miyajima, Kentaro Sano |
CLUSTER | 2 |
| 2021 | Efficient Queue-Balancing Switch for FPGAsabstractThis paper presents a novel FPGA-based switch design that achieves high algorithmic performance and an efficient FPGA implementation. Crossbar switches based on virtual output queues (VOQs) and variations have been rather popular for implementing switches on FPGAs, with applications to network-on-chip (NoC) routers and network switches. The efficiency of VOQs is well-documented on ASICs, though we show that their disadvantages can outweigh their advantages on FPGAs. Our proposed design uses an output-queued switch internally for simplifying scheduling, and a queue balancing technique to avoid queue fragmentation and reduce the need for memory-sharing VOQs. Our implementation approaches the scheduling performance of the state-of-the-art, while requiring considerably fewer FPGA resources. Philippos Papaphilippou, Kentaro Sano, Boma Anantasatya Adhi, Wayne Luk |
FPT | 2 |
| 2020 | A Template-based Framework for Exploring Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are being considered as a complementary addition to modern High-Performance Computing (HPC) systems. These reconfigurable devices overcome many of the limitations of the (more popular) FPGA, by providing higher operating frequency, denser compute capacity, and lower power consumption. Today, CGRAs have been used in several embedded applications, including automobile, telecommunication, and mobile systems, but the literature on CGRAs in HPC is sparse and the field full of research opportunities. In this work, we introduce our CGRA simulator infrastructure for use in evaluating future HPC CGRA systems. Our CGRA simulator is built on synthesizable VHDL and is highly parametrizable, including support for connectivity, SIMD, data-type width, and heterogeneity. Unlike other related work, our framework supports co-integration with third-party memory simulators or evaluation of future memory architecture, which is crucial to reason around memory-bound applications. We demonstrate how our framework can be used to explore the performance of multiple different kernels, showing the impact of different configuration and design-space options. Artur Podobas, Kentaro Sano, Satoshi Matsuoka |
ASAP | 2 |
| 2020 | Extending High-Level Synthesis with High-Performance Computing Performance VisualizationabstractThe recent maturity in High-Level Synthesis (HLS) has renewed the interest of using Field-Programmable Gate-Arrays (FPGAs) to accelerate High-Performance Computing (HPC) applications. Today, several studies have shown performance- and power-benefits of using FPGAs compared to existing approaches for a number of application kernels with ample room for improvements. Unfortunately, modern HLS tools offer little support to gain clarity and insight regarding why a certain application behaves as it does on the FPGA, and most experts rely on intuition or abstract performance models. In this work, we hypothesize that existing profiling and visualization tools used in the HPC domain are also usable for understanding performance on FPGAs. We extend an existing HLS tool-chain to support Paraver - a state-of-the-art visualization and profiling tool well-known in HPC. We describe how each of the events and states are collected, and empirically quantify its hardware overhead. Finally, we practically apply our contribution to two different applications, demonstrating how the tool can be used to provide unique insights into application execution and how it can be used to guide optimizations. In this work, we hypothesize that existing profiling and visualization tools used in the HPC domain are also usable for understanding performance on FPGAs. We extend an existing HLS tool-chain to support Paraver - a state-of-the-art visualization and profiling tool well-known in HPC. We describe how each of the events and states are collected, and empirically quantify its hardware overhead. Finally, we practically apply our contribution to two different applications, demonstrating how the tool can be used to provide unique insights into application execution and how it can be used to guide optimizations. Jens Huthmann, Artur Podobas, Lukas Sommer, Andreas Koch 0001, Kentaro Sano |
CLUSTER | 5 |
| 2020 | Performance Evaluation and Power Analysis of Teraflop-scale Fluid Simulation with Stratix 10 FPGAabstractStream computing is a suitable approach to improve both performance and power efficiency of numerical computations with FPGAs. To achieve further performance gain, temporal and spatial parallelism were exploited: the first one deepens and the latter duplicates pipelines of streamed computation cores. These two types of parallelism were previously evaluated with Arria 10 FPGA. However, it has not been verified if they are also effective for the latest FPGA, Stratix 10, which has a larger amount of logic elements (i.e., 2.4X of Arria 10) and is equipped with a new feature to improve the maximum clock frequency (i.e., HyperFlex architecture). To show the scalability for such state-of-the-art FPGAs, in this paper, we firstly implemented a streamed fluid simulation accelerator with both parallelism types for Stratix 10. We then thoroughly evaluated it by obtaining computational performance (FLOPS), power efficiency (FLOPS/W), resource utilization, and maximum clock frequency (Fmax). From the results, we found that this implementation excessively used DSP blocks due to inefficient mapping of floating-point operations, which reduced Fmax and the number of pipelined cores. To improve the scalability, we optimized the implementation to reduce the DSP block usage by utilizing a Multiply-Add function in a single DSP block. As a result, the optimized fluid simulation achieves 1.06 TFLOPS and 12.6 GFLOPS/W, which is 1.36X and 1.24X higher than the non-optimized version, respectively. Moreover, we estimate that the fluid simulation with Stratix 10 could outperform GPU-based implementation with Tesla V100 by optimizing it for HyperFlex architecture. Atsushi Koshiba, Kouki Watanabe, Takaaki Miyajima, Kentaro Sano |
FPGA | 4 |
| 2019 | FPGA implementation of a robot control algorithmabstractRobot control operation is always implemented onto a processor-based embedded system. However, if a robot operation was implemented onto an FPGA as hardware operation, the robot operation could be accelerated. Since currently high-level synthesis tool is available, a C++ designed operation can easily be translated onto the corresponding FPGA hardware. In this paper, we present the hardware implementation of control algorithm of a robot with 6 legs on an Arria10 SOC field programmable gate array using Cyber Work Bench as a high-level synthesis tool. Yusuke Takaki, Kohei Nagasu, Shin Abiko, Minoru Watanabe, Kentaro Sano |
ETFA | 5 |
| 2018 | Performance Estimation of Deeply Pipelined Fluid Simulation on Multiple FPGAs with High-speed Communication SubsystemabstractTo precisely evaluate the sustained performance and scalability of pipelined multiple FPGAs, this paper presents the implementation of a high-speed communication subsystem for a deeply pipelined stream computing platform. Internally, a pipeline of hardware modules for a domain-specific application is implemented in the FPGAs, where they are directly connected through their serial transceiver links. The necessary inter-FPGA communication subsystem with a flow control mechanism is implemented with Intel Arria 10 FPGAs, where the resource consumption and sustained network throughput are obtained, which averages at 7.92 GB/s. Performance estimation of a fluid simulation using the measured inter-FPGA network parameters is shown for a varied pipeline depth with explored temporal and spatial parallel options. Results show that the proposed platform with 16 FPGAs is estimated to achieve a sustained performance of 4.8 TFlops and may be scaled further by deepening the pipeline with even up to 128 FPGAs. Antoniette Mondigo, Kentaro Sano, Hiroyuki Takizawa |
ASAP | 2 |
| 2018 | Performance Analysis of Hardware-Based Numerical Data Compression on Various Data FormatsabstractThe amount of data processed in high-performance computing has been growing rapidly. Accordingly, the cost of data movement in a large-scale computing system further increases, which has a huge effect on computing performance. To reduce the data movement cost, we have proposed hardware-based data compression for numerical data streams that can greatly reduce overhead has been proposed. Although our proposed hardware compressor works well for scientific computation in previous studies, the compression performance heavily depends on a type of target data. For practical use, we need to know the characteristics of the data compression and the relationship between data types and compression performance. In this paper, we clarify difference in compression performance among different data types by investigating hardware-based compression performance for data types of double, single, and half-precision floating-point and fixed-point with results of numerical simulation. We also propose an area-saving hardware design for double-precision floating-point data. Tomohiro Ueno, Kentaro Sano, Takashi Furusawa |
DCC | 2 |
| 2018 | Enhancing Memory Bandwidth in a Single Stream Computation with Multiple FPGAsabstractStream computing is an area where FPGAs can be suitably utilized to meet high performance and high scalability demands. To achieve these, a deep computing pipeline is implemented on multiple FPGAs where stream computing is performed. This paper presents an approach to utilize two masters in a 1D ring network of multiple FPGAs for a single stream computation. Each master FPGA will be reading and writing to their respective DDR3 memories alternately, while streaming through the slave FPGAs. This is done in order to synchronize the computational results on physically separate memory units. Due to this, the aggregate memory bandwidth is doubled, which suggests enhanced performance. The introduction of this streaming concept lays the groundwork towards full utilization of memories in all the FPGAs, as an identified future work. Antoniette Mondigo, Kentaro Sano, Hiroyuki Takizawa |
FPT | 2 |
| 2017 | FPGA-based tsunami simulation: Performance comparison with GPUs, and roofline model for scalability analysis
Kohei Nagasu, Kentaro Sano, Fumiya Kono, Naohito Nakasato |
J. Parallel Distributed Comput. | 2 |
| 2017 | FPGA-Based Scalable and Power-Efficient Fluid Simulation using Floating-Point DSP BlocksabstractHigh-performance and low-power computation is required for large-scale fluid dynamics simulation. Due to the inefficient architecture and structure of CPUs and GPUs, they now have a difficulty in improving power efficiency for the target application. Although FPGAs become promising alternatives for power-efficient and high-performance computation due to their new architecture having floating-point (FP) DSP blocks, their relatively narrow memory bandwidth requires an appropriate way to fully exploit the advantage. This paper presents an architecture and design for scalable fluid simulation based on data-flow computing with a state-of-the-art FPGA. To exploit available hardware resources including FP DSPs, we introduce spatial and temporal parallelism to further scale the performance by adding more stream processing elements (SPEs) in an array. Performance modeling and prototype implementation allow us to explore the design space for both the existing Altera Arria10 and the upcoming Intel Stratix10 FPGAs. We demonstrate that Arria10 10AX115 FPGA achieves 519 GFlops at 9.67 GFlops/W only with a stream bandwidth of 9.0 GB/s, which is 97.9 percent of the peak performance of 18 implemented SPEs. We also estimate that Stratix10 FPGA can scale up to 6844 GFlops by combining spatial and temporal parallelism adequately. Kentaro Sano, Satoru Yamamoto |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Bandwidth Compression of Floating-Point Numerical Data Streams for FPGA-Based High-Performance ComputingabstractAlthough computational performance is often limited by insufficient bandwidth to/from an external memory, it is not easy to physically increase off-chip memory bandwidth. In this study, we propose a hardware-based bandwidth compression technique that can be applied to field-programmable gate array-- (FPGA) based high-performance computation with a logically wider effective memory bandwidth. Our proposed hardware approach can boost the performance of FPGA-based stream computations by applying a data compression technique to effectively transfer more data streams. To apply this data compression technique to bandwidth compression via hardware, several requirements must first be satisfied, including an acceptable level of compression performance and a sufficiently small hardware footprint. Our proposed hardware-based bandwidth compressor utilizes an efficient prediction-based data compression algorithm. Moreover, we propose a multichannel serializer and deserializer that enable applications to use multiple channels of computational data with the bandwidth compression. The serializer encodes compressed data blocks of multiple channels into a data stream, which is efficiently written to an external memory. Based on preliminary evaluation, we define an encoding format considering both high compression ratio and small hardware area. As a result, we demonstrate that our area saving bandwidth compressor increases performance of an FPGA-based fluid dynamics simulation by deploying more processing elements to exploit spatial parallelism with the enhanced memory bandwidth. Tomohiro Ueno, Kentaro Sano, Satoru Yamamoto |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2016 | Parallelism for High-Performance Tsunami Simulation with FPGA: Spatial or Temporal?abstractTo carry out fast but accurate tsunami simulation after a major earthquake, we have developed an FPGA-based custom computing machine for high-speed but low-power tsunami simulator. We design a stream processing element (SPE) which is hardware based on pipelining and data-flow for tsunami computation. This paper presents design-space exploration for spatial and temporal parallelism of SPEs. Kohei Nagasu, Kentaro Sano, Fumiya Kono, Naohito Nakasato, Alexander Vazhenin, Stanislav G. Sedukhin |
FCCM | 2 |
| 2014 | Bandwidth compression of multiple numerical data streams for high performance custom computingabstractBandwidth compression improves the performance of stream computing by enhancing an effective bandwidth. To apply the bandwidth compression to numerical applications such as numerical simulations, a compressor has to handle multiple data streams. In this paper, we describe a design of an FPGA-based bandwidth compressor for high performance stream computation. For synchronization of original data in multiple compressed streams with different bit-rate, we propose a data block transmission scheduler and explore a design space to reduce the size of their barrel shifters. Tomohiro Ueno, Ryo Ito, Kentaro Sano, Satoru Yamamoto |
ASAP | 3 |
| 2014 | Multi-FPGA Accelerator for Scalable Stencil Computation with Constant Memory BandwidthabstractStencil computation is one of the important kernels in scientific computations. However, sustained performance is limited owing to restriction on memory bandwidth, especially on multicore microprocessors and graphics processing units (GPUs) because of their small operational intensity. In this paper, we present a custom computing machine (CCM), called a scalable streaming-array (SSA), for high-performance stencil computations with multiple field-programmable gate arrays (FPGAs). We design SSA based on a domain-specific programmable concept, where CCMs are programmable with the minimum functionality required for an algorithm domain. We employ a deep pipelining approach over successive iterations to achieve linear scalability for multiple devices with a constant memory bandwidth. Prototype implementation using nine FPGAs demonstrates good agreement with a performance model, and achieves 260 and 236 GFlop/s for 2D and 3D Jacobi computation, which are 87.4 and 83.9 percent of the peak, respectively, with a memory bandwidth of only 2.0 GB/s. We also evaluate the performance of SSA for state-of-the-art FPGAs. Kentaro Sano, Yoshiaki Hatsuda, Satoru Yamamoto |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | Parallel and scalable custom computing for real-time fluid simulation on a cluster node with four tightly-coupled FPGAsabstractSummary form only given. Numerical simulation based on computational fluid dynamics (CFD) is now an indispensable technique especially in industry due to its acquisition capability of various data at a lower cost than experiments using a wind tunnel. The lattice Boltzmann method (LBM) is one of the CFD schemes, which is used to compute various problems including multiphase flow. LBM has good parallelism, but simultaneously requires many data to compute each lattice point, resulting in a low operational intensity. Consequently, the sustained performance of LBM is limited by memory bandwidth rather than arithmetic performance when computed by using general-purpose processors and GPUs. To make matters worse, insufficient bandwidth and high-latency of an interconnection network cause a relatively big overhead in parallel computing, especially in the case of strong-scaling. Kentaro Sano, Ryo Ito, Hayato Suzuki, Yoshiaki Kono |
FPL | 1 |
| 2012 | Scalability analysis of tightly-coupled FPGA-cluster for lattice Boltzmann computationabstractThis paper presents a performance model of an LBM accelerator to be implemented on a tightly-coupled FPGA cluster. In strong scaling, each accelerator node has a smaller computation as the nodes increase, and consequently communication overhead becomes apparent and limits the scalability. Our tightly-coupled FPGA cluster has the 1D ring of the accelerator-domain network (ADN) which allows FPGAs to send and receive data with low communication overhead. We propose the LBM accelerator architecture and its stream computation appropriate to use ADN. We formulate a sustained-performance model of the accelerator, which consists of three cases depending on one of the resource availability, the network bandwidth and the size of shift-registers. With the model, we show that the network bandwidth is much more important than the memory bandwidth. The wider the network bandwidth is, the more FPGAs can scale the sustained performance in computing a constant size of a lattice. This result demonstrates the importance of ADN in the tightly-coupled FPGA cluster. Yoshiaki Kono, Kentaro Sano, Satoru Yamamoto |
FPL | 2 |
| 2011 | Scalable Streaming-Array of Simple Soft-Processors for Stencil Computations with Constant Memory-BandwidthabstractStencil computation is one of the important kernels in scientific computations, however, the sustained performance is limited by memory bandwidth especially on multi-core microprocessors and GPGPUs due to its small operationalintensity. In this paper, we propose a scalable streaming-array (SSA) of simple soft-processors for high-performance stencil computation on multiple FPGAs. The SSA architecture allows a multi-device system to have linear scalability of computing performance by deeply pipelining with a constant bandwidth of an external-memory. We present an array-structure of programmable cores optimized for stencil computations and formulate a performance model of pipelined execution on the array. For Jacobi computations, SSA implemented on nine Stratix III FPGAs with the memory bandwidth of only 2 GB/s achieves 260 GFlop/s, corresponding to 87.4 % of its peak performance, at 1.3 GFlop/sW. We demonstrate that SSA provides almost linear speedup for larger than medium-sized computation as expected by the performance model. These high utilization and scalability show a big potential of custom computing on reconfigurable devices as a power-efficient and high-performance computing platform. Kentaro Sano, Yoshiaki Hatsuda, Satoru Yamamoto |
FCCM | 1 |
| 2011 | SW and HW co-design of Connect6 accelerator with scalable streaming coresabstractThis paper presents software and hardware co-design of an FPGA-based Connect6 solver with scalable streaming cores. The solver searches a game tree by using the miniMax algorithm with alpha-beta pruning. Since evaluation of board situations is the most time-consuming part, we adopted an approach to accelerate it with dedicated hardware while other parts are executed by software. We design a custom accelerator composed of multiple streaming-cores, each of which independently evaluates possible connectabilities with six stones. We use the ALTERA's SOPC (system on programmable chip) development tool to implement the system, which contains an NIOS II processor and the custom accelerator. The implemented system operates at 100 MHz on ALTERA Cyclone IV EP4CE115 FPGA of DE2-115 board. We could implement up to 32 streaming-cores on the FPGA. In a preliminary experiment where only a single core is used to search a game tree with a depth of 1, the FPGA-based solver wins against the given software opponent at a rate of 100 % with 33.2 stones on average. Kentaro Sano |
FPT | 1 |
| 2010 | FPGA-based lossless compressors of floating-point data streams to enhance memory bandwidthabstractThis paper presents an FPGA-based lossless compressor which directly compresses floating-point data streams to enhance the actual memory bandwidth of lattice Boltzmann method (LBM) accelerators. We show that the compression algorithms based on the 1D polynomial prediction are suitable for high-throughput hardware design. Moreover we show that integer operations provide comparable prediction performance to a floating-point predictor, while an integer predictor is expected to have smaller circuits than a floating-point one. We evaluate the compression ratio, the operating frequency and the resource consumption of the compressors with integer-based predictors through their prototype implementation using ALTERA Stratix III FPGA. We demonstrate that the implemented compressors dominate only 0.15 to 0.23 % of the entire logic resources and operate at 95 to 174 MHz to provide the compression ratio of up to 3.5, which means that we can enhance the memory bandwidth by a factor of 3.5 on average. Kazuya Katahira, Kentaro Sano, Satoru Yamamoto |
ASAP | 2 |
| 2010 | Segment-Parallel Predictor for FPGA-Based Hardware Compressor and Decompressor of Floating-Point Data Streams to Enhance Memory I/O BandwidthabstractThis paper presents segment-parallel prediction for high-throughput compression and decompression of floating-point data streams on an FPGA-based LBM accelerator. In order to enhance the actual memory I/O bandwidth of the accelerator, we focus on the prediction-based compression of floating-point data streams. Although hardware implementation is essential to high-throughput compression, the feedback loop in the decompressor is a bottleneck due to sequential predictions necessary for bit reconstruction. We introduce a segment-parallel approach to the 1D polynomial predictor to achieve the required throughput for decompression. We evaluate the compression ratio of the segment-parallel cubic prediction with various encoders of prediction difference. Kentaro Sano, Kazuya Katahira, Satoru Yamamoto |
DCC | 1 |
| 2010 | Local-and-global stall mechanism for systolic computational-memory array on extensible multi-FPGA systemabstractSo far we have proposed the systolic computational-memory (SCM) architecture for high-performance and scalable computation based on the finite difference methods. Although the SCM architecture has a completely parallel array structure, a lot of semiconductor devices are required to build a larger SCM array in the real world, which prefers a globally asynchronous and locally synchronous (GALS) design with different clock domains for system extensibility. This paper presents the local-and-global stall mechanism (LGSM) for an SCM array implemented over multiple FPGAs to guarantee the data-synchronization among FPGAs operating at different clocks. Prototype implementation with ALTERA Stratix III FPGAs shows that the proposed design does not give a big overhead to operating frequency and hardware resource utilization. We also evaluate the scalability of the SCM array over multiple FPGAs considering actual stall cycles. Luzhou Wang, Kentaro Sano, Satoru Yamamoto |
FPT | 2 |
| 2010 | FPGA-Array with Bandwidth-Reduction Mechanism for Scalable and Power-Efficient Numerical Simulations Based on Finite Difference MethodsabstractFor scientific numerical simulation that requires a relatively high ratio of data access to computation, the scalability of memory bandwidth is the key to performance improvement, and therefore custom-computing machines (CCMs) are one of the promising approaches to provide bandwidth-aware structures tailored for individual applications. In this article, we propose a scalable FPGA-array with bandwidth-reduction mechanism (BRM) to implement high-performance and power-efficient CCMs for scientific simulations based on finite difference methods. With the FPGA-array, we construct a systolic computational-memory array (SCMA), which is given a minimum of programmability to provide flexibility and high productivity for various computing kernels and boundary computations. Since the systolic computational-memory architecture of SCMA provides scalability of both memory bandwidth and arithmetic performance according to the array size, we introduce a homogeneously partitioning approach to the SCMA so that it is extensible over a 1D or 2D array of FPGAs connected with a mesh network. To satisfy the bandwidth requirement of inter-FPGA communication, we propose BRM based on time-division multiplexing. BRM decreases the required number of communication channels between the adjacent FPGAs at the cost of delay cycles. We formulate the trade-off between bandwidth and delay of inter-FPGA data-transfer with BRM. To demonstrate feasibility and evaluate performance quantitatively, we design and implement the SCMA of 192 processing elements over two ALTERA Stratix II FPGAs. The implemented SCMA running at 106MHz has the peak performance of 40.7 GFlops in single precision. We demonstrate that the SCMA achieves the sustained performances of 32.8 to 35.7 GFlops for three benchmark computations with high utilization of computing units. The SCMA has complete scalability to the increasing number of FPGAs due to the highly localized computation and communication. In addition, we also demonstrate that the FPGA-based SCMA is power-efficient: it consumes 69% to 87% power and requires only 2.8% to 7.0% energy of those for the same computations performed by a 3.4-GHz Pentium4 processor. With software simulation, we show that BRM works effectively for benchmark computations, and therefore commercially available low-end FPGAs with relatively narrow I/O bandwidth can be utilized to construct a scalable FPGA-array. Kentaro Sano, Luzhou Wang, Yoshiaki Hatsuda, Takanori Iizuka, Satoru Yamamoto |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2008 | Evaluating power and energy consumption of FPGA-based custom computing machines for scientific floating-point computationabstractThis paper evaluates the actual power consumption and the total energy for scientific floating-point computations accelerated by FPGA-based custom computing machines. With our FPGA-based machines: the streaming accelerator for computational fluid dynamics and the programmable systolic-array processor for numerical simulations based on difference schemes, we measure the power of the entire systems including a host PC and an FPGA board, and obtain the total energy for each computation. We report that the FPGAs perform the same computation with 5% to 30% of the total energy consumed by a microprocessor, while the FPGAs accelerate the computation. Kentaro Sano, Takeshi Nishikawa, Takayuki Aoki, Satoru Yamamoto |
FPT | 1 |
| 2007 | Systolic Architecture for Computational Fluid Dynamics on FPGAsabstractThis paper presents an FPGA-based flow solver based on the systolic architecture. We show that the fractional-step method employing central difference schemes can be expressed as a systolic algorithm, and therefore the systolic architecture is suitable for a dedicated processor to the flow solver. We have designed a 2D systolic array of cells, each of which has a micro-programmable data-path containing a MAC (multiplication and accumulation) unit and a local memory to store necessary data for computational fluid dynamics. With ALTERA Stratix II FPGA, we implemented 96(= 12 times 8) cells running at 60 MHz. Since the MAC unit has both an adder and a multiplier for single-precision floating-point numbers, the total peak performance is 11.5(= 96times60 MHztimes2) GFlops. We made a choice of 2D square driven cavity flow as a benchmark computation based on the fractional-step method. For this computation, the FPGA-based processor running only at 60 MHz achieved 7.14 and 6.41 times faster computations than Pentium4 processor at 3.2 GHz and Itanium2 at 1.4 GHz, respectively. Kentaro Sano, Takanori Iizuka, Satoru Yamamoto |
FCCM | 1 |
| 2007 | FPGA-based Streaming Computation for Lattice Boltzmann MethodabstractThis paper presents an FPGA-based streaming computation for the lattice Boltzmann method (LBM) to simulate fluid flow with floating-point calculations. LBM is suitable for streaming computation because of its parallelism and regularity. We optimize the equations of LBM, and then formulate a streaming computation. To design an efficient data-path for throughput and hardware resource utilization, we introduce multiple cycle inputs and computing-unit sharing to the streaming data-path. The streaming accelerator implemented on a Virtex-4 FPGA with PCTExpress x8 interface achieves 2.93 and 2.46 times faster computation than a 3.4 GHz Pentium4 processor and a 2.2 GHz Opteron processor, respectively, for 2-dimensional time-dependent fluid dynamics problems. Kentaro Sano, Oliver Pell, Wayne Luk, Satoru Yamamoto |
FPT | 1 |
| 2004 | Parallel competitive learning algorithm for fast codebook design on partitioned spaceabstractVector quantization (VQ) is an attractive technique for lossy data compression, which is a key technology for data storage and/or transfer. So far, various competitive learning (CL) algorithms have been proposed to design optimal codebooks presenting quantization with minimized errors. However, their practical use has been limited for large scale problems, due to the computational complexity of competitive learning. This work presents a parallel competitive learning algorithm for fast code-book design based on space partitioning. The algorithm partitions input-vector space into some subspaces, and independently designs corresponding subcodebooks for these subspaces with computational complexity reduced. Independent processing on different subspaces can be processed in parallel without synchronization overhead, resulting in high scalability. We perform experiments of parallel codebook design on a commodity PC cluster with 8 nodes. Experimental results show that the high speedup of the codebook design is obtained without increase of quantization errors. Shintaro Momose, Kentaro Sano, Ken-Ichi Suzuki, Tadao Nakamura |
CLUSTER | 2 |
| 2004 | Differential coding scheme for efficient parallel image composition on a PC cluster system
Kentaro Sano, Yusuke Kobayashi 0005, Tadao Nakamura |
Parallel Comput. | 1 |
| 2004 | Efficient parallel processing of competitive learning algorithms
Kentaro Sano, Shintaro Momose, Hiroyuki Takizawa, Hiroaki Kobayashi, Tadao Nakamura |
Parallel Comput. | 1 |
| 2001 | 3DCGiRAM: An Intelligent Memory Architecture for Photo-Realistic Image SynthesisabstractThis paper proposes an intelligent memory architecture for photo-realistic image synthesis, named 3DCGiRAM. The 3DCGiRAM has a hardware-accelerated 3D line generator which finds objects that are likely to intersect traced rays. It also has functional memory cells, each of which is composed of graphics logic and its local memory to detect intersecting objects and to calculate intensities. A distributed frame buffer is employed to alleviate the access conflicts of functional cells to the frame buffer as well as to compose globally illuminated intensities at screen pixels. As the graphics processing capability is localized to data in the 3DCGiRAM through memory-logic merged LSI technology, a scalability and modularity similar to those of conventional memory modules can be expected. The experimental results show that a single 3DCGiRAM module running at 200 MHz with a memory bandwidth of 6.4 GB will be able to synthesize a ray-traced walk-through animation at a rate of one frame per second. Hiroaki Kobayashi, Ken-Ichi Suzuki, Kentaro Sano, Yoshiyuki Kaeriyama, Yasumasa Saida, Nobuyuki Ohba, Tadao Nakamura |
ICCD | 3 |