VLDB 2026 Research / reviewers in the wild / expert
Satoru Yamamoto
dblp:91/5244
· DBLP profile ↗
25ranked-venue papers
6as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12Applied, interdisciplinary, general and emerging computing · 12 · 6 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | First Cross- and Inter-Band Calibrations of the Hyperspectral Imager Suite Using Off-Nadir Quasi-Simultaneous Overpass CounterpartsabstractThe hyperspectral imager suite (HISUI) is a recently commissioned spaceborne hyperspectral sensor that is expected to provide higher-quality information for terrestrial and aquatic monitoring at a global scale. To maximize the utility of the detailed spectra observed by HISUI, calibration and validation activities have been intensively conducted for the past several years. However, research on long-term monitoring of sensor radiometric characteristics and checking inter-band radiometric consistency is still limited. This study provided the first results of cross- and inter-band calibrations using five scenes acquired over a desert calibration site during 2021–2023, and calculated relative cross-calibration coefficients (RCCCs) to calibrate all the HISUI bands. To make full use of the limited amount of data archived over the calibration site, off-nadir scenes obtained by wide-swath counterpart sensors [i.e., the moderate resolution imaging spectroradiometer (MODIS)] were used for the cross-calibration, with correction by a bidirectional reflectance distribution function. The obtained RCCCs successfully mitigated the inter-band radiometric variability of HISUI data, which was consistent with the in situ surface reflectance data observed by the radiometric calibration network (RadCalNet). Time-series analysis further revealed a notable discrepancy between the visible to near-infrared and short-wave infrared (SWIR) spectral domains after the International Space Station (ISS) maneuvers in spring 2022. This article is also a prompt report of this phenomenon and will soon contribute to HISUI radiometric calibration, particularly after the maneuver events. Hiroki Mizuochi, Satoshi Tsuchida, Satoru Yamamoto, Minoru Urai, Moe Matsuoka, Ayame Ikeda, Koki Iwao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Analysis Based on Onboard Lamp and Lunar Vicarious Calibrations for Sensitivity Degradation of a Hyperspectral SensorabstractWe assessed the degradation of the radiance lamp in the onboard calibration system of the Spectral Profiler, a hyperspectral sensor onboard the Japanese Selenological and Engineering Explorer (SELENE), which was launched in September 2007 and observed the Moon until June 2009. We found that the output of the radiance lamp soon after launch increased by 3.2%–4.6% as compared with the prelaunch output, after which it decreased over time. The total decrease during the mission lifetime of about 1.5 years after the increase of 3.2%–4.6% was ~1.1%–1.5%, which is slightly larger than the$0.8\,\mu \text{m}$. The trend of the wavelength dependence was shown to be consistent with a model of blackbody radiance with a decrease in temperature of the lamp of about −5 K, suggesting that the degradation of the lamp was mostly due to temperature decrease of about −3 K/yr. Our results may provide an important constraint on the quantitative assessment of the degradation of the sensitivity of remote-sensing sensors using an onboard calibration system, such as the Advanced Spaceborne Thermal Emission and Reflection Radiometer. Satoru Yamamoto, Hiroki Mizuochi, Moe Matsuoka, Koki Iwao, Satoshi Tsuchida |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Initial Analysis of Spectral Smile Calibration of Hyperspectral Imager Suite (HISUI) Using Atmospheric Absorption BandsabstractThis paper reports on the initial analysis of spectral smile calibration of the Hyperspectral Imager Suite (HISUI) onboard the International Space Station, which has been continuously acquiring data since September 4, 2020. HISUI is an optical hyperspectral imager consisting of two subsystems: VNIR covering 400 to 980nm at intervals of 10 nm, and SWIR covering 895 to 2481nm at 12.5nm intervals. Based on the atmospheric correction for actual observation images, we assessed cross-track dependences of the wavelength deviation (spectral smile) and the full-width at half-maximum (FWHM) of the HISUI response function. We found that significant spectral smile was observed, with maximum variations of 1.8nm in VNIR and 4.3–4.5nm in SWIR. In addition, the cross-track variation of FWHM was observed with maximum variations of 5.0nm for VNIR and 2.5–3.5nm for SWIR.We used the results to model the smile functions to update a smile correction table in the internal calibration system of HISUI. Then, we evaluated how the smile functions reduce the spectral smile in the data acquired after the update on September 27, 2021.We confirmed that VNIR showed a nearly flat profile within 0.25nm with a nearly constant FWHM. For SWIR, although a slight amount of spectral smile and a variation of FWHM were still observed partly due to wavelength dependence in the spectral smile, the spectral smile was reduced to <~2.2 nm. This study demonstrated that wavelength calibration using actual observation images for ground surfaces is important for the characterization of hyperspectral sensors. Satoru Yamamoto, Satoshi Tsuchida, Minoru Urai, Hiroki Mizuochi, Koki Iwao, Akira Iwasaki |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Initial Onboard Calibration Results of the HISUI Hyperspectral SensorabstractHISUI, the Japanese hyperspectral sensor, was launched on December 6, 2019 and the first image was taken on September 4, 2020. The first HISUI calibration was conducted on September 11. Wavelength and radiance calibration were conducted to evaluate characteristic change during the launch process. We found 2.3 nm and 5.5 nm wavelength shifts for VNIR and SWIR, respectively using SRM2025a filter onboard HISUI calibration system. Sensitivity changes during the launch process were less than 1.0 % that is smaller than the radiometric accuracy specifications. Minoru Urai, Satoshi Tsuchida, Satoru Yamamoto, Tetsushi Tachikawa, Akira Iwasaki, Juntaro Ishii |
IGARSS | 3 |
| 2017 | An Automated Method for Crater Counting Using Rotational Pixel Swapping MethodabstractWe develop a fully automated algorithm for determining the geological ages by crater counting from the digital terrain model (DTM) and the digital elevation model (DEM) taken by remote-sensing observations. The algorithm is based on the rotational pixel swapping method, which uses a multiplication operation between the original DTM/DEM data and the rotated data to detect impact craters. Our method does not need binarization and/or noise reduction, because noise components are automatically erased. We show that our method can detect not only simple craters but also complex circular structures such as imperfect, degraded, or overlapping craters. We demonstrate that this method succeeds in the automatic detection of hundreds to thousands of impact craters, and the estimated ages are consistent with those by manual counting in previous works. In addition, it is shown that the calculation time by this method is more than several hundred times faster than by previous methods. Satoru Yamamoto, Tsuneo Matsunaga, Ryosuke Nakamura, Yasuhito Sekine, Naru Hirata, Yasushi Yamaguchi 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | FPGA-Based Scalable and Power-Efficient Fluid Simulation using Floating-Point DSP BlocksabstractHigh-performance and low-power computation is required for large-scale fluid dynamics simulation. Due to the inefficient architecture and structure of CPUs and GPUs, they now have a difficulty in improving power efficiency for the target application. Although FPGAs become promising alternatives for power-efficient and high-performance computation due to their new architecture having floating-point (FP) DSP blocks, their relatively narrow memory bandwidth requires an appropriate way to fully exploit the advantage. This paper presents an architecture and design for scalable fluid simulation based on data-flow computing with a state-of-the-art FPGA. To exploit available hardware resources including FP DSPs, we introduce spatial and temporal parallelism to further scale the performance by adding more stream processing elements (SPEs) in an array. Performance modeling and prototype implementation allow us to explore the design space for both the existing Altera Arria10 and the upcoming Intel Stratix10 FPGAs. We demonstrate that Arria10 10AX115 FPGA achieves 519 GFlops at 9.67 GFlops/W only with a stream bandwidth of 9.0 GB/s, which is 97.9 percent of the peak performance of 18 implemented SPEs. We also estimate that Stratix10 FPGA can scale up to 6844 GFlops by combining spatial and temporal parallelism adequately. Kentaro Sano, Satoru Yamamoto |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Bandwidth Compression of Floating-Point Numerical Data Streams for FPGA-Based High-Performance ComputingabstractAlthough computational performance is often limited by insufficient bandwidth to/from an external memory, it is not easy to physically increase off-chip memory bandwidth. In this study, we propose a hardware-based bandwidth compression technique that can be applied to field-programmable gate array-- (FPGA) based high-performance computation with a logically wider effective memory bandwidth. Our proposed hardware approach can boost the performance of FPGA-based stream computations by applying a data compression technique to effectively transfer more data streams. To apply this data compression technique to bandwidth compression via hardware, several requirements must first be satisfied, including an acceptable level of compression performance and a sufficiently small hardware footprint. Our proposed hardware-based bandwidth compressor utilizes an efficient prediction-based data compression algorithm. Moreover, we propose a multichannel serializer and deserializer that enable applications to use multiple channels of computational data with the bandwidth compression. The serializer encodes compressed data blocks of multiple channels into a data stream, which is efficiently written to an external memory. Based on preliminary evaluation, we define an encoding format considering both high compression ratio and small hardware area. As a result, we demonstrate that our area saving bandwidth compressor increases performance of an FPGA-based fluid dynamics simulation by deploying more processing elements to exploit spatial parallelism with the enhanced memory bandwidth. Tomohiro Ueno, Kentaro Sano, Satoru Yamamoto |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2015 | Rotational Pixel Swapping Method for Detection of Circular Features in Binary ImagesabstractWe propose a new automatic method called the rotational pixel swapping (RPSW) method to detect circular features in binary images of remote sensing images. The method is based on a multiplication operation between the original image and the rotated images. We show that the RPSW selectively enhances rotational symmetric patterns and weakens nonrotational symmetric patterns, including noise components, without any noise reduction processes. The method can detect not only simple circles but also more complex circular features such as incomplete ring structures or several concentric rings. Furthermore, we demonstrate that the RPSW provides the stable detection of circular features such as terrestrial impact structures, which are irregular imperfect circular shapes, in binary images based on Earth-observation satellite images. The RPSW would provide a potential method of future surveys or statistical studies using huge data sets of multiband or hyperspectral images obtained by Earth-observation satellites. Satoru Yamamoto, Tsuneo Matsunaga, Ryosuke Nakamura, Yasuhito Sekine, Naru Hirata, Yasushi Yamaguchi 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Bandwidth compression of multiple numerical data streams for high performance custom computingabstractBandwidth compression improves the performance of stream computing by enhancing an effective bandwidth. To apply the bandwidth compression to numerical applications such as numerical simulations, a compressor has to handle multiple data streams. In this paper, we describe a design of an FPGA-based bandwidth compressor for high performance stream computation. For synchronization of original data in multiple compressed streams with different bit-rate, we propose a data block transmission scheduler and explore a design space to reduce the size of their barrel shifters. Tomohiro Ueno, Ryo Ito, Kentaro Sano, Satoru Yamamoto |
ASAP | 4 |
| 2014 | Effective observation planning and its simulation of a Japanese spaceborne sensor: Hyperspectral imager suite (HISUI)abstractHyperspectral Imager Suite (HISUI) is a Japanese future spaceborne hyperspectral instrument being developed by Ministry of Economy, Trade, and Industry (METI) and will be launched in 2017 or later. In HISUI project, observation strategy is important especially for hyperspectral sensor, and relationship between the limitations of sensor operation and the planned observation scenarios have to be studied. Using observation coverage simulation program and we estimate progress of observation coverage of image with days after launch. We found that HISUI can make 4 times repeated observations for protected area (20 million km2in the world). And about 70 % of land surface can be observed over 5 years. We also found that the developed rules to avoid cloudy are will improve the area coverage up to 2.4 %. Kenta Ogawa, Tsuneo Matsunaga, Satoru Yamamoto, Osamu Kashimura, Tetsushi Tachikawa, Satoshi Tsuchida, Jun Tanii, Shuichi Rokugawa |
IGARSS | 3 |
| 2014 | Calibration of NIR 2 of Spectral Profiler Onboard Kaguya/SELENEabstractThe Spectral Profiler (SP) is a visible-near infrared spectrometer onboard the Japanese Selenological and Engineering Explorer (SELENE), which was launched in 2007 and observed the Moon until June 2009. The SP consists of two gratings and three linear-array detectors: VIS (0.5-1.0 μm), NIR 1 (0.9- 1.7 μm), and NIR 2 (1.7-2.6 μm). In this paper, we propose a new method for radiometric calibration of NIR 2, specifically for the dark output (background) estimate, which is different from the previous method used for VIS and NIR 1. We show that the reflectance spectra of NIR 2 derived from the new radiometric calibration show less noise than those of the previous method. Based on an analysis of the reflectance spectra at exposure sites of the end-member minerals on the lunar surface, we demonstrated that the spectral features of the 2-μm band in the NIR 2 spectra are consistent with those expected from the minerals inferred from the features of the 1-μm band in the VIS and NIR 1 spectra. Finally, we examined the repeatability of the radiometric calibration of NIR 2 using the SP data near the Apollo 16 landing site observed at four different times. The typical difference in the reflectance at wavelengths <;~2.1 μm was a few percent, which is within the uncertainty due to the error in the background estimate, suggesting that there was no significant change in the sensitivity of NIR 2 over the mission period. Satoru Yamamoto, Tsuneo Matsunaga, Yoshiko Ogawa, Ryosuke Nakamura, Yasuhiro Yokota, Makiko Ohtake, Jun'ichi Haruyama, Tomokatsu Morota, Chikatoshi Honda, Takahiro Hiroi, Shinsuke Kodama |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Multi-FPGA Accelerator for Scalable Stencil Computation with Constant Memory BandwidthabstractStencil computation is one of the important kernels in scientific computations. However, sustained performance is limited owing to restriction on memory bandwidth, especially on multicore microprocessors and graphics processing units (GPUs) because of their small operational intensity. In this paper, we present a custom computing machine (CCM), called a scalable streaming-array (SSA), for high-performance stencil computations with multiple field-programmable gate arrays (FPGAs). We design SSA based on a domain-specific programmable concept, where CCMs are programmable with the minimum functionality required for an algorithm domain. We employ a deep pipelining approach over successive iterations to achieve linear scalability for multiple devices with a constant memory bandwidth. Prototype implementation using nine FPGAs demonstrates good agreement with a performance model, and achieves 260 and 236 GFlop/s for 2D and 3D Jacobi computation, which are 87.4 and 83.9 percent of the peak, respectively, with a memory bandwidth of only 2.0 GB/s. We also evaluate the performance of SSA for state-of-the-art FPGAs. Kentaro Sano, Yoshiaki Hatsuda, Satoru Yamamoto |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Usability of lunar reflectance model based on SELENE/SP for planned HISUI radiometric calibrationabstractWe have developed a method for radiometric calibration of HISUI's hyper and multi-spectral sensors using a lunar reflectance model developed from SELENE SP data, which involves lunar surface reflectance and photometric properties. For evaluating the utilization of the model, we simulated a lunar observation conducted by ASTER of its three visible and infrared bands and confirmed the model describes lunar surface photometric properties correctly because correlation coefficients of observed and modeled radiance exceed 0.99 for all bands. Although absolute radiance shows some discrepancy between the observed and the simulated Moon in visible band, the model is, at least, useful to evaluate relative degradation of sensors. Toru Kouyama, Yoshiaki Ishihara, Ryosuke Nakamura, Satoshi Tsuchida, Tsuneo Matsunaga, Fumihiro Sakuma, Yasuhiro Yokota, Hirokazu Yamamoto, Satoru Yamamoto |
IGARSS | 9 |
| 2013 | Observation planning and its coverage simulation of a Japanese spaceborne sensor: Hyperspectral Imager Suite (HISUI)abstractAs mentioned above, the simulation program is useful to investigate observation-planning strategy and rules for improving efficiency of observations. Kenta Ogawa, Tsuneo Matsunaga, Satoru Yamamoto, Osamu Kashimura, Tetsushi Tachikawa, Satoshi Tsuchida, Jun Tanii, Shuichi Rokugawa |
IGARSS | 3 |
| 2012 | Scalability analysis of tightly-coupled FPGA-cluster for lattice Boltzmann computationabstractThis paper presents a performance model of an LBM accelerator to be implemented on a tightly-coupled FPGA cluster. In strong scaling, each accelerator node has a smaller computation as the nodes increase, and consequently communication overhead becomes apparent and limits the scalability. Our tightly-coupled FPGA cluster has the 1D ring of the accelerator-domain network (ADN) which allows FPGAs to send and receive data with low communication overhead. We propose the LBM accelerator architecture and its stream computation appropriate to use ADN. We formulate a sustained-performance model of the accelerator, which consists of three cases depending on one of the resource availability, the network bandwidth and the size of shift-registers. With the model, we show that the network bandwidth is much more important than the memory bandwidth. The wider the network bandwidth is, the more FPGAs can scale the sustained performance in computing a constant size of a lattice. This result demonstrates the importance of ADN in the tightly-coupled FPGA cluster. Yoshiaki Kono, Kentaro Sano, Satoru Yamamoto |
FPL | 3 |
| 2012 | USage of cloud climate data in operation misson plan simulation for Japanese future hyperspectral and multispectral senor: HISUIabstractHyperspectral Imager Suite (HISUI) is a Japanese future spaceborne hyperspectral instrument [1][2] being developed by Ministry of Economy, Trade, and Industry (METI) and will be launched in 2015 or later. HISUI's operation strategic study is described in this paper. In HISUI project, Operation Mission Planning (OMP) team will make long term and short term strategy of the observation and sensor operation plan. OMP is important for HISUI especially for hyperspectral sensor, and relationship between the limitations of sensor operation and the planned observation scenarios have to be studied. Major factors of the limitations are the combinations of downlink rate, observation time (15 minutes per orbit) and the swath of the sensor (30 km). The achievements of global mapping or repeated observations of specific site need to be simulated precisely before launch [3]. We have prepared daily global high resolution (30 second in latitude and longitude) climate data for the simulation. Kenta Ogawa, Tsuneo Matsunaga, Satoru Yamamoto, Osamu Kashimura, Tetsushi Tachikawa, Satoshi Tsuchida, Jun Tanii, Shuichi Rokugawa |
IGARSS | 3 |
| 2011 | Scalable Streaming-Array of Simple Soft-Processors for Stencil Computations with Constant Memory-BandwidthabstractStencil computation is one of the important kernels in scientific computations, however, the sustained performance is limited by memory bandwidth especially on multi-core microprocessors and GPGPUs due to its small operationalintensity. In this paper, we propose a scalable streaming-array (SSA) of simple soft-processors for high-performance stencil computation on multiple FPGAs. The SSA architecture allows a multi-device system to have linear scalability of computing performance by deeply pipelining with a constant bandwidth of an external-memory. We present an array-structure of programmable cores optimized for stencil computations and formulate a performance model of pipelined execution on the array. For Jacobi computations, SSA implemented on nine Stratix III FPGAs with the memory bandwidth of only 2 GB/s achieves 260 GFlop/s, corresponding to 87.4 % of its peak performance, at 1.3 GFlop/sW. We demonstrate that SSA provides almost linear speedup for larger than medium-sized computation as expected by the performance model. These high utilization and scalability show a big potential of custom computing on reconfigurable devices as a power-efficient and high-performance computing platform. Kentaro Sano, Yoshiaki Hatsuda, Satoru Yamamoto |
FCCM | 3 |
| 2011 | Preflight and In-Flight Calibration of the Spectral Profiler on Board SELENE (Kaguya)abstractThe Spectral Profiler (SP) is a visible-near infrared spectrometer on board the Japanese Selenological and Engineering Explorer, which was launched in 2007 and observed the Moon until June 2009. The SP consists of two gratings and three linear-array detectors: VIS (0.5-1.0 μm ), NIR 1 (0.9-1.7 μm), and NIR 2 (1.7-2.6 μm). In this paper, we characterize the radiometric and spectral properties of VIS and NIR 1 using in-flight observational data as well as preflight data derived in laboratory experiments using a calibrated integrating sphere. We also proposed new methods for radiometric calibration, specifically methods for nonlinearity correction, wavelength correction, and the correction of the radiometric calibration coefficients affected by the water vapor. After all the corrections, including the photometric correction, we obtained the reflectance spectra for the lunar surface. Finally, we examined the stability of the SP using the SP data near the Apollo 16 landing site observed at four different times. The difference in reflectance among these four observations was less than ~ ±1% for most of the bands, suggesting that the degradation of the SP is not significant over the mission period. Satoru Yamamoto, Tsuneo Matsunaga, Yoshiko Ogawa, Ryosuke Nakamura, Yasuhiro Yokota, Makiko Ohtake, Jun'ichi Haruyama, Tomokatsu Morota, Chikatoshi Honda, Takahiro Hiroi, Shinsuke Kodama |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2010 | FPGA-based lossless compressors of floating-point data streams to enhance memory bandwidthabstractThis paper presents an FPGA-based lossless compressor which directly compresses floating-point data streams to enhance the actual memory bandwidth of lattice Boltzmann method (LBM) accelerators. We show that the compression algorithms based on the 1D polynomial prediction are suitable for high-throughput hardware design. Moreover we show that integer operations provide comparable prediction performance to a floating-point predictor, while an integer predictor is expected to have smaller circuits than a floating-point one. We evaluate the compression ratio, the operating frequency and the resource consumption of the compressors with integer-based predictors through their prototype implementation using ALTERA Stratix III FPGA. We demonstrate that the implemented compressors dominate only 0.15 to 0.23 % of the entire logic resources and operate at 95 to 174 MHz to provide the compression ratio of up to 3.5, which means that we can enhance the memory bandwidth by a factor of 3.5 on average. Kazuya Katahira, Kentaro Sano, Satoru Yamamoto |
ASAP | 3 |
| 2010 | Segment-Parallel Predictor for FPGA-Based Hardware Compressor and Decompressor of Floating-Point Data Streams to Enhance Memory I/O BandwidthabstractThis paper presents segment-parallel prediction for high-throughput compression and decompression of floating-point data streams on an FPGA-based LBM accelerator. In order to enhance the actual memory I/O bandwidth of the accelerator, we focus on the prediction-based compression of floating-point data streams. Although hardware implementation is essential to high-throughput compression, the feedback loop in the decompressor is a bottleneck due to sequential predictions necessary for bit reconstruction. We introduce a segment-parallel approach to the 1D polynomial predictor to achieve the required throughput for decompression. We evaluate the compression ratio of the segment-parallel cubic prediction with various encoders of prediction difference. Kentaro Sano, Kazuya Katahira, Satoru Yamamoto |
DCC | 3 |
| 2010 | Local-and-global stall mechanism for systolic computational-memory array on extensible multi-FPGA systemabstractSo far we have proposed the systolic computational-memory (SCM) architecture for high-performance and scalable computation based on the finite difference methods. Although the SCM architecture has a completely parallel array structure, a lot of semiconductor devices are required to build a larger SCM array in the real world, which prefers a globally asynchronous and locally synchronous (GALS) design with different clock domains for system extensibility. This paper presents the local-and-global stall mechanism (LGSM) for an SCM array implemented over multiple FPGAs to guarantee the data-synchronization among FPGAs operating at different clocks. Prototype implementation with ALTERA Stratix III FPGAs shows that the proposed design does not give a big overhead to operating frequency and hardware resource utilization. We also evaluate the scalability of the SCM array over multiple FPGAs considering actual stall cycles. Luzhou Wang, Kentaro Sano, Satoru Yamamoto |
FPT | 3 |
| 2010 | FPGA-Array with Bandwidth-Reduction Mechanism for Scalable and Power-Efficient Numerical Simulations Based on Finite Difference MethodsabstractFor scientific numerical simulation that requires a relatively high ratio of data access to computation, the scalability of memory bandwidth is the key to performance improvement, and therefore custom-computing machines (CCMs) are one of the promising approaches to provide bandwidth-aware structures tailored for individual applications. In this article, we propose a scalable FPGA-array with bandwidth-reduction mechanism (BRM) to implement high-performance and power-efficient CCMs for scientific simulations based on finite difference methods. With the FPGA-array, we construct a systolic computational-memory array (SCMA), which is given a minimum of programmability to provide flexibility and high productivity for various computing kernels and boundary computations. Since the systolic computational-memory architecture of SCMA provides scalability of both memory bandwidth and arithmetic performance according to the array size, we introduce a homogeneously partitioning approach to the SCMA so that it is extensible over a 1D or 2D array of FPGAs connected with a mesh network. To satisfy the bandwidth requirement of inter-FPGA communication, we propose BRM based on time-division multiplexing. BRM decreases the required number of communication channels between the adjacent FPGAs at the cost of delay cycles. We formulate the trade-off between bandwidth and delay of inter-FPGA data-transfer with BRM. To demonstrate feasibility and evaluate performance quantitatively, we design and implement the SCMA of 192 processing elements over two ALTERA Stratix II FPGAs. The implemented SCMA running at 106MHz has the peak performance of 40.7 GFlops in single precision. We demonstrate that the SCMA achieves the sustained performances of 32.8 to 35.7 GFlops for three benchmark computations with high utilization of computing units. The SCMA has complete scalability to the increasing number of FPGAs due to the highly localized computation and communication. In addition, we also demonstrate that the FPGA-based SCMA is power-efficient: it consumes 69% to 87% power and requires only 2.8% to 7.0% energy of those for the same computations performed by a 3.4-GHz Pentium4 processor. With software simulation, we show that BRM works effectively for benchmark computations, and therefore commercially available low-end FPGAs with relatively narrow I/O bandwidth can be utilized to construct a scalable FPGA-array. Kentaro Sano, Luzhou Wang, Yoshiaki Hatsuda, Takanori Iizuka, Satoru Yamamoto |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2008 | Evaluating power and energy consumption of FPGA-based custom computing machines for scientific floating-point computationabstractThis paper evaluates the actual power consumption and the total energy for scientific floating-point computations accelerated by FPGA-based custom computing machines. With our FPGA-based machines: the streaming accelerator for computational fluid dynamics and the programmable systolic-array processor for numerical simulations based on difference schemes, we measure the power of the entire systems including a host PC and an FPGA board, and obtain the total energy for each computation. We report that the FPGAs perform the same computation with 5% to 30% of the total energy consumed by a microprocessor, while the FPGAs accelerate the computation. Kentaro Sano, Takeshi Nishikawa, Takayuki Aoki, Satoru Yamamoto |
FPT | 4 |
| 2007 | Systolic Architecture for Computational Fluid Dynamics on FPGAsabstractThis paper presents an FPGA-based flow solver based on the systolic architecture. We show that the fractional-step method employing central difference schemes can be expressed as a systolic algorithm, and therefore the systolic architecture is suitable for a dedicated processor to the flow solver. We have designed a 2D systolic array of cells, each of which has a micro-programmable data-path containing a MAC (multiplication and accumulation) unit and a local memory to store necessary data for computational fluid dynamics. With ALTERA Stratix II FPGA, we implemented 96(= 12 times 8) cells running at 60 MHz. Since the MAC unit has both an adder and a multiplier for single-precision floating-point numbers, the total peak performance is 11.5(= 96times60 MHztimes2) GFlops. We made a choice of 2D square driven cavity flow as a benchmark computation based on the fractional-step method. For this computation, the FPGA-based processor running only at 60 MHz achieved 7.14 and 6.41 times faster computations than Pentium4 processor at 3.2 GHz and Itanium2 at 1.4 GHz, respectively. Kentaro Sano, Takanori Iizuka, Satoru Yamamoto |
FCCM | 3 |
| 2007 | FPGA-based Streaming Computation for Lattice Boltzmann MethodabstractThis paper presents an FPGA-based streaming computation for the lattice Boltzmann method (LBM) to simulate fluid flow with floating-point calculations. LBM is suitable for streaming computation because of its parallelism and regularity. We optimize the equations of LBM, and then formulate a streaming computation. To design an efficient data-path for throughput and hardware resource utilization, we introduce multiple cycle inputs and computing-unit sharing to the streaming data-path. The streaming accelerator implemented on a Virtex-4 FPGA with PCTExpress x8 interface achieves 2.93 and 2.46 times faster computation than a 3.4 GHz Pentium4 processor and a 2.2 GHz Opteron processor, respectively, for 2-dimensional time-dependent fluid dynamics problems. Kentaro Sano, Oliver Pell, Wayne Luk, Satoru Yamamoto |
FPT | 4 |