EDBT 2026 Demo / reviewers in the wild / expert
Akira Nukada
dblp:74/1996
· DBLP profile ↗
28ranked-venue papers
8as first author
4since 2021 · last 2025
0000-0001-7959-6975ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
High-performance computing · 57% GPUs and heterogeneous computing · 32% Interconnection networks and networks-on-chip · 11% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU computing |
0.4 | 4 | 2011 | Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputer · SC 2011 An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production Code · SC 2010 Auto-tuning 3-D FFT library for CUDA GPUs · SC 2009 |
High-performance computing
scientific computing systems |
0.3 | 4 | 2011 | Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputer · SC 2011 An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production Code · SC 2010 Auto-tuning 3-D FFT library for CUDA GPUs · SC 2009 |
High-performance computing › fast fourier transform
3D FFT |
0.2 | 2 | 2009 | Auto-tuning 3-D FFT library for CUDA GPUs · SC 2009 Bandwidth intensive 3-D FFT kernel for GPUs using CUDA · SC 2008 |
Interconnection networks and networks-on-chip › interprocessor communication
all-to-all communication |
0.1 | 1 | 2012 | Scalable multi-GPU 3-D FFT for TSUBAME 2.0 supercomputer · SC 2012 |
GPUs and heterogeneous computing
multi-GPU computing |
0.1 | 1 | 2012 | Scalable multi-GPU 3-D FFT for TSUBAME 2.0 supercomputer · SC 2012 |
High-performance computing
supercomputing |
0.1 | 1 | 2012 | Scalable multi-GPU 3-D FFT for TSUBAME 2.0 supercomputer · SC 2012 |
High-performance computing › scientific computing systems
phase field simulation |
0.1 | 1 | 2011 | Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputer · SC 2011 |
High-performance computing › scientific computing systems
weather prediction |
0.1 | 1 | 2010 | An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production Code · SC 2010 |
High-performance computing › performance optimization
auto-tuning |
0.1 | 1 | 2009 | Auto-tuning 3-D FFT library for CUDA GPUs · SC 2009 |
High-performance computing
numerical libraries |
0.1 | 1 | 2006 | Poster reception - Scalable software infrastructure project · SC 2006 |
Interconnection networks and networks-on-chip › cluster interconnect
infiniband |
0.0 | 1 | 2012 | Scalable multi-GPU 3-D FFT for TSUBAME 2.0 supercomputer · SC 2012 |
Methods — techniques the papers use, named apart from their topics
auto-tuning · 0.2MPI · 0.1CUDA memory copy · 0.1phase-field method · 0.1FFT · 0.1CUDA · 0.1shared memory bank conflict avoidance · 0.1thread and register optimization · 0.1stride memory access avoidance · 0.1shared memory utilization · 0.1matrix computation framework · 0.1iterative solver · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating General Relativistic Radiation Magnetohydrodynamic Simulations with GPUs
Ryohei Kobayashi 0001, Hiroyuki R. Takahashi, Akira Nukada, Yuta Asahina, Taisuke Boku, Ken Ohsuga |
HPC Asia | 3 |
| 2024 | FCUFS: Core-Level Frequency Tuning for Energy Optimization on Intel ProcessorsabstractSacrificing minor performance for better energy efficiency is effective in reducing the energy consumption of supercomputers. Recent studies have utilized some frequency tuning and power capping features of Intel processors to decrease energy consumption in supercomputers. However, two main issues persist: 1) current methods do not account for variability between cores, leading to insufficient energy savings in mixed workloads; and 2) they fail to control the extent of performance loss following frequency tuning. To address these issues, we developed the FCUFS framework, which includes two components: 1) a neural network for predicting performance and power; and 2) a strategy optimization algorithm for selecting frequencies with controllable performance loss. We evaluated FCUFS on Intel mainstream processors across 15 dedicated workloads and 5 mixed workloads. On a single dual-socket Ice Lake-SP server, at a 5% performance loss target, the average energy savings were 11.4% for dedicated workloads and 13.8% for mixed workloads, with average performance losses of 2.9% and 3.5%, respectively. At a 10% performance loss target, the energy savings increased to 14.2% for dedicated workloads and 14.4% for mixed workloads, with performance losses of 8.7% and 9.4%, respectively. When scaling up to 2,048 cores, the average energy savings were 9.7% with 4.4% performance loss. The results show that FCUFS achieves consistent energy savings across dedicated and mixed workloads while maintaining controllable performance loss. Hongjian Zhang, Akira Nukada, Qiucheng Liao |
CLUSTER | 2 |
| 2023 | Efficient checkpoint/Restart of CUDA applications
Akira Nukada, Taichiro Suzuki, Satoshi Matsuoka |
Parallel Comput. | 1 |
| 2021 | Performance Optimization of Allreduce Operation for Multi-GPU SystemsabstractAllreduce is one of the important collective commu-nications used in distributed deep learning. We present a novel hybrid allreduce algorithm optimized for multi-GPU systems with NV-Link, which is current main computing platform for data-parallel distributed deep learning. The hybrid algorithm efficiently utilizes direct access to memory of peer GPU devices via the NV-Link. In addition to NV-Link, the hybrid algorithm employs PCI-Express network as extra bandwidth between GPU devices. To use the PCI-Express network we need to explicitly copy the data to host memory and these operations are software-pipelined. By selecting optimal parameters, the hybrid allreduce algorithm outperforms NVIDIA’s NCCL library. Akira Nukada |
IEEE BigData | 1 |
| 2019 | Batched Sparse Matrix Multiplication for Accelerating Graph Convolutional NetworksabstractGraph Convolutional Networks (GCNs) are recently getting much attention in bioinformatics and chemoinformatics as a state-of-the-art machine learning approach with high accuracy. GCNs process convolutional operations along with graph structures, and GPUs are used to process enormous operations including sparse-dense matrix multiplication (SpMM) when the graph structure is expressed as an adjacency matrix with sparse matrix format. However, the SpMM operation on small graph, where the number of nodes is tens or hundreds, hardly exploits high parallelism or compute power of GPU. Therefore, SpMM becomes a bottleneck of training and inference in GCNs applications. In order to improve the performance of GCNs applications, we propose new SpMM algorithm especially for small sparse matrix and Batched SpMM, which exploits high parallelism of GPU by processing multiple SpMM operations with single CUDA kernel. To the best of our knowledge, this is the first work of batched approach for SpMM. We evaluated the performance of the GCNs application on TSUBAME3.0 implementing NVIDIA Tesla P100 GPU, and our batched approach shows significant speedups of up to 1.59x and 1.37x in training and inference, respectively. Yusuke Nagasaka, Akira Nukada, Ryosuke Kojima, Satoshi Matsuoka |
CCGRID | 2 |
| 2018 | Optimizing Preconditioned Conjugate Gradient on TaihuLight for OpenFOAMabstractPorting the domain-specific software OpenFOAM onto the TaihuLight supercomputer is a challenging task, due to the highly memory-bound nature of both the supercomputer's processor (SW26010) and the software's liner solvers. Our study tackles this technical challenge, in three steps, by optimizing the linear solvers, such as Preconditioned Conjugate Gradient (PCG), on the SW26010. First, in order to minimize the all_reduce communication cost of PCG, we developed a new algorithm RNPCG, a non-blocking PCG leveraging the on-chip register communication. Second, we optimized three key kernels of the PCG, including proposing a localized version of the Diagonal-based Incomplete Cholesky (LDIC) preconditioner. Third, to scale the RNPCG on TaihuLight, we designed the three-level non-blocking all_reduce operations. With these three steps, we implemented the RNPCG in OpenFOAM. The experimental results on TaihuLight show that 1) compared with the default implementations of OpenFOAM, the RNPCG and the LDIC on a single-core group of SW26010 can achieve a maximum speedup of 8.9X and 3.1X, respectively; 2) the scalable RNPCG can outperform the standard PCG both in the strong and the weak scaling up to 66,560 cores. James Lin 0001, Minhua Wen, Delong Meng, Xin Liu 0020, Akira Nukada, Satoshi Matsuoka |
CCGrid | 5 |
| 2018 | Efficient Solving of Scan Primitive on Multi-GPU SystemsabstractGPUs fulfill high computation demands, but it is necessary to develop code carefully, selecting algorithms well suited to the GPU architecture and applying different optimizations. This article presents a GPU-suitable algorithm and a tuning strategy for performing the scan primitive over large problem sizes in CUDA. This tuning strategy defines different performance premises to find the GPU execution parameters that maximize performance. Taking these premises into consideration, we easily develop the kernels using CUDA skeletons to ensure efficiency and portability. Based on this, we describe an optimal proposal analyzed over different multiple GPU environments, the first multiple-GPU batch scan proposal to the best of our knowledge. The resulting implementations outperform other well-known libraries in most cases, such as CUDPP, ModernGPU, Thrust, CUB and LightScan. Adrián Pérez Diéguez, Margarita Amor, Ramón Doallo, Akira Nukada, Satoshi Matsuoka |
IPDPS | 4 |
| 2018 | Evaluating the SW26010 many-core processor with a micro-benchmark suite for performance optimizations
James Lin 0001, Zhigeng Xu, Linjin Cai, Akira Nukada, Satoshi Matsuoka |
Parallel Comput. | 4 |
| 2017 | Optimizations of Two Compute-Bound Scientific Kernels on the SW26010 Many-Core ProcessorabstractThe home-grown SW26010 many-core processor enabled the production of China's first independently developed number-one ranked supercomputer - the Sunway TaihuLight. The design of the limited off-chip memory bandwidth, however, renders the SW26010 a highly memory-bound processor. To compensate for this limitation, the processor was designed with a unique hardware feature, "Register Level Communication" (RLC), to share register data among its 8 × 8 computing processing elements (CPEs) via a 2D onchip network. Such a radical architecture has sparked global researchers' concerns regarding the programming challenges this may cause. To address these concerns, we adopted two compute-bound scientific kernels as benchmarks to identify the potential programming challenges. The first kernel is doubleprecision general matrix-multiplication (DGEMM). An RLCfriendly algorithm was designed for this kernel to reuse the data that already reside in the registers of 64 CPEs. This novel optimization enables the kernel to achieve up to 88.7% efficiency in one core group of the SW26010. This paper reveals, for the first time, the details of how the highly efficient DGEMM is implemented on the home-grown processor. The second kernel that we used is N-body. Due to the inefficient hardware support for transcendental operations on the SW26010, we replaced the reciprocal square root (rsqrt) instruction of N-body with a software routine to tackle the problem. Based on the programming challenges identified through these two optimized kernels, we proposed a three-level programming guideline for the SW26010. The paper concludes with our crucial finding that the critical step towards bridging the ninja performance gap on the SW26010 is to design an RLC-friendly algorithm to increase arithmetic intensity. James Lin 0001, Zhigeng Xu, Akira Nukada, Naoya Maruyama, Satoshi Matsuoka |
ICPP | 3 |
| 2017 | High-Performance and Memory-Saving Sparse General Matrix-Matrix Multiplication for NVIDIA Pascal GPUabstractSparse general matrix-matrix multiplication (SpGEMM) is one of the key kernels of preconditioners such as algebraic multigrid method or graph algorithms. However, the performance of SpGEMM is quite low on modern processors due to random memory access to both input and output matrices. As well as the number and the pattern of non-zero elements in the output matrix, important for achieving locality, are unknown before the execution. Moreover, the state-of-the-art GPU implementations of SpGEMM requires large amounts of memory for temporary results, limiting the matrix size computable on fast GPU device memory. We propose a new fast SpGEMM algorithm requiring small amount of memory and achieving high performance. Calculation of the pattern and value in output matrix is optimized by using GPU's on-chip shared memory and a hash table. Additionally, our algorithm launches multiple kernels running concurrently to improve the utilization of GPU resources. The kernels for the calculation of each row of output matrix are chosen based on the number of non-zero elements. Performance evaluation using matrices from the Sparse Matrix Collection of University Florida on NVIDIA's Pascal generation GPU shows that our approach achieves speedups of up to x4.3 in single precision and x4.4 in double precision compared to existing SpGEMM libraries. Furthermore, the memory usage is reduced by 14.7% in single precision and 10.9% in double precision on average, allowing larger matrices to be computed. Yusuke Nagasaka, Akira Nukada, Satoshi Matsuoka |
ICPP | 2 |
| 2015 | Modeling Gather and Scatter with Hardware Performance Counters for Xeon PhiabstractIntel Initial Many-Core Instructions (IMCI) for Xeon Phi introduces hardware-implemented Gather and Scatter (G/S) load/store contents of SIMD registers from/to non-contiguous memory locations. However, they can be one of key performance bottlenecks for Xeon Phi. Modelling G/S can provide insights to the performance on Xeon Phi, however, the existing solution needs a hand-written assembly implementation. Therefore, we modeled G/S with hardware performance counters which can be profiled by the tools like PAPI. We profiled Address Generation Interlock (AGI) events as the number of G/S, estimated the average latency of G/S with VPU_DATA_READ, and combined them to model the total latencies of G/S. We applied our model to the 3D 7-point stencil and the result showed G/S spent nearly 40% of total kernel time. We also validated the model by implementing a G/S- free version with intrinsics. The contribution of the work is a performance model for G/S built with hardware counters. We believe the model can be generally applicable to CPU as well. James Lin 0001, Akira Nukada, Satoshi Matsuoka |
CCGRID | 2 |
| 2015 | Efficient Execution of Multiple CUDA Applications Using Transparent Suspend, Resume and Migration
Taichiro Suzuki, Akira Nukada, Satoshi Matsuoka |
Euro-Par | 2 |
| 2014 | TSUBAME-KFC: A modern liquid submersion cooling prototype towards exascale becoming the greenest supercomputer in the worldabstractModern supercomputer performance is principally limited by power. TSUBAME-KFC is a state-of-the-art prototype for our next-generation TSUBAME3.0 supercomputer and towards future exascale. In collaboration with Green Revolution Cooling and others, TSUBAME-KFC submerges compute nodes configured with extremely high processor/component density, into non-toxic, low viscosity oil with high 260 Celsius flash point, and cooled using ambient / evaporative cooling tower. This minimizes cooling power while all semiconductor components kept at low temperature to lower leakage current. Numerous off-line in addition to on-line power and temperature sensors are facilitated throughout and constantly monitored to immediately observe the effect of voltage/frequency control. As a result, TSUBAME-KFC achieved world No.1 on the Green500 in Nov. 2013 and Jun. 2014, by over 20% c.f. the nearest competitors. Toshio Endo, Akira Nukada, Satoshi Matsuoka |
ICPADS | 2 |
| 2014 | Cache-aware sparse matrix formats for Kepler GPUabstractScientific simulations often require solving extremely large sparse linear equations, whose dominant kernel is sparse matrix vector multiplication. On modern many-core processors such as GPU or MIC, the operation has been known to pose significant bottleneck and thus would result in extremely poor efficiency, because of limited processor-to-memory bandwidth and low cache hit ratio due to random access to the input vector. Our family of new sparse matrix formats for many-core processors significantly increases the cache hit ratio and thus performance by segmenting the matrix along the columns, dividing the work among the many core up to the internal cache capacity, and aggregating the result later on. Performance studies show that we achieve up to x3.0 speedup in SpMV and x1.68 in multi-node CG, compared to the best vendor libraries and competing new formats that have been recently proposed such as SELL-C-σ. Yusuke Nagasaka, Akira Nukada, Satoshi Matsuoka |
ICPADS | 2 |
| 2012 | Scalable multi-GPU 3-D FFT for TSUBAME 2.0 supercomputerabstractFor scalable 3-D FFT computation using multiple GPUs, efficient all-to-all communication between GPUs is the most important factor in good performance. Implementations with point-to-point MPI library functions and CUDA memory copy APIs typically exhibit very large overheads especially for small message sizes in all-to-all communications between many nodes. We propose several schemes to minimize the overheads, including employment of lower-level API of InfiniBand to effectively overlap intra- and inter-node communication, as well as auto-tuning strategies to control scheduling and determine rail assignments. As a result we achieve very good strong scalability as well as good performance, up to 4.8TFLOPS using 256 nodes of TSUBAME 2.0 Supercomputer (768 GPUs) in double precision. Akira Nukada, Kento Sato, Satoshi Matsuoka |
SC | 1 |
| 2011 | Hamming Color Code for Dense and Robust One-shot 3D ScanningabstractWe propose a novel color code, Hamming color code, designed for rapid 3D shape acquisition using structured light projection. The Hamming color code has several properties which are desirable for practical 3D acquisition as follows. First, the Hamming distance of adjacent colors is always 1, which makes the color detection robust to color blending due to defocusing, subsurface scattering, or chromatic aberration. Second, the substrings of a certain length is guaranteed to be unique. In other words, the Hamming code can be viewed as a subset of de Bruijn sequence. Third, a one-dimensional coordinate can be encoded for each pixel, which enables dense 3D reconstruction from a single pattern projection. Thanks to the uniqueness and robustness of the substrings, the structured light can be decoded stably by dynamic programming. We have implemented parallel dynamic programming on GPU and achieved the speed-up by a factor of 630 compared to the CPU-based implementation, and accomplished video-rate 3D acquisition using commodity hardware. Several experiments have been conducted to demonstrate the stability and performance of our algorithm. Finally we discuss the limitation and future direction of this work. Shuntaro Yamazaki, Akira Nukada, Masaaki Mochimaru |
BMVC | 2 |
| 2011 | Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputerabstractThe mechanical properties of metal materials largely depend on their intrinsic internal microstructures. To develop engineering materials with the expected properties, predicting patterns in solidified metals would be indispensable. The phase-field simulation is the most powerful method known to simulate the micro-scale dendritic growth during solidification in a binary alloy. To evaluate the realistic description of solidification, however, phase-field simulation requires computing a large number of complex nonlinear terms over a fine-grained grid. Due to such heavy computational demand, previous work on simulating three-dimensional solidification with phase-field methods was successful only in describing simple shapes. Our new simulation techniques achieved scales unprecedentedly large, sufficient for handling complex dendritic structures required in material science. Our simulations on the GPU-rich TSUBAME 2.0 supercomputer at the Tokyo Institute of Technology have demonstrated good weak scaling and achieved 1.017 PFlops in single precision for our largest configuration, using 4,000 GPUs along with 16,000 CPU cores. Takashi Shimokawabe, Takayuki Aoki, Tomohiro Takaki, Toshio Endo, Akinori Yamanaka, Naoya Maruyama, Akira Nukada, Satoshi Matsuoka |
SC | 7 |
| 2010 | Low-overhead diskless checkpoint for hybrid computing systemsabstractAs the size of new supercomputers scales to tens of thousands of sockets, the mean time between failures (MTBF) is decreasing to just several hours and long executions need some kind of fault tolerance method to survive failures. Checkpoint\Restart is a popular technique used for this purpose; but writing the state of a big scientific application to remote storage will become prohibitively expensive in the near future. Diskless checkpoint was proposed as a solution to avoid the I/O bottleneck of disk-based checkpoint. However, the complex time-consuming encoding techniques hinder its scalability. At the same time, heterogeneous computing is becoming more and more popular in high performance computing (HPC), with new clusters combining CPUs and graphic processing units (GPUs). However, hybrid applications cannot always use all the resources available on the nodes, leaving some idle resources suc h us GPUs or CPU cores. In this work, we propose a hybrid diskless checkpoint (HDC) technique for GPU-accelerated clusters, that can checkpoint CPU/GPU applications, does not require spare nodes and can tolerate up to 50% of process failures with a low, sometimes negligible, checkpoint overhead. Leonardo Arturo Bautista-Gomez, Akira Nukada, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
HiPC | 2 |
| 2010 | Linpack evaluation on a supercomputer with heterogeneous acceleratorsabstractWe report Linpack benchmark results on the TSUBAME supercomputer, a large scale heterogeneous system equipped with NVIDIA Tesla GPUs and ClearSpeed SIMD accelerators. With all of 10,480 Opteron cores, 640 Xeon cores, 648 ClearSpeed accelerators and 624 NVIDIA Tesla GPUs, we have achieved 87.01TFlops, which is the third record as a heterogeneous system in the world. This paper describes careful tuning and load balancing method required to achieve this performance. On the other hand, since the peak speed is 163 TFlops, the efficiency is 53%, which is lower than other systems. This paper also analyses this gap from the aspect of system architecture. Toshio Endo, Akira Nukada, Satoshi Matsuoka, Naoya Maruyama |
IPDPS | 2 |
| 2010 | A high-performance fault-tolerant software framework for memory on commodity GPUsabstractAs GPUs are increasingly used to accelerate HPC applications by allowing more flexibility and programmability, their fault tolerance is becoming much more important than before when they were used only for graphics. The current generation of GPUs, however, does not have standard error detection and correction capabilities, such as SEC-DED ECC for DRAM, which is almost always exercised in HPC servers. We present a high-performance software framework to enhance commodity off-the-shelf GPUs with DRAM fault tolerance. It combines data coding for detecting bit-flip errors and checkpointing for recovering computations when such errors are detected. We analyze performance of data coding in GPUs and present optimizations geared toward memory-intensive GPU applications. We present performance studies of the prototype implementation of the framework and show that the proposed framework can be realized with negligible overheads in compute intensive applications such as N-body problem and matrix multiplication, and as low as 35% in a highly-efficient memory intensive 3-D FFT kernel. Naoya Maruyama, Akira Nukada, Satoshi Matsuoka |
IPDPS | 2 |
| 2010 | An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production CodeabstractRegional weather forecasting demands fast simulation over fine-grained grids, resulting in extremely memory- bottlenecked computation, a difficult problem on conventional supercomputers. Early work on accelerating mainstream weather code WRF using GPUs with their high memory performance, however, resulted in only minor speedup due to partial GPU porting of the huge code. Our full CUDA porting of the high- resolution weather prediction model ASUCA is the first such one we know to date; ASUCA is a next-generation, production weather code developed by the Japan Meteorological Agency, similar to WRF in the underlying physics (non-hydrostatic model). Benchmark on the 528 (NVIDIA GT200 Tesla) GPU TSUBAME Supercomputer at the Tokyo Institute of Technology demonstrated over 80-fold speedup and good weak scaling achieving 15.0 TFlops in single precision for 6956 x 6052 x 48 mesh. Further benchmarks on TSUBAME 2.0, which will embody over 4000 NVIDIA Fermi GPUs and deployed in October 2010, will be presented. Takashi Shimokawabe, Takayuki Aoki, Chiashi Muroi, Junichi Ishida, Kohei Kawano, Toshio Endo, Akira Nukada, Naoya Maruyama, Satoshi Matsuoka |
SC | 7 |
| 2009 | Aspects of GPU for general purpose high performance computingabstractWe discuss hardware and software aspects of GPGPU, specifically focusing on NVIDIA cards and CUDA, from the viewpoints of parallel computing. The major weak points of GPU against newest supercomputers are identified to be and summarized as only four points: large SIMD vector length, small memory, absence of fast L2 cache, and high register spill penalty. As software concerns, we derive optimal scheduling algorithm for latency hiding of host-device data transfer, and discuss SPMD parallelism on GPUs. Reiji Suda, Takayuki Aoki, Shoichi Hirasawa, Akira Nukada, Hiroki Honda, Satoshi Matsuoka |
ASP-DAC | 4 |
| 2009 | Auto-tuning 3-D FFT library for CUDA GPUsabstractExisting implementations of FFTs on GPUs are optimized for specific transform sizes like powers of two, and exhibit unstable and peaky performance i.e., do not perform as well in other sizes that appear in practice. Our new auto-tuning 3-D FFT on CUDA generates high performance CUDA kernels for FFTs of varying transform sizes, alleviating this problem. Although auto-tuning has been implemented on GPUs for dense kernels such as DGEMM and stencils, this is the first instance that has been applied comprehensively to bandwidth intensive and complex kernels such as 3-D FFTs. Bandwidth intensive optimizations such as selecting the number of threads and inserting padding to avoid bank conflicts on shared memory are systematically applied. Our resulting autotuner is fast and results in performance that essentially beats all 3-D FFT implementations on a single processor to date, and moreover exhibits stable performance irrespective of problem sizes or the underlying GPU hardware. Akira Nukada, Satoshi Matsuoka |
SC | 1 |
| 2008 | Bandwidth intensive 3-D FFT kernel for GPUs using CUDAabstractMost GPU performance “hypes” have focused around tightly-coupled applications with small memory bandwidth requirements e.g., N-body, but GPUs are also commodity vector machines sporting substantial memory bandwidth; however, effective programming methodologies thereof have been poorly studied. Our new 3-D FFT kernel, written in NVIDIA CUDA, achieves nearly 80 GFLOPS on a top-end GPU, being more than three times faster than any existing FFT implementations on GPUs including CUFFT. Careful programming techniques are employed to fully exploit modern GPU hardware characteristics while overcoming their limitations, including on-chip shared memory utilization, optimizing the number of threads and registers through appropriate localization, and avoiding low-speed stride memory accesses. Our kernel applied to real applications achieves orders of magnitude boost in power&cost vs. performance metrics. The off-card bandwidth limitation is still an issue, which could be alleviated somewhat with application kernels confinement within the card, while ideal solution being facilitation of faster GPU interfaces. Akira Nukada, Yasuhiko Ogata, Toshio Endo, Satoshi Matsuoka |
SC | 1 |
| 2007 | High Performance FFT on SGI Altix 3700
Akira Nukada, Daisuke Takahashi, Reiji Suda, Akira Nishida |
HPCC | 1 |
| 2007 | High Performance 3D Convolution for Protein Docking on IBM Blue Gene
Akira Nukada, Yuichiro Hourai, Akira Nishida, Yutaka Akiyama |
ISPA | 1 |
| 2006 | FFTSS: A High Performance Fast Fourier Transform LibraryabstractIn this paper, we introduce a new fast Fourier transform (FFT) library. In developing this software, we focus on the efficient execution of the floating-point operation instructions. To achieve high performance on various processors, we provide the source code which compilers can optimize easily. Since the compilers provided by processor vendors have powerful optimizers for loop sentences, the code generated by them will run very fast as long as the iteration count of the innermost loop is large enough. In such a case, the library outperforms other libraries even provided by processor vendors Akira Nukada |
ICASSP (3) | 1 |
| 2006 | Poster reception - Scalable software infrastructure projectabstractRecent progress of science and technology has made numerical simulation an important approach for studies in various fields. Although scalable and high performance numerical libraries on large scale computing resources are indispensable tools for handling various multiscale phenomena, few projects for integrating these numerical libraries have been reported. The object of this project is the development of a basic library of solutions and algorithms required for large scale scientific simulations, which have been developed separately in each fields, and its integration into a scalable software infrastructure. The components include a scalable iterative solvers library Lis, having a number of solvers, preconditioners, and matrix storage formats that are flexibly combinable, a fast Fourier transform library FFTSS for various superscalar architectures with SIMD instructions, which outperforms some vendor-provided FFT libraries, and a language- and computing environment-independent matrix computation framework SILC. We show some highlights of our achievements on leading high performance computers. Akira Nishida, Hisashi Kotakemori, Tamito Kajiyama, Akira Nukada |
SC | 4 |