EDBT 2026 Demo / reviewers in the wild / expert
Takayuki Aoki
dblp:05/3528
· DBLP profile ↗
15ranked-venue papers
0as first author
3since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 63% GPUs and heterogeneous computing · 33% Performance modeling and evaluation · 4% |
Topics — the 5 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
scientific computing systems |
0.4 | 3 | 2014 | High-Productivity Framework on GPU-Rich Supercomputers for Operational Weather Prediction Code ASUCA · SC 2014 Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputer · SC 2011 An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production Code · SC 2010 |
High-performance computing › scientific computing systems
weather prediction |
0.3 | 2 | 2014 | High-Productivity Framework on GPU-Rich Supercomputers for Operational Weather Prediction Code ASUCA · SC 2014 An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production Code · SC 2010 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 2 | 2011 | Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputer · SC 2011 An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production Code · SC 2010 |
High-performance computing
stencil computation |
0.2 | 1 | 2014 | High-Productivity Framework on GPU-Rich Supercomputers for Operational Weather Prediction Code ASUCA · SC 2014 |
High-performance computing › scientific computing systems
phase field simulation |
0.1 | 1 | 2011 | Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputer · SC 2011 |
Methods — techniques the papers use, named apart from their topics
empirical analysis · 0.2compiler transformation · 0.2peer-to-peer direct access · 0.2MPI · 0.2phase-field method · 0.1CUDA · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Multi-GPU Scaling of a Conservative Weakly Compressible Solver for Large-Scale Two-Phase Flow Simulation
Takayuki Aoki |
PDCAT | 2 |
| 2021 | GPU Acceleration of Multigrid Preconditioned Conjugate Gradient Solver on Block-Structured Cartesian GridabstractWe develop a multigrid preconditioned conjugate gradient (MG-CG) solver for the pressure Poisson equation in a two-phase flow CFD code JUPITER. The JUPITER code is redesigned to realize efficient CFD simulations including complex boundaries and objects based on a block-structured Cartesian grid system. The code is written in CUDA, and is tuned to achieve high performance on GPU based supercomputers. The main kernels of the MG-CG solver achieve more than 90% of the roofline performance. The MG preconditioner is constructed based on the geometric MG method with a three-stage V-cycle, and a red-black SOR (RB-SOR) smoother and its variant with cache-reuse optimization (CR-SOR) are applied at each stage. The numerical experiments are conducted for two-phase flows in a fuel bundle of a nuclear reactor. Thanks to the block-structured data format, grids inside fuel pins are removed without performance degradation, and the total number of grids is reduced to 2.26 × 109, which is about 70% of the original Cartesian grid. The MG-CG solvers with the RB-SOR and CR-SOR smoothers reduce the number of iterations to less than 15% and 9% of the original preconditioned CG method, leading to 3.1- and 5.9-times speedups, respectively. In the strong scaling test, the MG-CG solver with the CR-SOR smoother is accelerated by 2.1 times between 64 and 256 GPUs. The obtained performance indicates that the MG-CG solver designed for the block-structured grid is highly efficient and enables large-scale simulations of two-phase flows on GPU based supercomputers. Naoyuki Onodera, Yasuhiro Idomura, Yuta Hasegawa, Susumu Yamashita, Takashi Shimokawabe, Takayuki Aoki |
HPC Asia | 6 |
| 2021 | Tree cutting approach for domain partitioning on forest-of-octrees-based block-structured static adaptive mesh refinement with lattice Boltzmann methodabstractThe aerodynamics simulation code based on the lattice Boltzmann method (LBM) using forest-of-octrees-based block-structured adaptive mesh refinement (AMR) with temporary-fixed refinement was implemented, and its performance was evaluated on GPU-based supercomputers. Although the Space-Filling-Curve-based (SFC) domain partitioning algorithm for the octree-based AMR has been widely used on conventional CPU-based supercomputers, accelerated computation on GPU-based supercomputers revealed a bottleneck due to costly halo data communication. Our new tree cutting approach adopts a hybrid domain partitioning with the coarse structured block decomposition and the SFC partitioning in each block. This hybrid approach improved the locality and the topology of the partitioned sub-domains and reduced the amount of the halo communication to one-third of the original SFC approach. In the strong scaling test, the code achieved maximum ×1.82 speedup at the performance of 2207 MLUPS (mega-lattice update per second) on 128 GPUs (NVIDIA® Tesla® V100). In the weak scaling test, the code achieved 9620 MLUPS at 128 GPUs with 4.473 billion grid points, while keeping the parallel efficiency of 93.4% from 8 to 128 GPUs. Yuta Hasegawa, Takayuki Aoki, Hiromichi Kobayashi, Yasuhiro Idomura, Naoyuki Onodera |
Parallel Comput. | 2 |
| 2020 | A domain partitioning method using a multi-phase-field model for block-based AMR applications
Seiya Watanabe, Takayuki Aoki, Tomohiro Takaki |
Parallel Comput. | 2 |
| 2017 | A Stencil Framework to Realize Large-Scale Computations Beyond Device Memory Capacity on GPU SupercomputersabstractStencil-based applications such as CFD have succeeded in obtaining high performance on GPU supercomputers. The problem sizes of these applications are limited by the GPU device memory capacity, which is typically smaller than the host memory. On GPU supercomputers, a locality improvement technique using temporal blocking method with memory swapping between host and device enables large computation beyond the device memory capacity. However, because the loop management of temporal blocking with data movement across these memories increase programming difficulty, the applying this methodology to the real stencil applications demands substantially higher programming cost. Our high-productivity stencil framework automatically applies temporal blocking to boundary exchange required for stencil computation and supports automatic memory swapping provided by a MPI/CUDA wrapper library. The framework-based application for the airflow in an urban city maintains 80% performance even with the twice larger than the GPU memory capacity and have demonstrated good weak scalability on the TSUBAME 2.5 supercomputer. Takashi Shimokawabe, Toshio Endo, Naoyuki Onodera, Takayuki Aoki |
CLUSTER | 4 |
| 2016 | Daino: a high-level framework for parallel and efficient AMR on GPUsabstractAdaptive Mesh Refinement methods reduce computational requirements of problems by increasing resolution for only areas of interest. However, in practice, efficient AMR implementations are difficult considering that the mesh hierarchy management must be optimized for the underlying hardware. Architecture complexity of GPUs can render efficient AMR to be particularity challenging in GPU-accelerated supercomputers. This paper presents a compiler-based high-level framework that can automatically transform serial uniform mesh code annotated by the user into parallel adaptive mesh code optimized for GPU-accelerated supercomputers. We also present a method for empirical analysis of a uniform mesh to project an upper-bound on achievable speedup of a GPU-optimized AMR code. We show experimental results on three production applications. The speedups of code generated by our framework are comparable to hand-written AMR code while achieving good and weak scaling up to 1000 GPUs. Mohamed Wahib, Naoya Maruyama, Takayuki Aoki |
SC | 3 |
| 2015 | Performance modeling and analysis of heterogeneous lattice Boltzmann simulations on CPU-GPU clusters
Christian Feichtinger, Johannes Habich, Harald Köstler, Ulrich Rüde, Takayuki Aoki |
Parallel Comput. | 5 |
| 2014 | High-Productivity Framework on GPU-Rich Supercomputers for Operational Weather Prediction Code ASUCAabstractThe weather prediction code demands large computational performance to achieve fast and high-resolution simulations. Skillful programming techniques are required for obtaining good parallel efficiency on GPU supercomputers. Our framework-based weather prediction code ASUCA has achieved good scalability with hiding complicated implementation and optimizations required for distributed GPUs, contributing to increasing the maintainability, ASUCA is a next-generation high resolution meso-scale atmospheric model being developed by the Japan Meteorological Agency. Our framework automatically translates user-written stencil functions that update grid points and generates both GPU and CPU codes. User-written codes are parallelized by MPI with intra-node GPU peer-to-peer direct access. These codes can easily utilize optimizations such as overlapping technique to hide communication overhead by computation. Our simulations on the GPU-rich supercomputer TSUBAME 2.5 at the Tokyo Institute of Technology have demonstrated good strong and weak scalability achieving 209.6 TFlops in single precision for our largest model using 4,108 NVIDIA K20X GPUs. Takashi Shimokawabe, Takayuki Aoki, Naoyuki Onodera |
SC | 2 |
| 2011 | Peta-scale phase-field simulation for dendritic solidification on the TSUBAME 2.0 supercomputerabstractThe mechanical properties of metal materials largely depend on their intrinsic internal microstructures. To develop engineering materials with the expected properties, predicting patterns in solidified metals would be indispensable. The phase-field simulation is the most powerful method known to simulate the micro-scale dendritic growth during solidification in a binary alloy. To evaluate the realistic description of solidification, however, phase-field simulation requires computing a large number of complex nonlinear terms over a fine-grained grid. Due to such heavy computational demand, previous work on simulating three-dimensional solidification with phase-field methods was successful only in describing simple shapes. Our new simulation techniques achieved scales unprecedentedly large, sufficient for handling complex dendritic structures required in material science. Our simulations on the GPU-rich TSUBAME 2.0 supercomputer at the Tokyo Institute of Technology have demonstrated good weak scaling and achieved 1.017 PFlops in single precision for our largest configuration, using 4,000 GPUs along with 16,000 CPU cores. Takashi Shimokawabe, Takayuki Aoki, Tomohiro Takaki, Toshio Endo, Akinori Yamanaka, Naoya Maruyama, Akira Nukada, Satoshi Matsuoka |
SC | 2 |
| 2011 | Multi-GPU performance of incompressible flow computation by lattice Boltzmann method on GPU cluster
Takayuki Aoki |
Parallel Comput. | 2 |
| 2010 | An 80-Fold Speedup, 15.0 TFlops Full GPU Acceleration of Non-Hydrostatic Weather Model ASUCA Production CodeabstractRegional weather forecasting demands fast simulation over fine-grained grids, resulting in extremely memory- bottlenecked computation, a difficult problem on conventional supercomputers. Early work on accelerating mainstream weather code WRF using GPUs with their high memory performance, however, resulted in only minor speedup due to partial GPU porting of the huge code. Our full CUDA porting of the high- resolution weather prediction model ASUCA is the first such one we know to date; ASUCA is a next-generation, production weather code developed by the Japan Meteorological Agency, similar to WRF in the underlying physics (non-hydrostatic model). Benchmark on the 528 (NVIDIA GT200 Tesla) GPU TSUBAME Supercomputer at the Tokyo Institute of Technology demonstrated over 80-fold speedup and good weak scaling achieving 15.0 TFlops in single precision for 6956 x 6052 x 48 mesh. Further benchmarks on TSUBAME 2.0, which will embody over 4000 NVIDIA Fermi GPUs and deployed in October 2010, will be presented. Takashi Shimokawabe, Takayuki Aoki, Chiashi Muroi, Junichi Ishida, Kohei Kawano, Toshio Endo, Akira Nukada, Naoya Maruyama, Satoshi Matsuoka |
SC | 2 |
| 2009 | Aspects of GPU for general purpose high performance computingabstractWe discuss hardware and software aspects of GPGPU, specifically focusing on NVIDIA cards and CUDA, from the viewpoints of parallel computing. The major weak points of GPU against newest supercomputers are identified to be and summarized as only four points: large SIMD vector length, small memory, absence of fast L2 cache, and high register spill penalty. As software concerns, we derive optimal scheduling algorithm for latency hiding of host-device data transfer, and discuss SPMD parallelism on GPUs. Reiji Suda, Takayuki Aoki, Shoichi Hirasawa, Akira Nukada, Hiroki Honda, Satoshi Matsuoka |
ASP-DAC | 2 |
| 2008 | Evaluating power and energy consumption of FPGA-based custom computing machines for scientific floating-point computationabstractThis paper evaluates the actual power consumption and the total energy for scientific floating-point computations accelerated by FPGA-based custom computing machines. With our FPGA-based machines: the streaming accelerator for computational fluid dynamics and the programmable systolic-array processor for numerical simulations based on difference schemes, we measure the power of the entire systems including a host PC and an FPGA board, and obtain the total energy for each computation. We report that the FPGAs perform the same computation with 5% to 30% of the total energy consumed by a microprocessor, while the FPGAs accelerate the computation. Kentaro Sano, Takeshi Nishikawa, Takayuki Aoki, Satoru Yamamoto |
FPT | 3 |
| 2003 | Implementation Methods of Class Based Queueing with Dynamic Bandwidth Decision Method for Network ProcessorsabstractDiffServ (differentiated services) is proposed to define the priority level in an IP header. For the priority controls, CBQ (class based queueing) is also presented. In CBQ, a router prepares a queue for each class of the priority level, and assigns available bandwidth to the class. However, the upper limit of the available bandwidth for each class is fixed in CBQ, and any class usually, cannot obtain extra bandwidth beyond the upper limit. A lower priority class can borrow a part of bandwidth from a higher priority class, only when the average packet transfer rate in the higher priority class must be smaller than the expected rate defined by a certain threshold. Thus, if the higher priority class constantly has data to be transferred, the lower priority class cannot borrow any bandwidth, even if the higher priority class has enough queue space. From this problem, we introduced the dynamic bandwidth decision mechanism to CBQ. In the mechanism, the upper limit of the available bandwidth for each class is dynamically modified due to the size of the queue space in each class. The proposal method provides more flexible bandwidth control, and then increases the effectiveness of the whole network from the simulation experiments. This paper presents implementation methods of the dynamic bandwidth decision method for network processors. Finally, the methods are evaluated on the real network environments. Shigetomo Kimura, Takayuki Aoki, Hideho Gomi, Yoshihiko Ebihara |
AINA | 2 |
| 2002 | First Light of the Earth Simulator and Its PC Cluster ApplicationsabstractThe Earth Simulator (ES) is the largest parallel vector processor in the world that is mainly dedicated to large-scale simulation studies of global change. Development of the ES system started in 1997 and was completed at the end of February, 2002. The system consists of 640 processor nodes that are connected via a very fast single-stage crossbar network (12.3 GB/s). The total peak performance and main memory of the system are 40 TFLOPS and 10 TB, respectively. Studies to evaluate the performance of the ES were made using an atmospheric circulation model Afes (Atmospheric General Circulation Model for ES) and LINPACK benchmark test. The sustained performance of Afes for T1279L96 (the equivalent horizontal resolution given by T1279 is about 10 km and the total number of layers is 96) was as high as 14.5 TFLOPS on a half system of the ES with 2,560 PEs (320 nodes). The sustained-to-peak performance ratio was 70.8%. The ES also achieved a LINPACK world record of 35.86 TFLOPS. This rating exceeded the previous record, set by the ASCI White, by about 5 times. The Earth Simulator is now running. Huge amounts of output data will arise from the huge computer system. For example, the data volume of simulation results from the Afes is of the order of 10-100 TB. In the phase of operation, management of huge output datafiles and interactive visual monitoring of many terabytes of simulation results are extremely important for the ES. The ES has introduced a prototype PC cluster to seek the best solution to these problems. The PC cluster comprises 64 PCs that are interconnected with a Myrinet2000 switch. Each PC has a Pentium III (1 GHz), 1 GB of main memory and 120 GB of disk space. An outline of the Earth Simulator system, recent results on performance evaluation using real applications and the LINPACK benchmark test, and an outline of the PC cluster system are presented. Keiji Tani, Takayuki Aoki, Satoshi Matsuoka, Satoru Ohkura, Hitoshi Uehara, Tetsuo Aoyagi |
CLUSTER | 2 |