Tetsu Narumi

dblp:64/4590 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
High-performance computing · 56% Performance modeling and evaluation · 18% GPUs and heterogeneous computing · 18%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Computational science and engineering · 69% Bioinformatics and computational biology · 31%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
0.242006
Gordon Bell finalists II - A 55 TFLOPS simulation of amyloid-forming peptides from yeast prion Sup35 with the special-purpose computer system MDGRAPE-3 · SC 2006
Protein Explorer: A Petaflops Special-Purpose Computer System for Molecular Dynamics Simulations · SC 2003
An 8.61 Tflop/s molecular dynamics simulation for NaCl with a special-purpose computer: MDM · SC 2001
Computational science and engineering › computational chemistry › molecular simulation
molecular dynamics
0.142006
Gordon Bell finalists II - A 55 TFLOPS simulation of amyloid-forming peptides from yeast prion Sup35 with the special-purpose computer system MDGRAPE-3 · SC 2006
An 8.61 Tflop/s molecular dynamics simulation for NaCl with a special-purpose computer: MDM · SC 2001
1.34 Tflops Molecular Dynamics Simulation for NaCl with a Special-Purpose Computer: MDM · SC 2000
Performance modeling and evaluation › numerical algorithms
fast multipole method
0.112009
42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster
0.112009
42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009
High-performance computing › n-body simulation
hierarchical n-body methods
0.112009
42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009
Bioinformatics and computational biology › structural bioinformatics
protein structure determination
0.112007
A 281 Tflops calculation for X-ray protein structure analysis with special-purpose computers MDGRAPE-3 · SC 2007
Hardware accelerators and domain-specific architectures › scientific computing accelerator
molecular dynamics accelerator
0.012003
Protein Explorer: A Petaflops Special-Purpose Computer System for Molecular Dynamics Simulations · SC 2003
High-performance computing › scientific computing systems
molecular dynamics simulation
0.012003
Protein Explorer: A Petaflops Special-Purpose Computer System for Molecular Dynamics Simulations · SC 2003
Computational science and engineering › computational fluid dynamics
turbulence simulation
0.012009
42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009
Mathematical optimization › evolutionary computation
genetic algorithm
0.012007
A 281 Tflops calculation for X-ray protein structure analysis with special-purpose computers MDGRAPE-3 · SC 2007

Methods — techniques the papers use, named apart from their topics

nonequispaced discrete fourier transformation · 0.2genetic algorithm · 0.2treecode · 0.2fast multipole method · 0.2ewald summation · 0.1special-purpose LSI · 0.1MDGRAPE-3 · 0.1
YearPublicationVenuePosition
2021 CUDA offloading for energy-efficient and high-frame-rate simulations using tablets
abstract
Summary The multiple sensors and touch capabilities of mobile devices are defining new methods of computer interaction. However, the computing power of such devices is not currently sufficient for new applications that require compute‐intensive applications. Using graphics processing units (GPUs) for general‐purpose computing with GPU programming models such as Compute Unified Device Architecture (CUDA) has been proved to accelerate simulations in supercomputers. Although, CUDA‐capable chips such as the Tegra K1 have been released on tablets can accelerate computer simulations, their absolute computing power and performance per watt are not comparable with ordinary GPUs. In this paper, we analyze a heterogeneous system composed of both of a tablet (client) and notebook with a low‐power GPU (server). Intensive computations on a tablet device are offloaded to a notebook GPU using the rCUDA middleware. Molecular dynamics (MD) simulations are performed using our test system, and the computing speed and performance per watt are reported. Implementing dynamic parallelism (DP) reduced the latency, doubling the total frames per second in some cases. Our system achieves better computational performance, and higher performance per watt than a tablet powered by a CUDA‐capable GPU. We achieved 21.7 Gflops/W by combining multiple client tablets and server, compared with 21.3 Gflops/W from the server itself.
Edgar Josafat Martinez-Noriega, Syunji Yazaki, Tetsu Narumi
Concurr. Comput. Pract. Exp.3
2011 Fast Calculation of Electrostatic Potentials on the GPU or the ASIC MD-GRAPE-3
abstract
Electrostatic potentials (ESPs) are frequently used in structural biology for the characterization of biomolecules. Here we study the potential employment of hardware accelerators like the graphics processing unit or the application-specific integrated circuit MD-GRAPE-3 for the purpose of efficient computation of ESPs. An algorithm closely coupled to the general description of molecular surfaces is ported to both specialized architectures. The high-level interface library MR1/3 is used, which greatly simplifies the porting process. Hardware-accelerated versions show significant Speed-Up factors reaching values of up to 27×. Once ESP computations have become a matter of seconds, the underlying application can be offered in the form of a web service. © The Author 2009. Published by Oxford University Press on behalf of The British Computer Society. All rights reserved.
Tetsu Narumi, Kenji Yasuoka, Makoto Taiji, Francesco Zerbetto, Siegfried Höfinger
Comput. J.1
2009 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence
abstract
As an entry for the 2009 Gordon Bell price/performance prize, we present the results of two different hierarchical N-body simulations on a cluster of 256 graphics processing units (GPUs). Unlike many previous N-body simulations on GPUs that scale as O(N2), the present method calculates the O(N log N) treecode and O(N) fast multipole method (FMM) on the GPUs with unprecedented efficiency. We demonstrate the performance of our method by choosing one standard application --a gravitational N-body simulation-- and one non-standard application --simulation of turbulence using vortex particles. The gravitational simulation using the treecode with 1,608,044,129 particles showed a sustained performance of 42.15 TFlops. The vortex particle simulation of homogeneous isotropic turbulence using the periodic FMM with 16,777,216 particles showed a sustained performance of 20.2 TFlops. The overall cost of the hardware was 228,912 dollars. The maximum corrected performance is 28.1TFlops for the gravitational simulation, which results in a cost performance of 124 MFlops/$. This correction is performed by counting the Flops based on the most efficient CPU algorithm. Any extra Flops that arise from the GPU implementation and parameter differences are not included in the 124 MFlops/$.
Tsuyoshi Hamada, Tetsu Narumi, Rio Yokota, Kenji Yasuoka, Keigo Nitadori, Makoto Taiji
SC2
2009 High-Performance Drug Discovery: Computational Screening by Combining Docking and Molecular Dynamics Simulations
abstract
Virtual compound screening using molecular docking is widely used in the discovery of new lead compounds for drug design. However, this method is not completely reliable and therefore unsatisfactory. In this study, we used massive molecular dynamics simulations of protein-ligand conformations obtained by molecular docking in order to improve the enrichment performance of molecular docking. Our screening approach employed the molecular mechanics/Poisson-Boltzmann and surface area method to estimate the binding free energies. For the top-ranking 1,000 compounds obtained by docking to a target protein, approximately 6,000 molecular dynamics simulations were performed using multiple docking poses in about a week. As a result, the enrichment performance of the top 100 compounds by our approach was improved by 1.6-4.0 times that of the enrichment performance of molecular dockings. This result indicates that the application of molecular dynamics simulations to virtual screening for lead discovery is both effective and practical. However, further optimization of the computational protocols is required for screening various target proteins.
Noriaki Okimoto, Noriyuki Futatsugi, Hideyoshi Fuji, Atsushi Suenaga, Gentaro Morimoto, Ryoko Yanai, Yousuke Ohno, Tetsu Narumi, Makoto Taiji
PLoS Comput. Biol.8
2008 Overheads in Accelerating Molecular Dynamics Simulations with GPUs
abstract
Molecular Dynamics (MD) simulation requires huge computational power, as each atom interacts with the others by long range forces such as the Coulomb or van der Waals forces. Recently, a video game computer, such as SONY PLAYSTATION 3 (PS3) or NVIDIApsilas Graphics Processing Unit (GPU) has become a candidate hardware for accelerating MD simulations as well as an MDGRAPE-3 special-purpose computer for their better performance than current CPU of the PC, and also for their cost-effectiveness. Especially the latest GPU has much more peak performance than a CPU of the PC or an MDGRAPE-3, though a GPU has much more overheads in accelerating MD simulations. When the number of particles is small or the calculation kernel becomes complicated, the performance of the GPU drops dramatically as low as that of the MDGRAPE-3. However, the acceleration ratio of the GPU and the PS3 per cost exceeds that of the MDGRAPE-3.
Tetsu Narumi, Ryuji Sakamaki, Shun Kameoka, Kenji Yasuoka
PDCAT1
2007 A 281 Tflops calculation for X-ray protein structure analysis with special-purpose computers MDGRAPE-3
abstract
We have achieved a sustained calculation speed of 281 Tflops for the optimization of the 3-D structures of proteins from the X-ray experimental data by the Genetic Algorithm - Direct Space (GA-DS) method. In this calculation we used MDGRAPE-3, special-purpose computer for molecular simulations, with the peak performance of 752 Tflops. In the GA-DS method, a set of selected parameters which define the crystal structures of proteins is optimized by the Genetic Algorithm. As a criterion to estimate the model parameters, we used the reliability factor R1 which indicates the statistical difference between the calculated and the measured diffraction data. To evaluate this factor it is necessary to reconstruct the diffraction patterns of the model structures every time the model is updated. Therefore, in this method the nonequispaced Discrete Fourier Transformation (DFT) used to calculate the diffraction patterns dominates most of the computation time. To accelerate DFT calculations, we used the special-purpose computer, MDGRAPE-3. A molecule, Carbamoyl-Phosphate Synthetase was investigated. The final reliability factors were much smaller than the typical values obtained in other methods such as the Molecular Replacement (MR) method. Our results successfully demonstrate that high-performance computing with GA-DS method on special-purpose computers is effective for the structure determination of biological molecules and the method has a potential to be widely used in near future.
Yousuke Ohno, Eiji Nishibori, Tetsu Narumi, Takahiro Koishi, Tahir H. Tahirov, Hideo Ago, Masashi Miyano, Ryutaro Himeno, Toshikazu Ebisuzaki, Makoto Sakata, Makoto Taiji
SC3
2006 Gordon Bell finalists II - A 55 TFLOPS simulation of amyloid-forming peptides from yeast prion Sup35 with the special-purpose computer system MDGRAPE-3
abstract
We have achieved a sustained performance of 55 TFLOPS for molecular dynamics simulations of the amyloid fibril formation of peptides from the yeast Sup35 in an aqueous solution. For performing the calculations, we used the MDGRAPE-3 system---a special-purpose computer system for molecular dynamics simulations. Its nominal peak performance was 415 TFLOPS for Coulomb force calculations; this is the highest-ever performance reported for classical molecular dynamics simulations. Amyloid fibril formation is known to be related to the occurrence of severe diseases such as Alzheimer's, Parkinson's, and Creutzfeldt-Jakob diseases. The Sup35 protein is a "yeast prion protein," which forms mini-crystals due to aggregation; it forms an effective platform for studying the formation process of amyloid fibrils. In these simulations, we first elucidate that the amyloid-forming peptides GNNQQNY aggregate at a higher frequency than non-amyloid-forming peptides SQNGNQQRG; further, the GNNQQNY peptides tend to form parallel two-stranded ß-sheets that would grow into a cross-ß amyloid nucleus. The results are consistent with those obtained experimentally. Furthermore, we could observe an early elongation of the amyloid nucleus. This result is expected to contribute toward a deeper understanding of the amyloid growth mechanism.
Tetsu Narumi, Yousuke Ohno, Noriaki Okimoto, Takahiro Koishi, Atsushi Suenaga, Noriyuki Futatsugi, Ryoko Yanai, Ryutaro Himeno, Shigenori Fujikawa, Makoto Taiji, Mitsuru Ikei
SC1
2003 Protein Explorer: A Petaflops Special-Purpose Computer System for Molecular Dynamics Simulations
abstract
We are developing the 'Protein Explorer' system, a petaflops special-purpose computer system for molecular dynamics simulations. The Protein Explorer is a PC cluster equipped with special-purpose engines that calculate nonbonded interactions between atoms, which is the most time-consuming part of the simulations. A dedicated LSI 'MDGRAPE-3 chip' performs these force calculations at a speed of 165 gigaflops or higher. The system will have 6,144 MDGRAPE-3 chips to achieve a nominal peak performance of one petaflop. The system will be completed in 2006. In this paper, we describe the project plans and the architecture of the Protein Explorer.
Makoto Taiji, Tetsu Narumi, Yousuke Ohno, Noriyuki Futatsugi, Atsushi Suenaga, Naoki Takada, Akihiko Konagaya
SC2
2001 An 8.61 Tflop/s molecular dynamics simulation for NaCl with a special-purpose computer: MDM
abstract
We performed molecular dynamics (MD) simulation of 33 million pairs of NaCl ions with the Ewald summation and obtained a calculation speed of 8.61 Tflop/s. In this calculation we used a special-purpose computer, MDM, which we have developed for the calculations of the Coulomb and van der Waals forces. The MDM enabled us to perform large scale MD simulations without truncating the Coulomb force. It is composed of MDGRAPE-2, WINE-2 and a host computer. MDGRAPE-2 accelerates the calculation for real-space part of the Coulomb and van der Waals forces. WINE-2 accelerates the calculation for wavenumber-space part of the Coulomb force. The host computer performs other calculations. With the completed MDM system we performed an MD simulation similar to what was the basis of our SC2000 submission for a Gordon Bell prize. With this large scale MD simulation, we can dramatically decrease the fluctuation of the temperature less than 0.1 Kelvin.
Tetsu Narumi, Atsushi Kawai, Takahiro Koishi
SC1
2000 1.34 Tflops Molecular Dynamics Simulation for NaCl with a Special-Purpose Computer: MDM
abstract
We performed molecular dynamics (MD) simulation of 9 million pairs of NaCl ions with the Ewald summation and obtained a calculation speed of 1.34 Tflops. In this calculation we used a special-purpose computer, MDM, which we are developing for the calculations of the Coulomb and van der Waals forces. The MDM enabled us to perform large scale MD simulations without truncating the Coulomb force. It is composed of WINE-2, MDGRAPE-2 and a host computer. WINE-2 accelerates the calculation for wavenumber-space part of the Coulomb force, while MDGRAPE-2 accelerates the calculation for real-space part of the Coulomb and van der Waals forces. The host computer performs other calculations. We performed MD simulation with the early version of the MDM system: 45 Tflops of WINE-2 and 1 Tflops of MDGRAPE-2. The peak performance of the final MDM system will reach 75 Tflops in total by the end of the year 2000.
Tetsu Narumi, Ryutaro Susukita, Takahiro Koishi, Kenji Yasuoka, Hideaki Furusawa, Atsushi Kawai, Toshikazu Ebisuzaki
SC1