EDBT 2026 Demo / reviewers in the wild / expert
Fumihiko Ino
dblp:64/201
· DBLP profile ↗
47ranked-venue papers
11as first author
9since 2021 · last 2026
0000-0002-5757-7631ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 9 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-authorArtificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A General-Purpose K-Nearest Neighbor Method with an Efficient Pruning Strategy for GPUs
Fumihiko Ino |
J. Parallel Distributed Comput. | 2 |
| 2025 | Overlapping Aware Data Placement Optimizations for LSM Tree-Based Store on ZNS SSDsabstractSolid State Drives (SSDs) based on the NVMe Zoned Namespaces (ZNS) interface can notably reduce the costs of address mapping, garbage collection, and over-provisioning by dividing the storage space into multiple zones for sequential writes and random reads. The Log-Structured Merge (LSM) tree, which is extensively used in key-value storage systems, converts random writes to sequential writes, hence a suitable scenario to utilize ZNS SSDs. However, LSM tree associated data significantly varies in lifetime due to the levels and merging mechanisms of the LSM tree. Therefore, without an accurate method to estimate data lifetime, data with disparate lifetimes may be placed in the same zone, thus causing low space utilization and high write amplification within the SSD. To address these issues, the article proposes two data overlapping aware optimizations to realize intelligent data placement: a zone allocation scheme and a garbage collection scheme. The key technique of these optimizations is an accurate data-lifetime estimation by considering both the associated tree level of the data and the data overlapping ratio between the data and those in the neighboring level. Using the estimation technique, the zone allocation optimization can place data with similar lifetimes in the same zone. Besides, the garbage collection optimization can reclaim zones in an adaptive manner based on overlapping ratios to reduce the amount of data migration. Experimental results demonstrate that the optimization schemes effectively reduce garbage collection-incurred data copy by average factors of 2.11× and 1.50× in comparison to a conventional work and a state-of-the-art work, respectively. Consequently, the proposed work successfully alleviates the write amplification effect by 18% and 6%, compared to the conventional work and the state-of-the-art work, respectively. Jingcheng Shen, Linbo Long, Zhenhua Tan, Congming Gao, Kan Zhong, Masao Okita, Fumihiko Ino |
ACM Trans. Archit. Code Optim. | 8 |
| 2025 | Lazy Qubit Reordering for Accelerating Parallel State-Vector-based Quantum Circuit SimulationabstractThis article proposes two quantum operation scheduling methods for accelerating parallel state-vector-based quantum circuit simulation using multiple graphics processing units (GPUs). The proposed methods reduce all-to-all communication caused by qubit reordering, which can dominate the overhead of parallel simulation. Our out-of-order approach eliminates redundant reorderings by introducing intentional delays in reordering communications such that multiple reorderings can be aggregated into a single reordering. The delays are carefully introduced based on the principles of time-space tiling, or a cache optimization technique for classical computers, which we use to arrange the execution order of quantum operations. Moreover, we develop these methods tailored for two primary procedures in variational quantum eigensolver simulation: quantum state update (QSU) and expectation value computation (EVC). Our QSU simulation takes an advantage of the hierarchical interconnection of GPU systems to avoid slow inter-node communication. On the other hand, our EVC simulation reduces the number of reorderings by diagonalization of Pauli strings. Experimental validation on 32-GPU executions demonstrates acceleration in QSU and EVC—up to 54× and 1,657×, respectively—compared to an inorder-based method. We believe that our out-of-order approach is useful for accelerating large-scale quantum circuit simulations, including QSU and/or EVC that operate qubits in a regular manner. Yusuke Teranishi, Shoma Hiraoka, Wataru Mizukami, Masao Okita, Fumihiko Ino |
ACM Trans. Quantum Comput. | 5 |
| 2024 | A theoretical and empirical exploration of TileTrans for effective tile pruning
Yanchen Li, Fumihiko Ino |
Knowl. Based Syst. | 2 |
| 2023 | PRF: A Fast Parallel Relaxed Flooding Algorithm for Voronoi Diagram Generation on GPUabstractThis paper introduces a novel parallel relaxed flooding (PRF) algorithm for Voronoi diagram generation. The algorithm takes a set of reference points extracted from an image as input and assigns each GPU thread a partition of the image domain to perform parallel flooding computation. Our PRF algorithm has three advantages as follows. (1) The PRF algorithm divides an image domain into subregions for concurrent flooding computation. To achieve high parallelism, a point selection method is incorporated to remove dependencies between different subregions. (2) We exploit the sparsity of the input point data with a k-d tree. With the k-d tree data structure, the point selection step achieves high efficiency, and the amount of CPU-GPU data transfer is reduced. (3) We propose a relaxed flooding method, which achieves more accurate results and decreases memory traffic compared to the traditional flooding method. In addition to these advantages, we provide an empirical method to determine the appropriate parameter in the point selection step for high performance, given an expected error rate. We evaluated the performance of our method on multiple datasets. Compared with the state-of-the-art parallel banding algorithm, our method achieved an average speed-up of 4.6× on the randomly generated datasets with a point density of 0.01%, and 6.8× on nuclei segmentation datasets. The code of the PRF algorithm is publicly available*. Fumihiko Ino, Jing Ke |
IPDPS | 2 |
| 2023 | A compression-based memory-efficient optimization for out-of-core GPU stencil computation
Jingcheng Shen, Linbo Long, Masao Okita, Fumihiko Ino |
J. Supercomput. | 5 |
| 2022 | A One-Shot Reparameterization Method for Reducing the Loss of Tile Pruning on DNNsabstractRecently, tile pruning has been widely studied to accelerate the inference of deep neural networks (DNNs). However, we found that the loss due to tile pruning, which can eliminate important elements together with unimportant elements, is large on trained DNNs. In this study, we propose a one-shot reparameterization method, called TileTrans, to reduce the loss of tile pruning. Specifically, we repermute the rows or columns of the weight matrix such that the model architecture can be kept unchanged after reparameterization. This repermutation realizes the reparameterization of the DNN model without any retraining. The proposed reparameterization method combines important elements into the same tile; thus, preserving the important elements after the tile pruning. Furthermore, TileTrans can be seamlessly integrated into existing tile pruning methods because it is a pre-processing method executed before pruning, which is orthogonal to most existing methods. The experimental results demonstrate that our method is essential in reducing the loss of tile pruning on DNNs. Specifically, the accuracy is improved by up to 17% for AlexNet while 5% for ResNet-34, where both models are pre-trained on ImageNet. Yanchen Li, Qingzhong Ai, Fumihiko Ino |
IJCNN | 3 |
| 2022 | Accelerating Imbalanced Many-to-Many Communication with Systematic Delay Insertion
Hirotoshi Yamada, Masao Okita, Fumihiko Ino |
PDCAT | 3 |
| 2021 | Accelerating GPU-Based Out-of-Core Stencil Computation with On-the-Fly Compression
Jingcheng Shen, Masao Okita, Fumihiko Ino |
PDCAT | 4 |
| 2020 | Accelerating Human Genome Phenotypic Analysis with Bitwise Search and Batched ComputationabstractWe propose an acceleration method for phenotypic analysis of monozygotic twins based on joint investigation of single-nucleotide polymorphisms (SNPs) and methylated cytosine-phosphate-guanine (mCpG) sites. The phenotypic analysis consists of the following two procedures: (1) identification of mCpG-SNP pairs closely located in base sequences and (2) t-tests on the identified mCpG-SNPs. The proposed method accelerates these procedures with bitwise search and batched computation. The bitwise search exploits the bit-level parallelism of the first procedure by simultaneously processing every 64 bases in sequences of over three billion bases. On the other hand, the batched computation exploits the data-parallelism of the second procedure, where millions of small t-tests must be processed for phenotypic analysis. Experimental results show that the proposed method on a single-core CPU is approximately 3.8 times faster than a previous method. Furthermore, the proposed method achieves a 14.6× speedup on two 10-core CPUs. We therefore believe that the proposed method is useful for accelerating genomewide association analysis, where large-scale data should be investigated in short time. Yuichiro Miyamoto, Masao Okita, Fumihiko Ino |
PDP | 3 |
| 2020 | Reducing the amount of out-of-core data access for GPU-accelerated randomized SVDabstractSummary We propose two acceleration methods, namely, Fused and Gram, for reducing out‐of‐core data access when performing randomized singular value decomposition (RSVD) on graphics processing units (GPUs). Out‐of‐core data here are data that are too large to fit into the GPU memory at once. Both methods accelerate GPU‐enabled RSVD using the following three schemes: (1) a highly tuned general matrix‐matrix multiplication (GEMM) scheme for processing out‐of‐core data on GPUs; (2) a data‐access reduction scheme based on one‐dimensional data partition; and (3) a first‐in, first‐out scheme that reduces CPU‐GPU data transfer using the reverse iteration. The Fused method further reduces the amount of out‐of‐core data access by merging two GEMM operations into a single operation. By contrast, the Gram method reduces both in‐core and out‐of‐core data access by explicitly forming the Gram matrix. According to our experimental results, the Fused and Gram methods improved the RSVD performance up to 1.7× and 5.2×, respectively, compared with a straightforward method that deploys schemes (1) and (2) on the GPU. In addition, we present a case study of deploying the Gram method for accelerating robust principal component analysis, a convex optimization problem in machine learning. Yuechao Lu, Ichitaro Yamazaki, Fumihiko Ino, Yasuyuki Matsushita, Stanimire Tomov, Jack J. Dongarra |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Transparent In-memory Cache Management in Apache Spark based on Post-Mortem AnalysisabstractThis paper proposes an extension to Apache Spark that provides automated and efficient in-memory cache management based on post-mortem dependency graph analysis. This extension allows programmers to focus on algorithmic issues without code modification for caching decision. We realized this extension with two techniques: (1) a selection algorithm for intermediate data to be cached and (2) a cache replacement algorithm looking ahead to the whole execution. For avoiding all recalculations in case of no cache replacement, the selection algorithm implicitly activates the necessary cache directives. When cache replacement is unavoidable, the cache replacement algorithm prevents frequent cache replacement that came from excessive cache directives. Experimental results demonstrate that machine learning applications on the extended runtime achieved competitive or higher performance up to 1.3 times compared to manually optimized programs, except for very simple cases. We expect that the proposed method is useful to reduce the cache management burden on programmers without regard to the amount of processing data. Atsuya Nasu, Kenji Yoneo, Masao Okita, Fumihiko Ino |
IEEE BigData | 4 |
| 2019 | GPU-based branch-and-bound method to solve large 0-1 knapsack problems with data-centric strategiesabstractSummary An out‐of‐core branch‐and‐bound (B&B) method to solve large 0‐1 knapsack problems on a graphics processing unit (GPU) is proposed. Given a large problem that produces many subproblems, the proposed method dynamically swaps subproblems to CPU memory. Because such a CPU‐centric subproblem management scheme increases CPU‐GPU data transfer, we adopt three data‐centric strategies to eliminate this side effect. The first is an out‐of‐order search (O3S) strategy that reduces the data transfer overhead by adaptively transferring subproblems between the CPU and GPU. The second is an explicitly‐managed pipelining strategy that hides the data transfer overhead by overlapping data transfer with GPU‐based B&B operations. The third is a GPU‐based stream compaction strategy that reduces the sparseness of arrays to be transferred. Experimental results demonstrate that the proposed out‐of‐core method stored 41 times as many subproblems as a previous in‐core method that manages subproblems in GPU memory, solving approximately twice as many problem instances on the GPU. In addition, compared to a previous breadth‐first search (BFS) strategy, the proposed O3S strategy achieved an average speedup of 7.5 times. Jingcheng Shen, Kentaro Shigeoka, Fumihiko Ino, Kenichi Hagihara |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Accelerating scoring computation of Smith-Waterman algorithm with mixed word lengthabstractIn this paper, we propose a graphics processing unit (GPU) accelerated Smith-Waterman (SW) method for aligning a short sequence with a long sequence. Our method deploys an 8-bit data structure to accelerate scoring computation, which can be classified into a memory-intensive operation. The proposed method is based on a mixed-word-length scheme capable of appropriately switching 8-bit mode and 32-bit mode to avoid integer overflow during computation. We integrate our data structure and scheme into a scan-based parallel approach to achieve high throughput for short sequences, which usually limit the parallelism inherent in the SW computation. Experimental results show that the proposed method is approximately three times faster than a previous scan-based approach that uniformly deploys 32-bit data structure. We also show that, for short sequences, scan-based parallel approaches are about twenty times faster than an antidiagonal-based parallel approach typically deployed for pairs of long sequences. Kazuki Yasui, Fumihiko Ino |
BIBM | 2 |
| 2017 | An Out-of-Core Branch and Bound Method for Solving the 0-1 Knapsack Problem on a GPU
Jingcheng Shen, Kentaro Shigeoka, Fumihiko Ino, Kenichi Hagihara |
ICA3PP | 3 |
| 2017 | Parallelizing Exact and Approximate String Matching via Inclusive Scan on a GPUabstractIn this study, to substantially improve the runtimes of exact and approximate string matching algorithms, we propose a tribrid parallel method for bit-parallel algorithms such as the Shift-Or and Wu-Manber algorithms. Our underlying idea is to interpret bit-parallel algorithms as inclusive-scan operations, which allow these bit-parallel algorithms to run efficiently on a graphics processing unit (GPU); we achieve this speed-up here because inclusive-scan operations not only eliminate duplicate searches between threads but also realize a GPU-friendly memory access pattern that maximizes memory read/write throughput. To realize our ideas, we first define two binary operators and then present a proof regarding the associativity of these operators, which is necessary for the parallelization of the inclusive-scan operations. Finally, we integrate the inclusive-scan scheme into a previous segmentation-based scheme to maximize search throughput, identifying the best tradeoff point between synchronization cost and duplicate work. Through our experiments, we compared our proposed method with previous segmentation-based methods and indexing-based sequence aligners. For online string matching, our proposed method performed 6.7-16.7 times faster than previous methods, achieving a search throughput of up to 1.88 terabits per second (Tbps) on a GeForce GTX TITAN X GPU. We therefore conclude that our proposed method is quite effective for decreasing the runtimes of online string matching of short patterns. Yasuaki Mitani, Fumihiko Ino, Kenichi Hagihara |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | An OpenACC Optimizer for Accelerating Histogram Computation on a GPUabstractThis paper presents a source-to-source OpenACC optimizer that automatically optimizes a histogram computation code for a graphics processing unit (GPU). Parallel histogram computation codes typically deploy multiple copies of histograms and update them with atomic operations. This duplication method can be implemented as an OpenACC code. However, the structure of sequential code blocks must be manually rewritten owing to the limitation on OpenACC directives. Such a rewritten code does not always achieve the highest performance on arbitrary platforms, and thus, the duplication method degrades the performance portability of the code. To tackle this issue, we propose an optimizer that identifies histogram-related blocks in a naive OpenACC code and automatically rewrites the detected blocks such that multiple copies of histograms can be exploited for acceleration. In experiments, we apply our optimizer to three practical applications and investigate their performance on three platforms: an NVIDIA GPU, an AMD GPU and an Intel CPU. Experimental results show that our automated approach is useful for OpenACC codes to maximize the performance of histogram computation, and thereby enhancing the performance portability of the code. Kei Ikeda, Fumihiko Ino, Kenichi Hagihara |
PDP | 2 |
| 2016 | Reducing memory usage by the lifting-based discrete wavelet transform with a unified buffer on a GPU
Takuya Ikuzawa, Fumihiko Ino, Kenichi Hagihara |
J. Parallel Distributed Comput. | 2 |
| 2015 | Accelerating the Smith-Waterman algorithm with interpair pruning and band optimization for the all-pairs comparison of base sequencesabstractBACKGROUND: The Smith-Waterman algorithm is known to be a more sensitive approach than heuristic algorithms for local sequence alignment algorithms. Despite its sensitivity, a greater time complexity associated with the Smith-Waterman algorithm prevents its application to the all-pairs comparisons of base sequences, which aids in the construction of accurate phylogenetic trees. The aim of this study is to achieve greater acceleration using the Smith-Waterman algorithm (by realizing interpair block pruning and band optimization) compared with that achieved using a previous method that performs intrapair block pruning on graphics processing units (GPUs). RESULTS: We present an interpair optimization method for the Smith-Waterman algorithm with the aim of accelerating the all-pairs comparison of base sequences. Given the results of the pairs of sequences, our method realizes efficient block pruning by computing a lower bound for other pairs that have not yet been processed. This lower bound is further used for band optimization. We integrated our interpair optimization method into SW#, a previous GPU-based implementation that employs variants of a banded Smith-Waterman algorithm and a banded Myers-Miller algorithm. Evaluation using the six genomes of Bacillus anthracis shows that our method pruned 88% of the matrix cells on a single GPU and 73% of the matrix cells on two GPUs. For the genomes of the human chromosome 21, the alignment performance reached 202 giga-cell updates per second (GCUPS) on two Tesla K40 GPUs. CONCLUSIONS: Efficient interpair pruning and band optimization makes it possible to complete the all-pairs comparisons of the sequences of the same species 1.2 times faster than the intrapair pruning method. This acceleration was achieved at the first phase of SW#, where our method significantly improved the initial lower bound. However, our interpair optimization was not effective for the comparison of the sequences of different species such as comparing human, chimpanzee, and gorilla. Consequently, our method is useful in accelerating the applications that require optimal local alignments scores for the same species. The source code is available for download from http://www-hagi.ist.osaka-u.ac.jp/research/code/. Daiki Okada, Fumihiko Ino, Kenichi Hagihara |
BMC Bioinform. | 2 |
| 2015 | A bit-parallel algorithm for searching multiple patterns with various lengths
Ko Kusudo, Fumihiko Ino, Kenichi Hagihara |
J. Parallel Distributed Comput. | 2 |
| 2014 | A parallel scheme for accelerating parameter sweep applications on a GPUabstractSUMMARY This paper proposes a parallel scheme for accelerating parameter sweep applications on a graphics processing unit. By using hundreds of cores on the graphics processing unit, we found that our scheme simultaneously processes multiple parameters rather than a single parameter. The simultaneous sweeps exploit the similarity of computing behaviors shared by different parameters, thus allowing memory accesses to be coalesced into a single access if similar irregularities appear among the parameters’ computational tasks. In addition, our scheme reduces the amount of off‐chip memory access by unifying the data that are commonly referenced by multiple parameters and by placing the unified data in the fast on‐chip memory. In several experiments, we applied our scheme to practical applications and found that our scheme can perform up to 8.5 times faster than a naive scheme that processes a single parameter at a time. We also include a discussion on application characteristics that are required for our scheme to outperform the naive scheme. Copyright © 2013 John Wiley & Sons, Ltd. Fumihiko Ino, Kentaro Shigeoka, Tomohiro Okuyama, Masaya Motokubota, Kenichi Hagihara |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Improving cache locality for GPU-based volume rendering
Yuki Sugimoto, Fumihiko Ino, Kenichi Hagihara |
Parallel Comput. | 2 |
| 2014 | Efficient Acceleration of Mutual Information Computation for Nonrigid Registration Using CUDAabstractIn this paper, we propose an efficient acceleration method for the nonrigid registration of multimodal images that uses a graphics processing unit. The key contribution of our method is efficient utilization of on-chip memory for both normalized mutual information (NMI) computation and hierarchical B-spline deformation, which compose a well-known registration algorithm. We implement this registration algorithm as a compute unified device architecture program with an efficient parallel scheme and several optimization techniques such as hierarchical data organization, data reuse, and multiresolution representation. We experimentally evaluate our method with four clinical datasets consisting of up to 512 × 512 × 296 voxels. We find that exploitation of on-chip memory achieves a 12-fold increase in speed over an off-chip memory version and, therefore, it increases the efficiency of parallel execution from 4% to 46%. We also find that our method running on a GeForce GTX 580 card is approximately 14 times faster than a fully optimized CPU-based implementation running on four cores. Some multimodal registration results are also provided to understand the limitation of our method. We believe that our highly efficient method, which completes an alignment task within a few tens of seconds, will be useful to realize rapid nonrigid registration. Kei Ikeda, Fumihiko Ino, Kenichi Hagihara |
IEEE J. Biomed. Health Informatics | 2 |
| 2012 | Cooperative multitasking for GPU-accelerated grid systemsabstractSUMMARY This paper presents a cooperative multitasking method for concurrent execution of scientific and graphics applications on the graphics processing unit (GPU). Our method is designed to accelerate compute unified device architecture‐based applications using idle GPU cycles in the office. To prevent significant slow‐down of graphics applications, the method divides scientific tasks into smaller pieces, which are then sequentially executed at the appropriate intervals. The method also has flexibility in finding the best tradeoff point between scientific applications and graphics applications. Experimental results show that the proposed method is useful to control the frame rate of the graphics application and the throughput of the scientific application. For example, biological sequence alignment can be processed at approximately 30% of the dedicated throughput while achieving interactive rendering at 58 frames per second. We also show that matrix multiplication can be efficiently processed at 60% of the dedicated throughput during word processing and web browsing. Copyright © 2011 John Wiley & Sons, Ltd. Fumihiko Ino, Akihiro Ogita, Kentaro Oita, Kenichi Hagihara |
Concurr. Comput. Pract. Exp. | 1 |
| 2012 | Sequence Homology Search Using Fine Grained Cycle Sharing of Idle GPUsabstractIn this paper, we propose a Fine Grained Cycle Sharing (FGCS) system capable of exploiting idle Graphics Processing Units (GPUs) for accelerating sequence homology search in local area network environments. Our system exploits short idle periods on GPUs by running small parts of guest programs such that each part can be completed within hundreds of milliseconds. To detect such short idle periods from the pool of registered resources, our system continuously monitors keyboard and mouse activities via event handlers rather than waiting for a screensaver, as is typically deployed in existing systems. Our system also divides guest tasks into small parts according to a performance model that estimates execution times of the parts. This task division strategy minimizes any disruption to the owners of the GPU resources. Experimental results show that our FGCS system running on two nondedicated GPUs achieves 111-116 percent of the throughput achieved by a single dedicated GPU. Furthermore, our system provides over two times the throughput of a screensaver-based system. We also show that the idle periods detected by our system constitute half of the system uptime. We believe that the GPUs hidden and often unused in office environments provide a powerful solution to sequence homology search. Fumihiko Ino, Yuma Munekawa, Kenichi Hagihara |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Accelerating Parameter Sweep Applications Using CUDAabstractThis paper proposes a parallelization scheme for parameter sweep (PS) applications using the compute unified device architecture (CUDA). Our scheme focuses on PS applications with irregular access patterns, which usually result in lower performance on the GPU. The key idea to resolve this irregularity is to exploit the similarity of data accesses between different parameters. That is, the scheme simultaneously processes multiple parameters instead of a single parameter. This simultaneous sweep allows data accesses to be coalesced into a single access if the irregularity appears similarly at every parameter. It also reduces the amount of off-chip memory access by using fast on-chip memory for the data commonly accessed for multiple parameters. As a result, the scheme achieves up to 4.5 times higher performance than a naive scheme that processes a single parameter by a kernel invocation. Masaya Motokubota, Fumihiko Ino, Kenichi Hagihara |
PDP | 2 |
| 2010 | Cooperative Multitasking for GPU-Accelerated Grid SystemsabstractExploiting the graphics processing unit (GPU) is useful to obtain higher performance with a less number of host machines in grid systems. One problem in GPU-accelerated grid systems is the lack of efficient multitasking mechanisms. In this paper, we propose a cooperative multitasking method capable of simultaneous execution of a graphics application and a CUDA-based scientific application on a single GPU. To prevent significant performance drop in frame rate, our method (1) divides scientific tasks into smaller subtasks and (2) serially executes them at the appropriate intervals. Experimental results show that the proposed method is useful to control the frame rate of the graphics application and the throughput of the scientific application. For example, matrix multiplication can be processed at 50% of the dedicated throughput while achieving interactive rendering at 54 frames per second. Fumihiko Ino, Akihiro Ogita, Kentaro Oita, Kenichi Hagihara |
CCGRID | 1 |
| 2010 | High-performance cone beam reconstruction using CUDA compatible GPUs
Yusuke Okitsu, Fumihiko Ino, Kenichi Hagihara |
Parallel Comput. | 2 |
| 2009 | Harnessing the power of idle GPUs for acceleration of biological sequence alignmentabstractThis paper presents a parallel system capable of accelerating biological sequence alignment on the graphics processing unit (GPU) grid. The GPU grid in this paper is a desktop grid system that utilizes idle GPUs and CPUs in the office and home. Our parallel implementation employs a master-worker paradigm to accelerate Liu's OpenGL-based algorithm that runs on a single GPU. We integrate this implementation into a screensaver-based grid system that detects idle resources on which the alignment code can run. We also show some experimental results comparing our implementation with three different implementations running on a single GPU, a single CPU, or multiple CPUs. As a result, we find that a single non-dedicated GPU can provide us almost the same throughput as two dedicated CPUs in our laboratory environment, where GPU-equipped machines are ordinarily used to develop GPU applications. Fumihiko Ino, Yuki Kotani, Kenichi Hagihara |
IPDPS | 1 |
| 2008 | Design and implementation of the Smith-Waterman algorithm on the CUDA-compatible GPUabstractThis paper describes a design and implementation of the Smith-Waterman algorithm accelerated on the graphics processing unit (GPU). Our method is implemented using compute unified device architecture (CUDA), which is available on the nVIDIA GPU. The method efficiently uses on-chip shared memory to reduce the data amount being transferred between off-chip memory and processing elements in the GPU. Furthermore, it reduces the number of data fetches by applying a data reuse technique to query and database sequences. We show some experimental results comparing the proposed method with an OpenGL-based method. As a result, the speedup over the OpenGL-based method reaches a factor of 6.4 when using amino acid sequence database.We also find that shared memory reduces the amount of data fetches to 1/140, providing a peak performance of 5.65 giga cell updates per second (GCUPS). This performance is approximately three times faster than a prior CUDA-based implementation. Yuma Munekawa, Fumihiko Ino, Kenichi Hagihara |
BIBE | 2 |
| 2008 | Accelerating Cone Beam Reconstruction Using the CUDA-Enabled GPU
Yusuke Okitsu, Fumihiko Ino, Kenichi Hagihara |
HiPC | 2 |
| 2008 | A Task Parallel Algorithm for Computing the Costs of All-Pairs Shortest Paths on the CUDA-Compatible GPUabstractThis paper proposes a fast method for computing the costs of all-pairs shortest paths (APSPs) on the graphics processing unit (GPU). The proposed method is implemented using compute unified device architecture (CUDA), which offers us a development environment for performing general-purpose computation on the GPU. Our method is based on Harish's iterative algorithm that computes the cost of the single-source shortest path (SSSP) for every source vertex. We present that exploiting task parallelism in the APSP problem allows us to efficiently use on-chip memory in the GPU, reducing the amount of data being transferred from relatively slower off-chip memory. Furthermore, our task parallel scheme is useful to exploit a higher parallelism, increasing the efficiency with highly threaded code. As a result, our method is 3.4--15 times faster than the prior method. Using on-chip memory, our method eliminates approximately 20% of data loads from off-chip memory. Tomohiro Okuyama, Fumihiko Ino, Kenichi Hagihara |
ISPA | 2 |
| 2008 | A decompression pipeline for accelerating out-of-core volume rendering of time-varying data
Daisuke Nagayasu, Fumihiko Ino, Kenichi Hagihara |
Comput. Graph. | 2 |
| 2008 | A Resource Selection System for Cycle Stealing in GPU Grids
Yuki Kotani, Fumihiko Ino, Kenichi Hagihara |
J. Grid Comput. | 2 |
| 2006 | A code motion technique for accelerating general-purpose computation on the GPUabstractGraphics processing units (GPUs) are providing increasingly higher performance with programmable internal processors, namely vertex processors (VPs) and fragment processors (FPs). Such newly added capabilities motivate us to perform general-purpose computation on GPUs (GPGPU) beyond graphics applications. Although VPs and FPs are connected in a pipeline, many GPGPU implementations utilize only FPs as a computational engine in the GPU. Therefore, such implementations may result in lower performance due to highly loaded FPs (as compared to VPs) being a performance bottleneck in the pipeline execution. The objective of our work is to improve the performance of GPGPU programs by eliminating this bottleneck. To achieve this, we present a code motion technique that is capable of reducing the FP workload by moving assembly instructions appropriately from the FP program to the VP program. We also present the definition of such movable instructions that do not change the I/O specification between the CPU and the GPU. The experimental results show that (1) our technique improves the performance of a Gaussian filter program with reducing execution time by approximately 40% and (2) it successfully reduces the FP workload in 10 out of 18 GPGPU programs. Takatoshi Ikeda, Fumihiko Ino, Kenichi Hagihara |
IPDPS | 2 |
| 2006 | A GPGPU Approach for Accelerating 2-D/3-D Rigid Registration of Medical Images
Fumihiko Ino, Jun Gomita, Yasuhiro Kawasaki, Kenichi Hagihara |
ISPA | 1 |
| 2005 | Performance Study of LU Decomposition on the Programmable GPU
Fumihiko Ino, Manabu Matsui, Keigo Goda, Kenichi Hagihara |
HiPC | 1 |
| 2005 | Performance Study of Nonrigid Registration Algorithm for Investigating Lung Disease on ClustersabstractThis paper presents a performance study of a nonrigid registration algorithm for investigating lung disease on clusters. Our algorithm combines two conventional acceleration techniques in order to achieve fast registration: a data-parallel processing technique for accelerating the registration procedure; and a precomputation technique for reducing the computational complexity. We perform some experiments on three clusters with different CPU and network performance in order to make clear what kinds of acceleration techniques and computing environments provide higher performance. The results show that a cluster with Gigabit Ethernet (GbE) network is the most cost effective solution that reduces registration time from ten hours to ten minutes with a linear speedup. Fumihiko Ino, Yuya Tanaka, Kenichi Hagihara, Hiroko Kitaoka |
PDCAT | 1 |
| 2005 | A data distributed parallel algorithm for nonrigid image registration
Fumihiko Ino, Kanrou Ooyama, Kenichi Hagihara |
Parallel Comput. | 1 |
| 2004 | Parallel Volume Rendering with Early Ray Termination for Visualizing Large-Scale Datasets
Manabu Matsui, Fumihiko Ino, Kenichi Hagihara |
ISPA | 2 |
| 2004 | Real-Time Estimation of Hip Range of Motion for Total Hip Replacement Surgery
Yasuhiro Kawasaki, Fumihiko Ino, Yoshinobu Sato, Nobuhiko Sugano, Hideki Yoshikawa, Shinichi Tamura, Kenichi Hagihara |
MICCAI (2) | 2 |
| 2004 | High-performance computing service over the Internet for intraoperative image processingabstractThis paper presents a framework for a cluster system that is suited for high-resolution image processing over the Internet during surgery. The system realizes high-performance computing (HPC) assisted surgery, which allows surgeons to utilize HPC resources remote from the operating room. One application available in the system is an intraoperative estimator for the range of motion (ROM) adjustment in total hip replacement (THR) surgery. In order to perform this computation-intensive estimation during surgery, we parallelize the ROM estimator on a cluster of 64 PCs, each with two CPUs. Acceleration techniques such as dynamic load balancing and data compression methods are incorporated into the system. The system also provides a remote-access service over the Internet with a secure execution environment. We applied the system to an actual THR surgery performed at Osaka University Hospital and confirmed that it realizes intraoperative ROM estimation without degrading the resolution of images and limiting the area for estimations. Yasuhiro Kawasaki, Fumihiko Ino, Yasuharu Mizutani, Noriyuki Fujimoto, Toshihiko Sasama, Yoshinobu Sato, Nobuhiko Sugano, Shinichi Tamura, Kenichi Hagihara |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2003 | An Emulation System for Predicting Master/Slave Program Performance
Yasuharu Mizutani, Fumihiko Ino, Kenichi Hagihara |
Euro-Par | 2 |
| 2003 | A High Performance Computing System for Medical Imaging in the Remote Operating Room
Yasuhiro Kawasaki, Fumihiko Ino, Yasuharu Mizutani, Noriyuki Fujimoto, Toshihiko Sasama, Yoshinobu Sato, Shinichi Tamura, Kenichi Hagihara |
HiPC | 2 |
| 2003 | Design and Implementation of Parallel Nonrigid Image Registration Using Off-the-Shelf Supercomputers
Fumihiko Ino, Kanrou Ooyama, Akira Takeuchi, Kenichi Hagihara |
MICCAI (1) | 1 |
| 2003 | An improved binary-swap compositing for sort-last parallel rendering on distributed memory multiprocessors
Akira Takeuchi, Fumihiko Ino, Kenichi Hagihara |
Parallel Comput. | 2 |
| 2001 | LogGPS: a parallel computational model for synchronization analysisabstractWe present a new parallel computational model, named LogGPS, which captures synchronization. Fumihiko Ino, Noriyuki Fujimoto, Kenichi Hagihara |
PPoPP | 1 |