VLDB 2026 Research / reviewers in the wild / expert
Tsuyoshi Hamada
dblp:61/5903
· DBLP profile ↗
15ranked-venue papers
5as first author
1since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-authorArtificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 32% GPUs and heterogeneous computing · 32% Performance modeling and evaluation · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › multi-GPU computing
GPU cluster |
0.2 | 2 | 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009 |
Interconnection networks and networks-on-chip › cluster interconnect
infiniband |
0.1 | 1 | 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010 |
High-performance computing › n-body simulation
treecode |
0.1 | 1 | 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010 |
Performance modeling and evaluation › numerical algorithms
fast multipole method |
0.1 | 1 | 2009 | 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009 |
High-performance computing › n-body simulation
hierarchical n-body methods |
0.1 | 1 | 2009 | 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009 |
Performance modeling and evaluation › design trade-off analysis
cost-performance analysis |
0.0 | 1 | 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUs · SC 2010 |
Computational science and engineering › computational fluid dynamics
turbulence simulation |
0.0 | 1 | 2009 | 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulence · SC 2009 |
Methods — techniques the papers use, named apart from their topics
treecode · 0.3fast multipole method · 0.2hierarchical n-body · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tumor-immune partitioning and clustering algorithm for identifying tumor-immune cell spatial interaction signatures within the tumor microenvironmentabstractBACKGROUND: Growing evidence supports the importance of characterizing the organizational patterns of various cellular constituents in the tumor microenvironment in precision oncology. Most existing data on immune cell infiltrates in tumors, which are based on immune cell counts or nearest neighbor-type analyses, have failed to fully capture the cellular organization and heterogeneity. METHODS: We introduce a computational algorithm, termed Tumor-Immune Partitioning and Clustering (TIPC), that jointly measures immune cell partitioning between tumor epithelial and stromal areas and immune cell clustering versus dispersion. As proof-of-principle, we applied TIPC to a prospective cohort incident tumor biobank containing 931 colorectal carcinoma cases. TIPC identified tumor subtypes with unique spatial patterns between tumor cells and T lymphocytes linked to certain molecular pathologic and prognostic features. T lymphocyte identification and phenotyping were achieved using multiplexed (multispectral) immunofluorescence. In a separate hepatocellular carcinoma cohort, we replaced the stromal component with specific immune cell types-CXCR3+CD68+ or CD8+-to profile their spatial relationships with CXCL9+CD68+ cells. RESULTS: Six unsupervised TIPC subtypes based on T lymphocyte distribution patterns were identified, comprising two cold and four hot subtypes. Three of the four hot subtypes were associated with significantly longer colorectal cancer (CRC)-specific survival compared to a reference cold subtype. Our analysis showed that variations in T-cell densities among the TIPC subtypes did not strictly correlate with prognostic benefits, underscoring the prognostic significance of immune cell spatial patterns. Additionally, TIPC revealed two spatially distinct and cell density-specific subtypes among microsatellite instability-high colorectal cancers, indicating its potential to upgrade tumor subtyping. TIPC was also applied to additional immune cell types, eosinophils and neutrophils, identified using morphology and supervised machine learning; here two tumor subtypes with similarly low densities, namely 'cold, tumor-rich' and 'cold, stroma-rich', exhibited differential prognostic associations. Lastly, we validated our methods and results using The Cancer Genome Atlas colon and rectal adenocarcinoma data (n = 570). Moreover, applying TIPC to hepatocellular carcinoma cases (n = 27) highlighted critical cell interactions like CXCL9-CXCR3 and CXCL9-CD8. CONCLUSIONS: Unsupervised discoveries of microgeometric tissue organizational patterns and novel tumor subtypes using the TIPC algorithm can deepen our understanding of the tumor immune microenvironment and likely inform precision cancer immunotherapy. Mai Chan Lau, Jennifer Borowsky, Juha P. Väyrynen, Koichiro Haruki, Melissa Zhao, Andressa Dias Costa, Simeng Gu, Annacarolina Da Silva, Tomotaka Ugai, Kota Arima, Minh N. Nguyen, Yasutoshi Takashima, Joe Yeong, David Tai, Tsuyoshi Hamada, Jochen K. Lennerz, Charles S. Fuchs, Catherine J. Wu, Jeffrey A. Meyerhardt, Shuji Ogino, Jonathan A. Nowak |
PLoS Comput. Biol. | 15 |
| 2017 | Biomarker correlation network in colorectal carcinoma by tumor anatomic locationabstractBACKGROUND: Colorectal carcinoma evolves through a multitude of molecular events including somatic mutations, epigenetic alterations, and aberrant protein expression, influenced by host immune reactions. One way to interrogate the complex carcinogenic process and interactions between aberrant events is to model a biomarker correlation network. Such a network analysis integrates multidimensional tumor biomarker data to identify key molecular events and pathways that are central to an underlying biological process. Due to embryological, physiological, and microbial differences, proximal and distal colorectal cancers have distinct sets of molecular pathological signatures. Given these differences, we hypothesized that a biomarker correlation network might vary by tumor location. RESULTS: We performed network analyses of 54 biomarkers, including major mutational events, microsatellite instability (MSI), epigenetic features, protein expression status, and immune reactions using data from 1380 colorectal cancer cases: 690 cases with proximal colon cancer and 690 cases with distal colorectal cancer matched by age and sex. Edges were defined by statistically significant correlations between biomarkers using Spearman correlation analyses. We found that the proximal colon cancer network formed a denser network (total number of edges, n = 173) than the distal colorectal cancer network (n = 95) (P < 0.0001 in permutation tests). The value of the average clustering coefficient was 0.50 in the proximal colon cancer network and 0.30 in the distal colorectal cancer network, indicating the greater clustering tendency of the proximal colon cancer network. In particular, MSI was a key hub, highly connected with other biomarkers in proximal colon cancer, but not in distal colorectal cancer. Among patients with non-MSI-high cancer, BRAF mutation status emerged as a distinct marker with higher connectivity in the network of proximal colon cancer, but not in distal colorectal cancer. CONCLUSION: In proximal colon cancer, tumor biomarkers tended to be correlated with each other, and MSI and BRAF mutation functioned as key molecular characteristics during the carcinogenesis. Our findings highlight the importance of considering multiple correlated pathways for therapeutic targets especially in proximal colon cancer. Reiko Nishihara, Kimberly Glass, Kosuke Mima, Tsuyoshi Hamada, Jonathan A. Nowak, Zhi Rong Qian, Peter Kraft, Edward L. Giovannucci, Charles S. Fuchs, Andrew T. Chan, John Quackenbush, Shuji Ogino, Jukka-Pekka Onnela |
BMC Bioinform. | 4 |
| 2016 | MrBayes tgMC3++: A High Performance and Resource-Efficient GPU-Oriented Phylogenetic Analysis MethodabstractMrBayes is a widespread phylogenetic inference tool harnessing empirical evolutionary models and Bayesian statistics. However, the computational cost on the likelihood estimation is very expensive, resulting in undesirably long execution time. Although a number of multi-threaded optimizations have been proposed to speed up MrBayes, there are bottlenecks that severely limit the GPU thread-level parallelism of likelihood estimations. This study proposes a high performance and resource-efficient method for GPU-oriented parallelization of likelihood estimations. Instead of having to rely on empirical programming, the proposed novel decomposition storage model implements high performance data transfers implicitly. In terms of performance improvement, a speedup factor of up to 178 can be achieved on the analysis of simulated datasets by four Tesla K40 cards. In comparison to the other publicly available GPU-oriented MrBayes, the tgMC3++ method (proposed herein) outperforms the tgMC3(v1.0), nMC3(v2.1.1) and oMC3(v1.00) methods by speedup factors of up to 1.6, 1.9 and 2.9, respectively. Moreover, tgMC3++ supports more evolutionary models and gamma categories, which previous GPU-oriented methods fail to take into analysis. Cheng Ling, Tsuyoshi Hamada, Jingyang Gao, Guoguang Zhao, Donghong Sun, Weifeng Shi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | Optimizing the Bayesian Inference of Phylogeny on Graphic ProcessorsabstractSearching for the evolutionary relationships between groups of organism has become a routine procedure in molecular biology. MrBayes is a popular model based phylogenetic inference tool using Bayesian statistics. Unfortunately, the computational cost is very high, resulting in undesirably long execution time. In this paper, we present what we believe the fastest solution of the MrBayes MC3 algorithm running on off-the-shelf graphic processors. The performance benefits are offered by the multi-granularity parallelism model, coarse-grained GPU kernel system, efficient thread arrangement strategy and GPU code level optimizations. MrBayes goMC3 (proposed herein) provides a significant performance improvement over the sequential MrBayes MC3 by a speedup of up to 48× when using single Tesla C2075 GPU card, whereas a speedup factor of 77× can be achieved when using dual GPUs. In comparison to the state-of-the-art version of other publicly available GPU implementations of MrBayes MC3, the cumulative optimizations adopted in goMC3 resulted in a speedup of up 2.5× over oMC3 (v1.0), 1.75× over tgMC3 (v1.0) and 1.46× over nMC3(v2.1.1) for realistic empirical biological datasets. Besides, experimental results indicated that goMC3 outstrips these GPU implementations on the analysis of simulated datasets composed of ultra-large-scale sequences. As a consequence, the reported performance improvement of goMC3 is significant and appears to scale well with increasing dataset sizes. Cheng Ling, Chunbao Zhou, Arong Luo, Guoguang Zhao, Tsuyoshi Hamada, Xiaoyan Zhu 0001 |
CCGRID | 5 |
| 2014 | A GRASS GIS parallel module for radio-propagation predictionsabstractGeographical information systems are ideal candidates for the application of parallel programming techniques, mainly because they usually handle large data sets. To help us deal with complex calculations over such data sets, we investigated the performance constraints of a classic master–worker parallel paradigm over a message-passing communication model. To this end, we present a new approach that employs an external database in order to improve the calculation–communication overlap, thus reducing the idle times for the worker processes. The presented approach is implemented as part of a parallel radio-coverage prediction tool for the Geographic Resources Analysis Support System (GRASS) environment. The prediction calculation employs digital elevation models and land-usage data in order to analyze the radio coverage of a geographical area. We provide an extended analysis of the experimental results, which are based on real data from an Long Term Evolution (LTE) network currently deployed in Slovenia. Based on the results of the experiments, which were performed on a computer cluster, the new approach exhibits better scalability than the traditional master–worker approach. We successfully tackled real-world-sized data sets, while greatly reducing the processing time and saturating the hardware utilization. Lucas Benedicic, Felipe A. Cruz, Tsuyoshi Hamada, Peter Korosec |
Int. J. Geogr. Inf. Sci. | 3 |
| 2010 | Highly efficient mapping of the Smith-Waterman algorithm on CUDA-compatible GPUsabstractThis paper describes a multi-threaded parallel design and implementation of the Smith-Waterman (SW) algorithm on graphic processing units (GPUs) with NVIDIA corporation's Compute Unified Device Architecture (CUDA). Central to this is a divide and conquer approach which divides the computation of a whole pairwise sequence alignment matrix into multiple sub-matrices (or parallelograms) each running efficiently on the available hardware resources of the GPU in hand, with temporary intermediate data stored in global memory. Moreover, we use thread warps and padding techniques in order to decrease the cost of thread synchronization, as well as loop unrolling in order to reduce the cost of conditional branches. While intermediate data is stored in global memory for large queries, the most inner loop in our implementation will only access shared memory and registers. As a result of these optimizations, our implementation of the SW algorithm achieves a throughput ranging between 9.09 GCUPS (Giga Cell Update per Second) and 12.71 GCUPS on a single-GPU version, and a throughput between 29.46 GCUPS and 43.05 GCUPS on a quad-GPU platform. Compared with the best GPU implementation of the SW algorithm reported to date, our implementation achieves up to 46 % improvement in speed. The source code of our implementation is available in the public domain for Bioinformaticians to benefit from its performance. Keisuke Dohi, Khaled Benkrid, Cheng Ling, Tsuyoshi Hamada, Yuichiro Shibata |
ASAP | 4 |
| 2010 | 190 TFlops Astrophysical N-body Simulation on a Cluster of GPUsabstractWe present the results of a hierarchical N-body simulation on DEGIMA, a cluster of PCs with 576 graphic processing units (GPUs) and using an InfiniBand interconnect. DEGIMA stands for DEstination for GPU Intensive MAchine, and is located at Nagasaki Advanced Computing Center (NACC), Nagasaki University. In this work, we have upgraded DEGIMA_s interconnect using InfiniBand. DEGIMA is composed by 144 nodes with 576 GT200 GPUs. An astrophysical N-body simulation with 3,278,982,596 particles using a treecode algorithm shows a sustained performance of 190.5 Tflops on DEGIMA. The overall cost of the hardware was $411,921 dollars. The maximum corrected performance is 104.8 Tflops for the simulation, resulting in a cost performance of 254.4 MFlops/$. This corrections is performed by counting the FLOPS based on the most efficient CPU algorithm. Any extra FLOPS that arise from the GPU implementation and parameter differences are not included in the 254.4 MFLOPS/$. Tsuyoshi Hamada, Keigo Nitadori |
SC | 1 |
| 2009 | Bayesian Multi-topic Microarray Analysis with Hyperparameter Reestimation
Tomonari Masada, Tsuyoshi Hamada, Yuichiro Shibata, Kiyoshi Oguri |
ADMA | 2 |
| 2009 | Dynamic hyperparameter optimization for bayesian topical trend analysisabstractThis paper presents a new Bayesian topical trend analysis. We regard the parameters of topic Dirichlet priors in latent Dirichlet allocation as a function of document timestamps and optimize the parameters by a gradient-based algorithm. Since our method gives similar hyperparameters to the documents having similar timestamps, topic assignment in collapsed Gibbs sampling is affected by timestamp similarities. We compute TFIDF-based document similarities by using a result of collapsed Gibbs sampling and evaluate our proposal by link detection task of Topic Detection and Tracking. Tomonari Masada, Daiji Fukagawa, Atsuhiro Takasu, Tsuyoshi Hamada, Yuichiro Shibata, Kiyoshi Oguri |
CIKM | 4 |
| 2009 | Accelerating Collapsed Variational Bayesian Inference for Latent Dirichlet Allocation with Nvidia CUDA Compatible Devices
Tomonari Masada, Tsuyoshi Hamada, Yuichiro Shibata, Kiyoshi Oguri |
IEA/AIE | 2 |
| 2009 | 42 TFlops hierarchical N-body simulations on GPUs with applications in both astrophysics and turbulenceabstractAs an entry for the 2009 Gordon Bell price/performance prize, we present the results of two different hierarchical N-body simulations on a cluster of 256 graphics processing units (GPUs). Unlike many previous N-body simulations on GPUs that scale as O(N2), the present method calculates the O(N log N) treecode and O(N) fast multipole method (FMM) on the GPUs with unprecedented efficiency. We demonstrate the performance of our method by choosing one standard application --a gravitational N-body simulation-- and one non-standard application --simulation of turbulence using vortex particles. The gravitational simulation using the treecode with 1,608,044,129 particles showed a sustained performance of 42.15 TFlops. The vortex particle simulation of homogeneous isotropic turbulence using the periodic FMM with 16,777,216 particles showed a sustained performance of 20.2 TFlops. The overall cost of the hardware was 228,912 dollars. The maximum corrected performance is 28.1TFlops for the gravitational simulation, which results in a cost performance of 124 MFlops/$. This correction is performed by counting the Flops based on the most efficient CPU algorithm. Any extra Flops that arise from the GPU implementation and parameter differences are not included in the 124 MFlops/$. Tsuyoshi Hamada, Tetsu Narumi, Rio Yokota, Kenji Yasuoka, Keigo Nitadori, Makoto Taiji |
SC | 1 |
| 2005 | Massively Parallel Processors Generator for Reconfigurable SystemabstractWe have developed PGR (processors generator for reconfigurable system) package which generate (a) a suitable configuration file for the FPGAs, (b) the C source code for interfacing with an FPGA-based accelerator, and (c) a software emulator from a high-level domain specific language. Using PGR package, we can easily produce high performance implementations for the particle-based simulation. Tsuyoshi Hamada, Naohito Nakasato |
FCCM | 1 |
| 2005 | Astrophysical Hydrodynamics Simulations on a Reconfigurable SystemabstractThe smoothed particle hydrodynamics (SPH) method is a widely used particle simulation scheme in astrophysical hydrodynamics simulations. Since a possible problem size of SPH simulations is limited by available computational resource, it is natural to implement a computation intensive core of the SPH simulations on a FPGA based reconfigurable board. We are now implementing SPH simulations on our FPGA based board PROGRAPE-3. Using our developed software, implementation of the SPH simulation becomes quite straightforward. Our "SPH processors", which compute particle interaction force by reduced accuracy floating-point operations, run at 66 MHz on our board. Obtained real performance of the SPH processors is 38 Gflops. Our results indicate using FPGA chips for floating-point intensive astrophysical simulations is being feasible. Naohito Nakasato, Tsuyoshi Hamada |
FCCM | 2 |
| 2005 | PGR: A Software Package for Reconfigurable Super-ComputingabstractIn this paper, we describe a methodology for implementing FPGA-based accelerator (FBA) from a high-level specification language. We have constructed a software package specially tuned for accelerating particle-based scientific computations with an FBA. Our software generates (a) a suitable configuration for the FPGA, (b) the C source code for interfacing with the FBA, and (c) a software emulator. The FPGA configuration is build by combining components from a library of parametrized arithmetic modules; these modules implement fixed-point, floating-point and logarithmic number system with flexible bitwidth and pipeline stages. To make certain our methodology is effective, we have applied our methodology to acceleration of astrophysical N-body application with two types of platforms. One is our PROGRAPE-3 with four XC2VP70-5 FPGAs and another is a minimum composition of CRAY-XD1 with one XC2VP50-7 FPGA. As the result, we have achieved peak performance of 324 Gflops with PROGRAPE-3 and 45 Gflops with the minimum CRAY-XD1, sustained performance of 236 Gflops with PROGRAPE-3 and 34 Gflops with the CRAY-XD1. Tsuyoshi Hamada, Naohito Nakasato |
FPL | 1 |
| 1998 | PROGRAPE-1: A Programmable Special-Purpose Computer for Many-Body SimulationsabstractWe have completed PROGRAPE-1, a programmable special-purpose computer for many-body simulations using FPGA (Field Programmable Gate Array). It has pipelines specialized for computations of interactions between particles. The peak performance of calculating gravitational/Coulomb force results in 2.4 Gflops. Tsuyoshi Hamada, Toshiyuki Fukushige, Atsushi Kawai, Junichiro Makino |
FCCM | 1 |