Kazuo Goda

dblp:15/628 · DBLP profile ↗
← Back
25ranked-venue papers in the field
3as first author
14since 2021 · last 2026
0000-0003-0618-4157ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 20 (2 first)Big Data, Cloud & Distributed Data Systems · 4Information Retrieval & Web Search · 1 (1 first)
YearPublicationVenuePosition
2026 Adaptive GPU Compute Resource Allocation for Efficient High-Utility Itemset Mining
Tarun Sreepada, Tsuyoshi Ozawa, Genki Kimura, R. Uday Kiran, Kazuo Goda
DASFAA (5)5
2026 Improving IO Efficiency for Storage-Scale Approximate Nearest Neighbor Search
Guoxing Wu, Tsuyoshi Ozawa, Yuto Hayamizu, Kazuo Goda
DEXA (2)4
2026 Scalable Concurrency Control for R-Trees on Three-Tier Storage with PMEM
Hirotaka Yoshioka, Kazuo Goda, Masaru Kitsuregawa
DEXA (2)2
2026 DPHIM: Efficient Parallel Mining of High-Utility Itemsets on Multicore Processors and Its Evaluation
abstract
High-utility itemset mining (HUIM) is an advanced problem of frequent itemset mining, considering the frequency of occurrence and quantitative criteria such as unit profit. Because HUIM can be applied to a broad spectrum of knowledge discovery work, various algorithmic improvements have been studied over the past two decades. On the other hand, limited efforts have been made to take advantage of hardware performance despite significant changes in hardware trends. This paper presents a novel parallelization method called DPHIM (Dynamic Parallelization for High-utility Itemset Mining). DPHIM dynamically decomposes a high-utility itemset mining task into subtasks to utilize logical parallelism and carefully assigns the subtasks and their related data to physical resources such as processing cores and nearby memory in a NUMA-aware manner. Through rigorous and diverse experiments, we found that DPHIM achieved speeds up to 72.7 times faster than the fully tuned serial execution, up to 23.5 times faster than static partitioning, and up to 2.5 times faster than the best case of alternative dynamic parallel executions for a variety of datasets and configurations on DRAM. We also demonstrated that DPHIM effectively worked on persistent memory; it offered similar thread scalability trends and was 1.1 to 2.4 times slower on persistent memory.
Genki Kimura, Yuto Hayamizu, R. Uday Kiran, Masaru Kitsuregawa, Kazuo Goda
IEEE Trans. Knowl. Data Eng.5
2025 Accelerating Fuzzy Frequent Pattern Mining on GPUs with GPU Direct Storage
Tarun Sreepada, Arjun Chakravarthi Pogaku, R. Uday Kiran, Kazuo Goda
IEEE Big Data4
2025 Incremental k-Anonymization for Continuously Growing Big Databases
Akifumi Kurumatani, Hiromasa Yoshimoto, Kazuo Goda
DEXA (2)3
2024 LakeHarbor: Making Structures First-Class Citizens in Data Lakes
abstract
This paper introduces LakeHarbor, a new data management paradigm that makes structures (e.g., indexes) first-class citizens in data lakes. The LakeHarbor paradigm enables a data lake system to flexibly construct structures based on registered access method functions and execute data processing jobs efficiently with the potential parallelism that the structures inherently hold by exploiting the functions while not sacrificing flexible data processing such as schema-on-read. This paper also presents ReDe, a prototype data processing engine that implements LakeHarbor, and a motivating evaluation and a case study of ReDe to explore the potential of LakeHarbor.
Hiroyuki Yamada, Masaru Kitsuregawa, Kazuo Goda
ICDE3
2023 Physical Database Design for Manufacturing Business Analytics
abstract
The manufacturing business field is accommodating a massive number of networked sensors, which are offering a new horizon of microscopic observability; every single piece of the production line is enabled to be monitored, transformed and organized into collective digital assets. Intensive analytics on the digital assets potentially offers new business solutions to improve the manufacturing productivity and secure the business continuity. For example, the capability of exhaustive production inspection would potentially remove the delivery of defective products. Another scenario is the capability of expeditiously identifying an impaired part in the production line, which would significantly reduce the yield loss. These business analytics work often hold the nature that they need to precisely identify a particular manufacturing event of interest (e.g., producing an intermediate material at a suspiciously faulty machine) and then trace back or forward every ancestor or descendant of the event thoroughly in the manufacturing pipeline spanning from the initial raw material to the final product. Yet, such microscopic business traceability has not been actively addressed in the technology community. This paper presents our experimental exploration of how to design relational database for manufacturing business analytics, which is technically differentiated by the involvement of microscopic business traceability from traditional business analytics. We have performed intensive experiments and cost-performance analyses to figure out the trade-offs associated with major options of physical database design with different benchmarks on a commercial database management system (DBMS) in an enterprise environment. The combination of columnar table stores and secondary indexes is recommended for manufacturing business analytics workloads, even though the pure use of columnar table stores is recommended for traditional analytics workloads. We believe that our experience would help those planning to work on microscopic business tractability in manufacturing and other fields.
Norifumi Nishikawa, Shinji Fujiwara, Yuto Hayamizu, Kazuo Goda
IEEE Big Data4
2023 Efficient Parallel Mining of High-utility Itemsets on Multicore Processors
abstract
High-utility itemset mining is a generalized problem of well-known frequent itemset mining, which considers not only the frequency of occurrence but also quantitative criteria such as unit profit. Because it can be applied to a wider spectrum of knowledge discovery work, various algorithmic improvements have been studied over the past two decades. On the other hand, limited efforts have been made to take advantage of hardware performance despite significant changes in hardware trends. This paper presents a novel parallelization method called DPHIM (Dynamic Parallelization for High-utility Itemset Mining). DPHIM dynamically decomposes the execution of high-utility itemset mining into subtasks in order to leverage logical data parallelism, and carefully assigns the subtasks and their related data to physical resources such as processing cores and nearby memory in the NUMA-aware manner. Our intensive and extensive experiments have confirmed that DPHIM performs up to 65.23 times faster than the fully-tuned serial execution, up to 23.54 times faster than static partitioning, and up to 2.51 times faster than the best case of alternative dynamic parallel executions for a variety of datasets and configurations on DRAM. As well, we have demonstrated that DPHIM effectively worked on persistent memory; it offered similar thread scalability trends and was 1.07 to 2.43 times slower on persistent memory.
Genki Kimura, Yuto Hayamizu, R. Uday Kiran, Masaru Kitsuregawa, Kazuo Goda
ICDE5
2023 Efficient Parallel Mining of High-utility Itemsets on Multicore Processors
abstract
High-utility itemset mining is a generalized problem of well-known frequent itemset mining, which considers not only the frequency of occurrence but also quantitative criteria such as unit profit. Because it can be applied to a wider spectrum of knowledge discovery work, various algorithmic improvements have been studied over the past two decades. On the other hand, limited efforts have been made to take advantage of hardware performance despite significant changes in hardware trends. This paper presents a novel parallelization method called DPHIM (Dynamic Parallelization for High-utility Itemset Mining). DPHIM dynamically decomposes the execution of high-utility itemset mining into subtasks in order to leverage logical data parallelism, and carefully assigns the subtasks and their related data to physical resources such as processing cores and nearby memory in the NUMA-aware manner. Our intensive and extensive experiments have confirmed that DPHIM performs up to 65.23 times faster than the fully-tuned serial execution, up to 23.54 times faster than static partitioning, and up to 2.51 times faster than the best case of alternative dynamic parallel executions for a variety of datasets and configurations on DRAM. As well, we have demonstrated that DPHIM effectively worked on persistent memory; it offered similar thread scalability trends and was 1.07 to 2.43 times slower on persistent memory.
Genki Kimura, Yuto Hayamizu, R. Uday Kiran, Masaru Kitsuregawa, Kazuo Goda
ICDE5
2023 Nested Loops Revisited Again
abstract
Hash joins and sort-merge joins have been considered the algorithms of choice for analytical relational queries in most parallel database systems because of their performance robustness and ease of parallelization. On the other hand, nested loop joins have been considered less attractive and are conservatively used. In this paper, we revisit the potential of nested loop joins in a cluster environment. We focus on exploring the parallelism aspect of nested loop joins because there could still be space for improvement by fully exploiting the parallelism of current commodity hardware, which could handle more than thousands of concurrent IOs. We also introduce scalable massively-parallel execution as one of the approaches for achieving massive parallelism in nested loop joins to explore how it widens the potential benefit of nested loop joins. Finally, we discuss future research directions based on our exploration.
Hiroyuki Yamada, Kazuo Goda, Masaru Kitsuregawa
ICDE2
2022 A Novel GPU-Accelerated Algorithm to Discover Periodic-Frequent Patterns in Temporal Databases
abstract
Periodic-frequent pattern mining is a vital knowledge discovery technique that aims to find all regularly occurring patterns in a temporal database. Previous studies focused on developing CPU-centric algorithms by disregarding the speedups offered by the GPUs. Furthermore, existing GPU-based frequent pattern mining algorithms cannot be employed to find periodic-frequent patterns because they ignore the items' temporal occurrence information in the database, and the multi-threaded sum-reduction technique cannot be employed to determine the periodicity of a pattern in a database. With this motivation, this paper proposes an efficient GPU-accelerated depth-first search algorithm, GPU Periodic Frequent-Miner (gPF-Miner), to find the desired patterns. Our algorithm employs a novel flattened array structure to effectively record the temporal occurrence information of every item in a database. Our algorithm also introduces a new multi-threaded parallelization technique to calculate the support and periodicity of a pattern in a GPU. This technique’s best and worst-case time complexities are O(1) and O(n), where n represents the data size. Experimental results demonstrate that gPF-Miner outperforms the existing CPU-based and naive GPU-based algorithms by a vast margin.
Tarun Sreepada, R. Uday Kiran, Yutaka Watanobe, Kazuo Goda
IEEE Big Data4
2022 μ-join: Efficient Join with Versioned Dimension Tables
Mika Takata, Kazuo Goda, Masaru Kitsuregawa
DASFAA (1)2
2022 Exploiting Embedded Synopsis for Exact and Approximate Query Processing
Hiroki Yuasa, Kazuo Goda, Masaru Kitsuregawa
DEXA (2)2
2020 Discovering Closed Periodic-Frequent Patterns in Very Large Temporal Databases
abstract
Periodic-frequent pattern mining (PFPM) is an important data mining model having many real-world applications. However, this model's prosperous industrial use has been hindered by the problem of combinatorial explosion of patterns, which is the generation of too many redundant patterns, most of which may be useless to the user. We propose a novel model of closed periodic-frequent patterns that may exist in a temporal database to address this problem. Closed periodic-frequent patterns represent a concise lossless subset that uniquely preserves the complete information of all periodic-frequent patterns in a database. An efficient depth-first search algorithm, called Closed Periodic-Frequent Pattern Miner (CPFP-Miner), has been introduced to find all the database's desired patterns. Experimental results demonstrate that CPFP-Miner is not only memory, runtime, and energy-efficient, but also highly scalable. The usefulness of our model has also been shown with a case study on traffic congestion analytics.
Likhitha Palla, Penugonda Ravikumar, R. Uday Kiran, Yuto Hayamizu, Kazuo Goda, Masashi Toyoda, Koji Zettsu, Sourabh Shrivastava
IEEE BigData5
2020 PhoeniQ: Failure-Tolerant Query Processing in Multi-node Environments
Yutaro Bessho, Yuto Hayamizu, Kazuo Goda, Masaru Kitsuregawa
DEXA (1)3
2020 Out-of-order Execution of Database Queries
abstract
Intra-query parallelism is a key for database software to offer acceptable responsiveness for data-intensive queries. Many researchers have studied how to achieve greater execution parallelism for database queries. Partitioning is a representative approach, which divides a query into multiple sub-tasks and executes them in parallel. However, given a new query, optimal division is not necessarily obvious. Database software utilizes heuristic rules or statistical information to decide how to divide the query before execution. As yet another approach to achieve execution parallelism, this paper presents out-of-order database execution (OoODE), a massively-parallel query execution method to offer significant speedup for database queries consistently. OoODE dynamically decomposes query work by making the best use of the exact knowledge of the potential execution parallelism for each operation ready to be performed during query execution. With OoODE, the database software is allowed to automatically squeeze out the execution parallelism that the query inherently holds. Hence, for a wide spectrum of queries, OoODE performs significantly faster than the serial (non-parallelized) execution, while it performs better than or comparably with alternative parallelizing methods without the need for dividing the query before execution. This paper presents the experiments that we conducted using the prototyped database software and demonstrates that OoODE is two to three orders of magnitude faster than the serial execution, whereas it is substantially (up to 2.07 times) faster than the best achievable case of partitioning. Besides, OoODE performs two to four orders of magnitude faster than major DBMSs.
Kazuo Goda, Yuto Hayamizu, Hiroyuki Yamada, Masaru Kitsuregawa
Proc. VLDB Endow.1
2019 A Prescription Trend Analysis using Medical Insurance Claim Big Data
abstract
Understanding the spread of diseases and the use of medicines is of practical importance for various organizations, such as medical providers, medical payers, and national governments. This study aims to detect the change in the prescription trends and to identify its cause through an analysis of Medical Insurance Claims (MICs), which comprise the specifications of medical fees charged to health insurers. Our approach is two-fold. (1) We propose a latent variable model that simulates the medication behavior of physicians to accurately reproduce monthly prescription time series from the MIC data, where prescription links between the diseases and medicines are missing. (2) We apply a state space model with intervention variables to decompose the monthly prescription time series into different components including seasonality and structural changes. Using a large dataset consisting of 3.5-year MIC records, we conduct experiments to evaluate our approach in terms of accuracy, usefulness, and efficiency. We also demonstrate three applications for our medical analysis.
Kazutoshi Umemoto, Kazuo Goda, Naohiro Mitsutake, Masaru Kitsuregawa
ICDE2
2018 Modeling Query Energy Costs in Analytical Database Systems with Processor Speed Scaling
Boming Luo, Yuto Hayamizu, Kazuo Goda, Masaru Kitsuregawa
DEXA (2)3
2016 Aging Locality Awareness in Cost Estimation for Database Query Optimization
Chihiro Kato, Yuto Hayamizu, Kazuo Goda, Masaru Kitsuregawa
DEXA (2)3
2015 An Experimental Study of Aging Influence on Query Cost Estimation
abstract
Many update queries on a database can eventually degrade the structural efficiency of the database and result in lower performance. This phenomenon is called aging. On aged databases, conventional cost-based query optimizers could choose non-optimal query execution plan because they are not aging-aware and could not accurately estimate query execution cost.
Chihiro Kato, Yuto Hayamizu, Kazuo Goda, Masaru Kitsuregawa
IDEAS3
2009 Evaluating Non-In-Place Update Techniques for Flash-Based Transaction Processing Systems
Yongkun Wang, Kazuo Goda, Masaru Kitsuregawa
DEXA2
2003 Effective Load-Balancing via Migration and Replication in Spatial Grids
Anirban Mondal, Kazuo Goda, Masaru Kitsuregawa
DEXA2
2002 Run-Time Load Balancing System on SAN-connected PC Cluster for Dynamic Injection of CPU and Disk Resource - A Case Study of Data Mining Application
Kazuo Goda, Takayuki Tamura, Masato Oguchi, Masaru Kitsuregawa
DEXA1
2001 Query Optimization for Vector Space Problems
abstract
We present performance measurement results for a parallel SQL based information retrieval system implemented on a PC cluster system. We used the Web-TREC dataset under a left-deep query execution plan. We achieved satisfactory speed up.
Kazuo Goda, Masaru Kitsuregawa, Takayuki Tamura, Ophir Frieder, Abdur Chowdhury
SIGIR1