EDBT 2026 Demo / reviewers in the wild / expert
Teruaki Hayashi
dblp:134/3473
· DBLP profile ↗
19ranked-venue papers in the field
5as first author
15since 2021 · last 2025
0000-0002-1806-5852ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 18 (4 first)Data Mining & Knowledge Discovery · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Contextual Graph Embeddings: Accounting for Data Characteristics in Heterogeneous Data Integration
Yuka Haruki, Shigeru Ishikura, Kazuya Demachi, Teruaki Hayashi |
IEEE Big Data | 4 |
| 2025 | LLM-Based Multi-Agent System for Simulating Strategic and Goal-Oriented Data Marketplaces
Jun Sashihara, Yukihisa Fujita, Kota Nakamura, Masahiro Kuwahara, Teruaki Hayashi |
IEEE Big Data | 5 |
| 2025 | Designing Reputation Systems for Manufacturing Data Trading Markets: A Multi-Agent Evaluation With Q-Learning and IRL-Estimated Utilities
Kenta Yamamoto, Teruaki Hayashi |
IEEE Big Data | 2 |
| 2025 | SHAP Distance: An Explainability-Aware Metric for Evaluating the Semantic Fidelity of Synthetic Tabular Data
Shigeru Ishikura, Yukari Usukura, Yuki Shigoku, Teruaki Hayashi |
IEEE Big Data | 5 |
| 2024 | Inferring Relationships between Tabular Data and Topics using LLM for a Dataset Search TaskabstractIn the big data era, data is called new oil and essential things in both business and academic fields. Both private and public data are continuously increasing and it is expected that such data are used for innovation and business improvement. However, finding datasets for specific purposes becomes a new challenge with the increase in the variety of data. To support the finding data task, there are some studies for dataset search; however, they are immature because of the lack of appropriate evaluation datasets, which consist of sets of queries, datasets, and labels, such as other machine learning areas. One of the major obstacles is that annotating the labels to a pair of queries and datasets requires specialized expertise and a substantial investment of time and effort. In this study, we aim to automate the data annotation process using a large language model (LLM) to reduce the manual effort required for creating annotated datasets for dataset search tasks. We propose an annotation framework that consists of LLM annotation with table compression and multiple results aggregation mechanism. We evaluated the proposed framework by human annotated datasets previously published for the dataset search challenge. The evaluation results show that the proposed framework outperformed the baseline method, which input simply a specific number of head rows to LLM. Moreover, our qualitative analysis gives insights into improving LLM-based systems and replacing human annotators. Yukihisa Fujita, Teruaki Hayashi, Masahiro Kuwahara |
IEEE Big Data | 2 |
| 2024 | Metadata-based Data Exploration with Retrieval-Augmented Generation for Large Language ModelsabstractDeveloping the capacity to effectively search for requisite datasets is an urgent requirement to assist data users in identifying relevant datasets considering the very limited available metadata. For this challenge, the utilization of third-party data is emerging as a valuable source for improvement. Our research introduces a new architecture for data exploration which employs a form of Retrieval-Augmented Generation (RAG) to enhance metadata-based data discovery. The system integrates large language models (LLMs) with external vector databases to identify semantic relationships among diverse types of datasets. The proposed framework offers a new method for evaluating semantic similarity among heterogeneous data sources and for improving data exploration. Our study includes experimental results on four critical tasks: 1) recommending similar datasets, 2) suggesting combinable datasets, 3) estimating tags, and 4) predicting variables. Our results demonstrate that RAG can enhance the selection of relevant datasets, particularly from different categories, when compared to conventional metadata approaches. However, performance varied across tasks and models, which confirms the significance of selecting appropriate techniques based on specific use cases. The findings suggest that this approach holds promise for addressing challenges in data exploration and discovery, although further refinement is necessary for estimation tasks. Teruaki Hayashi, Hiroki Sakaji, Jiayi Dai, Randy Goebel |
IEEE Big Data | 1 |
| 2024 | Metadata-less Dataset Recommendation Leveraging Dataset Embeddings by Pre-trained Tabular Language ModelsabstractThe acceleration of data-driven business and research through the use of third-party datasets has led to the emergence of data platforms that enable cross-disciplinary data exchange, thereby increasing the need for dataset recommendations. In this climate, traditional retrieval systems, such as explanatory information in metadata, have limitations in terms of the reliability of metadata descriptions and their creation cost. This study proposes a method for recommending datasets by leveraging actual dataset information without relying on meta-data. We performed two metric learning methods, unsupervised contrastive learning and table-query metric learning, utilizing pre-trained tabular language models (TaLMs). The experimental results suggest that these metric learning methods can obtain representations that are more consistent with the labels representing the topics of the datasets. Moreover, our embeddings without metadata performed as well as those with metadata, suggesting that appropriate data extraction and clustering can be performed even in cases where metadata are sparse or incomplete. Kosuke Manabe, Yukihisa Fujita, Masahiro Kuwahara, Teruaki Hayashi |
IEEE Big Data | 4 |
| 2024 | Impact of Buyer Strategy and Market Size on Data Marketplace Dynamics: A Network-Based Simulation StudyabstractOwing to the requisite costs and labor for data collection and analysis, the purchase and exchange of thirdparty data in the market has become a viable option for many institutions. However, the debate on the dynamic interactions among various types of data, regulations, and market participants remains unresolved. Thus, this study examined the impacts of buyer behaviors and market size on data distribution and pricing dynamics in data marketplaces. This study developed a model on data provided by marketplaces through a random walk on the co-occurrence networks of variables to reflect the characteristics of real-world data networks. Buyer agents with three purchasing strategies (Random, Related, and Ranking) were simulated across six scenarios by varying the number of datasets and buyers. The experimental results demonstrated that the proposed model generated data networks that reflected the principal characteristics of actual data marketplaces. Markets, wherein buyers acquired data related to their existing data, yielded increasingly prevalent datasets with more variables, resulting in marketplace imbalances. The Random strategy resulted in the highest utility for buyers across all scenarios, whereas the Related strategy led to the highest prices for popular datasets. As the number of buyers increased relative to the available datasets, disparities in the distribution based on the strategy became more pronounced. The strategic adjustment of buyer approaches in accordance with the market size could mitigate the emergence of disproportionately popular data and promote price stability. Thus, this study provides insights for data marketplace operators and institutional design through the modeling of realistic marketplace dynamics and an analysis of the outcomes across different market conditions. Jun Sashihara, Teruaki Hayashi |
IEEE Big Data | 2 |
| 2024 | Understanding User Interactions and Community Formation on Data Competition PlatformabstractAs increasing volumes of data have become available globally, there are growing expectations for the creation of value through the exchange of data across diverse fields and the enhancement of the value of existing services. Consequently, there is an increasing demand for markets and platforms to facilitate data trading and exchanges. However, unlike other well-established business ecosystems such as existing financial markets and service ecosystems, the overall structure and characteristics of the data exchange ecosystem remain largely unexplored. This study focuses on users who handle data, and aims to elucidate their behaviors and identify influential users on the platforms, thereby deepening our understanding of the data ecosystem. We conducted network analysis, community extraction, and time-series analysis of the behavior of Kaggle users. Our results indicate that the network of Kaggle users exhibits a scale-free structure similar to typical social networks and web links. This suggests that information dissemination among Kaggle users is likely to occur through a small number of hubs. Furthermore, we performed community extraction, analyzed the behavior of users in each community, and discovered that each community possessed distinct user behavior characteristics. Subsequently, we analyzed the number of influential users with numerous followers or high degree centrality and examined the time-series changes in the number of actions taken by these users. This analysis enabled us to identify the behaviors and formation patterns of those who played a central role in the data platform. The approach and findings of our study provide important considerations for stakeholders and potential participants in all the markets and platforms that handle data. Kenta Yamamoto, Teruaki Hayashi |
IEEE Big Data | 2 |
| 2023 | Topic-Based Search: Dataset Search without Metadata and Users' Knowledge about DataabstractWith the advancement of information technologies, we can obtain various kinds of data, which can be leveraged for various purposes. The availability of a large amount of data is a desirable situation. However, it makes dataset retrieval a time-consuming and complex task. Conventional dataset search methods require unified metadata and knowledge about keywords representing the datasets. In other words, they require user knowledge regarding the datasets, such as the terms used in the dataset and fields in the metadata. To address this issue, we propose a topic-based search method without metadata, especially for users lacking knowledge about the datasets. The topic-based search can find datasets by using not the exact keywords but abstract keywords described as topics. In this paper, we focus on table data, which contain column names and data values and are widely used for storing data. As preliminary analysis, we collected and analyzed public datasets available in Japanese data portals to clarify the features of datasets that should be searched through dataset search. The analysis results revealed the use of many general and common keywords as column names, but it is difficult to implement a dataset search using only column names. Therefore, based on the analysis results, we decided to use embeddings converted from the datasets to utilize both column names and data values to extract topics from datasets. The experimental results showed that we can extract topics from datasets by using the topic modeling method and obtain better search results when compared with the search method using exact keywords. Yukihisa Fujita, Teruaki Hayashi, Masahiro Kuwahara |
IEEE Big Data | 2 |
| 2023 | Exploring the Fundamental Units of Semantic Representation of Data Using Heterogeneous Variable Network in Data EcosystemsabstractThe value creation achieved through the exchange, distribution, and collaboration of data among different organizations has garnered significant attention as a new source of innovation. The mathematical treatment of the meaning of data helps measure its “quality” to formulate evaluation criteria for data exchange between stakeholders with distinct background knowledge in data ecosystems. This study examines the structure of data morphemes, the fundamental units of semantic representation of data, by conducting network and association analyses of variables present in metadata from diverse fields. Network analysis identifies the globally sparse and locally dense characteristics of variable co-occurrence networks and highlights essential relationships and core variables. Key findings include the discovery of “depth,” “sediment/rock,” and “sample code/label” as both universal variables and crucial nodes between datasets used in the experiment. Association analysis reveals vital variable pairs, such as “age” and “ring width” or “latitude” and “longitude.” This research may provide a understanding of the structure and meaningful representation of data, facilitating smooth data exchange and utilization practices among stakeholders with different domains, purposes of data use, and background knowledge in data ecosystems. Teruaki Hayashi, Yukihisa Fujita, Masahiro Kuwahara |
IEEE Big Data | 1 |
| 2023 | Variable-based Learning Considering Topic Specificity in Heterogeneous Data Clustering TasksabstractRecently, data mining via interdisciplinary co-creation has attracted considerable social attention, and various data have been published, for free or for a fee. Data publication encourages the exchange and combination of data between different institutions, which is helpful for interdisciplinary data collaboration. However, issues pertaining to designing high-quality data for interdisciplinary data discovery remain in data search. Variables are frameworks for data, and reflect the data topics and intent of the data design. In this study, the relationships between data topics and variables for a large dataset were quantitatively investigated to provide suggestions for data design and exploration. The probability of occurrence of variables and their pairs for each topic was determined to elucidate the relationship between the topics and variables; subsequently, clustering was applied based on these relationships. Kosuke Manabe, Yukihisa Fujita, Masahiro Kuwahara, Teruaki Hayashi |
IEEE Big Data | 4 |
| 2022 | A Model of Pricing Data and Their Constituent Variables Traded in Two-Sided Markets with Resale: A Subject ExperimentabstractThis note presents a simple model of pricing data and their constituent variables traded in two-sided markets, where resale of data is allowed. The prices of those variables are exogenously set at the initial round, and in each round those prices are updated for the next round immediately at the end of the round, based on the outcomes of data transactions. If traders behave in accordance with the backward induction, then the initial prices never move in any rounds. In the subject experiment, this property was not observed but the average prices of those variables were not far from the initial values. We also examined whether information provision of the gross profits the user and non-user receive from data transactions affects the social welfare measured by the amounts of producer surplus. Toshihiko Nanba, Kazuhito Ogawa, Naoki Watanabe 0001, Teruaki Hayashi, Hiroki Sakaji |
IEEE Big Data | 4 |
| 2021 | Growing Process of Communities on Data Platforms: Case Analysis of a COVID-19 DatasetabstractIn recent years, there have been growing expectations for the creation of new businesses and the improvement of the value of existing services by exchanging data in different fields. Data stored in-house within organizations have become a new source of innovation. While there is a high need for the value creation of data, determining the data value is not an easy task, as there is a wide range of factors to be considered, such as data pricing, acquisition cost, usage value, and update frequency. In this study, we observe communication, such as the sharing of know-hows in data exchange and analysis, and discuss the growing process of a community on the data platform. For the experiment, we focused on the data community in the COVID-19 disaster and used a unique dataset from the data platform Kaggle, which is the data analysis competition service. The results suggest that user actions differ in the discussion of the dataset and analysis. Moreover, providing topics, user participation, and activating actions in the early stages after the dataset is released are essential for forming a data community. We argue that the actions on the data analysis, such as comments and votes, are also crucial for fostering a common understanding of the data value. Teruaki Hayashi, Takumi Shimizu, Yoshiaki Fukami, Hiroki Sakaji, Hiroyasu Matsushima |
IEEE BigData | 1 |
| 2021 | Retrieving of Data Similarity using Metadata on a Data Analysis Competition PlatformabstractIn recent years, instead of closing data and analysis skills in-house, there has been much interest in widely releasing data analysis knowledge on the web. A data exchange platform is a type of digital platform that exchanges data between stakeholders, e.g., data owners, users, and analysts. However, the datasets handled on such platforms are independently acquired and stored by the data providers for their own purposes. These datasets are not based on the premise of coordination and combination, and there is currently little information available to discuss the systematic organization and combination of these datasets. In this study, we focus on a metadata, summary information of data, and examine the similarity of data on a data exchange platform using natural language processing. In our experiments, we use the metadata from the data exchange platform Kaggle. To compare the similarity of the data, our method employs word2vec and BERT as vectorize methods and converts data descriptions to vectors. Then, our method measures the distances of each vector by calculating cosine similarities between each vector. From experimental results, we found that Kaggle has the same character as other data exchange platforms. Additionally, the results indicated the usability of the natural language processing-based method for extracting similar data pairs. Hiroki Sakaji, Teruaki Hayashi, Yoshiaki Fukami, Takumi Shimizu, Hiroyasu Matsushima, Kiyoshi Izumi |
IEEE BigData | 2 |
| 2020 | Data Requests and Scenarios for Data Design of Unobserved Events in Corona-related Confusion Using TEEDAabstractDue to the global violence of the novel coronavirus, various industries have been affected and the breakdown between systems has been apparent. To understand and overcome the phenomenon related to this unprecedented crisis caused by the coronavirus infectious disease (COVID-19), the importance of data exchange and sharing across fields has gained social attention. In this study, we use the interactive platform called treasuring every encounter of data affairs (TEEDA) to externalize data requests from data users, which is a tool to exchange not only the information on data that can be provided but also the call for data-what data users want and for what purpose. Further, we analyze the characteristics of missing data in the corona-related confusion stemming from both the data requests and the providable data obtained in the workshop. We also create three scenarios for the data design of unobserved events focusing on variables. Teruaki Hayashi, Nao Uehara, Daisuke Hase, Yukio Ohsawa |
IEEE BigData | 1 |
| 2020 | Verification of Data Similarity using Metadata on a Data Exchange PlatformabstractWith the development of computers and the rise of data exchange, the expectations for innovation by combining data from different industries, and data exchange platforms that handle different types of data are sprouting up. However, the data handled on such platforms have been obtained and are stored independently by data providers with different purposes, and the maintenance of data catalogs and the spread of schemata are not currently insufficient, making it difficult to understand the relationships between the data on such platforms. In this study, in order to derive the relationships among datasets and to discuss the similarity and combinability of the data on the data exchange platform, we analyze and discuss the relationships between datasets on the platform. We focussed on the data outlines and variables using the metadata of the data exchange platform service, D-Ocean, and found that the similarity of data cannot be measured by a single indicator, and that the items to be referred to calculate the similarity depending on the indicator. Hiroki Sakaji, Teruaki Hayashi, Kiyoshi Izumi, Yukio Ohsawa |
IEEE BigData | 2 |
| 2020 | Hierarchical Graph Convolutional Network for Data Evaluation of Dynamic GraphsabstractAs data are being generated at an incredible speed and scale, the market of data that aims to fully discover the value of data is becoming increasingly important. Data evaluation is a vital part of the data market. It can discover the data's flaws and the value hidden in the data. Nevertheless, there have been few studies investigating how to evaluate data in the data market. Existing methods seldom utilize the multi-level structure in data. Our work proposes a novel hierarchical graph convolutional network for the data evaluation of dynamic graphs, following the anomaly detection paradigm. Our model performs significantly better than existing models on several benchmark datasets. This study provides a more powerful tool for processing dynamic graphs. It could also guide the direction of data evaluation for a broader range of data categories in the data market. Teruaki Hayashi, Yukio Ohsawa |
IEEE BigData | 2 |
| 2017 | Matrix-Based Method for Inferring Variable Labels Using Outlines of Data in Data Jackets
Teruaki Hayashi, Yukio Ohsawa |
PAKDD (2) | 1 |