EDBT 2026 Demo / reviewers in the wild / expert
Yukihisa Fujita
dblp:59/2660
· DBLP profile ↗
6ranked-venue papers in the field
2as first author
6since 2021 · last 2025
0000-0002-0581-5116ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 6 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLM-Based Multi-Agent System for Simulating Strategic and Goal-Oriented Data Marketplaces
Jun Sashihara, Yukihisa Fujita, Kota Nakamura, Masahiro Kuwahara, Teruaki Hayashi |
IEEE Big Data | 2 |
| 2024 | Inferring Relationships between Tabular Data and Topics using LLM for a Dataset Search TaskabstractIn the big data era, data is called new oil and essential things in both business and academic fields. Both private and public data are continuously increasing and it is expected that such data are used for innovation and business improvement. However, finding datasets for specific purposes becomes a new challenge with the increase in the variety of data. To support the finding data task, there are some studies for dataset search; however, they are immature because of the lack of appropriate evaluation datasets, which consist of sets of queries, datasets, and labels, such as other machine learning areas. One of the major obstacles is that annotating the labels to a pair of queries and datasets requires specialized expertise and a substantial investment of time and effort. In this study, we aim to automate the data annotation process using a large language model (LLM) to reduce the manual effort required for creating annotated datasets for dataset search tasks. We propose an annotation framework that consists of LLM annotation with table compression and multiple results aggregation mechanism. We evaluated the proposed framework by human annotated datasets previously published for the dataset search challenge. The evaluation results show that the proposed framework outperformed the baseline method, which input simply a specific number of head rows to LLM. Moreover, our qualitative analysis gives insights into improving LLM-based systems and replacing human annotators. Yukihisa Fujita, Teruaki Hayashi, Masahiro Kuwahara |
IEEE Big Data | 1 |
| 2024 | Metadata-less Dataset Recommendation Leveraging Dataset Embeddings by Pre-trained Tabular Language ModelsabstractThe acceleration of data-driven business and research through the use of third-party datasets has led to the emergence of data platforms that enable cross-disciplinary data exchange, thereby increasing the need for dataset recommendations. In this climate, traditional retrieval systems, such as explanatory information in metadata, have limitations in terms of the reliability of metadata descriptions and their creation cost. This study proposes a method for recommending datasets by leveraging actual dataset information without relying on meta-data. We performed two metric learning methods, unsupervised contrastive learning and table-query metric learning, utilizing pre-trained tabular language models (TaLMs). The experimental results suggest that these metric learning methods can obtain representations that are more consistent with the labels representing the topics of the datasets. Moreover, our embeddings without metadata performed as well as those with metadata, suggesting that appropriate data extraction and clustering can be performed even in cases where metadata are sparse or incomplete. Kosuke Manabe, Yukihisa Fujita, Masahiro Kuwahara, Teruaki Hayashi |
IEEE Big Data | 2 |
| 2023 | Topic-Based Search: Dataset Search without Metadata and Users' Knowledge about DataabstractWith the advancement of information technologies, we can obtain various kinds of data, which can be leveraged for various purposes. The availability of a large amount of data is a desirable situation. However, it makes dataset retrieval a time-consuming and complex task. Conventional dataset search methods require unified metadata and knowledge about keywords representing the datasets. In other words, they require user knowledge regarding the datasets, such as the terms used in the dataset and fields in the metadata. To address this issue, we propose a topic-based search method without metadata, especially for users lacking knowledge about the datasets. The topic-based search can find datasets by using not the exact keywords but abstract keywords described as topics. In this paper, we focus on table data, which contain column names and data values and are widely used for storing data. As preliminary analysis, we collected and analyzed public datasets available in Japanese data portals to clarify the features of datasets that should be searched through dataset search. The analysis results revealed the use of many general and common keywords as column names, but it is difficult to implement a dataset search using only column names. Therefore, based on the analysis results, we decided to use embeddings converted from the datasets to utilize both column names and data values to extract topics from datasets. The experimental results showed that we can extract topics from datasets by using the topic modeling method and obtain better search results when compared with the search method using exact keywords. Yukihisa Fujita, Teruaki Hayashi, Masahiro Kuwahara |
IEEE Big Data | 1 |
| 2023 | Exploring the Fundamental Units of Semantic Representation of Data Using Heterogeneous Variable Network in Data EcosystemsabstractThe value creation achieved through the exchange, distribution, and collaboration of data among different organizations has garnered significant attention as a new source of innovation. The mathematical treatment of the meaning of data helps measure its “quality” to formulate evaluation criteria for data exchange between stakeholders with distinct background knowledge in data ecosystems. This study examines the structure of data morphemes, the fundamental units of semantic representation of data, by conducting network and association analyses of variables present in metadata from diverse fields. Network analysis identifies the globally sparse and locally dense characteristics of variable co-occurrence networks and highlights essential relationships and core variables. Key findings include the discovery of “depth,” “sediment/rock,” and “sample code/label” as both universal variables and crucial nodes between datasets used in the experiment. Association analysis reveals vital variable pairs, such as “age” and “ring width” or “latitude” and “longitude.” This research may provide a understanding of the structure and meaningful representation of data, facilitating smooth data exchange and utilization practices among stakeholders with different domains, purposes of data use, and background knowledge in data ecosystems. Teruaki Hayashi, Yukihisa Fujita, Masahiro Kuwahara |
IEEE Big Data | 2 |
| 2023 | Variable-based Learning Considering Topic Specificity in Heterogeneous Data Clustering TasksabstractRecently, data mining via interdisciplinary co-creation has attracted considerable social attention, and various data have been published, for free or for a fee. Data publication encourages the exchange and combination of data between different institutions, which is helpful for interdisciplinary data collaboration. However, issues pertaining to designing high-quality data for interdisciplinary data discovery remain in data search. Variables are frameworks for data, and reflect the data topics and intent of the data design. In this study, the relationships between data topics and variables for a large dataset were quantitatively investigated to provide suggestions for data design and exploration. The probability of occurrence of variables and their pairs for each topic was determined to elucidate the relationship between the topics and variables; subsequently, clustering was applied based on these relationships. Kosuke Manabe, Yukihisa Fujita, Masahiro Kuwahara, Teruaki Hayashi |
IEEE Big Data | 2 |