Kosuke Manabe

dblp:355/0786 · DBLP profile ↗
← Back
2ranked-venue papers in the field
2as first author
2since 2021 · last 2024
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2 (2 first)
YearPublicationVenuePosition
2024 Metadata-less Dataset Recommendation Leveraging Dataset Embeddings by Pre-trained Tabular Language Models
abstract
The acceleration of data-driven business and research through the use of third-party datasets has led to the emergence of data platforms that enable cross-disciplinary data exchange, thereby increasing the need for dataset recommendations. In this climate, traditional retrieval systems, such as explanatory information in metadata, have limitations in terms of the reliability of metadata descriptions and their creation cost. This study proposes a method for recommending datasets by leveraging actual dataset information without relying on meta-data. We performed two metric learning methods, unsupervised contrastive learning and table-query metric learning, utilizing pre-trained tabular language models (TaLMs). The experimental results suggest that these metric learning methods can obtain representations that are more consistent with the labels representing the topics of the datasets. Moreover, our embeddings without metadata performed as well as those with metadata, suggesting that appropriate data extraction and clustering can be performed even in cases where metadata are sparse or incomplete.
Kosuke Manabe, Yukihisa Fujita, Masahiro Kuwahara, Teruaki Hayashi
IEEE Big Data1
2023 Variable-based Learning Considering Topic Specificity in Heterogeneous Data Clustering Tasks
abstract
Recently, data mining via interdisciplinary co-creation has attracted considerable social attention, and various data have been published, for free or for a fee. Data publication encourages the exchange and combination of data between different institutions, which is helpful for interdisciplinary data collaboration. However, issues pertaining to designing high-quality data for interdisciplinary data discovery remain in data search. Variables are frameworks for data, and reflect the data topics and intent of the data design. In this study, the relationships between data topics and variables for a large dataset were quantitatively investigated to provide suggestions for data design and exploration. The probability of occurrence of variables and their pairs for each topic was determined to elucidate the relationship between the topics and variables; subsequently, clustering was applied based on these relationships.
Kosuke Manabe, Yukihisa Fujita, Masahiro Kuwahara, Teruaki Hayashi
IEEE Big Data1