EDBT 2026 Demo / reviewers in the wild / expert
Kosuke Manabe
dblp:355/0786
· DBLP profile ↗
2ranked-venue papers in the field
2as first author
2since 2021 · last 2024
—ORCID · none
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 2 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Metadata-less Dataset Recommendation Leveraging Dataset Embeddings by Pre-trained Tabular Language ModelsabstractThe acceleration of data-driven business and research through the use of third-party datasets has led to the emergence of data platforms that enable cross-disciplinary data exchange, thereby increasing the need for dataset recommendations. In this climate, traditional retrieval systems, such as explanatory information in metadata, have limitations in terms of the reliability of metadata descriptions and their creation cost. This study proposes a method for recommending datasets by leveraging actual dataset information without relying on meta-data. We performed two metric learning methods, unsupervised contrastive learning and table-query metric learning, utilizing pre-trained tabular language models (TaLMs). The experimental results suggest that these metric learning methods can obtain representations that are more consistent with the labels representing the topics of the datasets. Moreover, our embeddings without metadata performed as well as those with metadata, suggesting that appropriate data extraction and clustering can be performed even in cases where metadata are sparse or incomplete. Kosuke Manabe, Yukihisa Fujita, Masahiro Kuwahara, Teruaki Hayashi |
IEEE Big Data | 1 |
| 2023 | Variable-based Learning Considering Topic Specificity in Heterogeneous Data Clustering TasksabstractRecently, data mining via interdisciplinary co-creation has attracted considerable social attention, and various data have been published, for free or for a fee. Data publication encourages the exchange and combination of data between different institutions, which is helpful for interdisciplinary data collaboration. However, issues pertaining to designing high-quality data for interdisciplinary data discovery remain in data search. Variables are frameworks for data, and reflect the data topics and intent of the data design. In this study, the relationships between data topics and variables for a large dataset were quantitatively investigated to provide suggestions for data design and exploration. The probability of occurrence of variables and their pairs for each topic was determined to elucidate the relationship between the topics and variables; subsequently, clustering was applied based on these relationships. Kosuke Manabe, Yukihisa Fujita, Masahiro Kuwahara, Teruaki Hayashi |
IEEE Big Data | 1 |