VLDB 2026 Research / reviewers in the wild / expert
Masafumi Oyamada
dblp:28/11004
· DBLP profile ↗
23ranked-venue papers
4as first author
16since 2021 · last 2025
0000-0002-4045-7350ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 15 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On Synthesizing Data for Context Attribution in Question AnsweringabstractGorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon, Sebastien Nicolas, Chia-Chien Hung, Timo Sztyler, Verena Heußer, Wiem Ben Rim, Masafumi Enomoto, Kunihiro Takeoka, Masafumi Oyamada, Goran Glavaš, Carolin Lawrence. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Gorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon, Sebastien Nicolas, Chia-Chien Hung, Timo Sztyler, Verena Heußer, Wiem Ben Rim, Masafumi Enomoto, Kunihiro Takeoka, Masafumi Oyamada, Goran Glavas, Carolin Lawrence |
ACL (1) | 12 |
| 2025 | LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM AgentsabstractLarge Language Models (LLMs) have demonstrated exceptional performance across a wide range of tasks.To further tailor LLMs to specific domains or applications, post-training techniques such as Supervised Fine-Tuning (SFT), Preference Learning, and model merging are commonly employed.While each of these methods has been extensively studied in isolation, the automated construction of complete post-training pipelines remains an underexplored area.Existing approaches typically rely on manual design or focus narrowly on optimizing individual components, such as data ordering or merging strategies.In this work, we introduce LaMDAgent (short for Language Model Developing Agent), a novel framework that autonomously constructs and optimizes full post-training pipelines through the use of LLM-based agents.LaMDAgent systematically explores diverse model generation techniques, datasets, and hyperparameter configurations, leveraging task-based feedback to discover high-performing pipelines with minimal human intervention.Our experiments show that LaMDAgent improves tool-use accuracy by 9.0 points while preserving instructionfollowing capabilities.Moreover, it uncovers effective post-training strategies that are often overlooked by conventional human-driven exploration.We further analyze the impact of data and model size scaling to reduce computational costs on the exploration, finding that model size scalings introduces new challenges, whereas scaling data size enables cost-effective pipeline discovery. Taro Yano, Yoichi Ishibashi, Masafumi Oyamada |
EMNLP | 3 |
| 2025 | Can Large Language Models Invent Algorithms to Improve Themselves?abstractYoichi Ishibashi, Taro Yano, Masafumi Oyamada. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yoichi Ishibashi, Taro Yano, Masafumi Oyamada |
NAACL (Long Papers) | 3 |
| 2025 | DISC: Dynamic Decomposition Improves LLM Inference ScalingabstractInference scaling methods for LLMs often rely on decomposing problems into steps (or groups of tokens), followed by sampling and selecting the best next steps. However, these steps and their sizes are often predetermined or manually designed based on domain knowledge. We propose dynamic decomposition, a method that adaptively and automatically partitions solution and reasoning traces into manageable steps during inference. By more effectively allocating compute -- particularly through subdividing challenging steps and prioritizing their sampling -- dynamic decomposition significantly improves inference efficiency. Experiments on benchmarks such as APPS, MATH, and LiveCodeBench demonstrate that dynamic decomposition outperforms static approaches, including token-level, sentence-level, and single-step decompositions, reducing the pass@10 error rate by 5.0%, 6.7%, and 10.5% respectively. These findings highlight the potential of dynamic decomposition to improve a wide range of inference scaling techniques. Jonathan Light, Wei Cheng 0002, Benjamin Rivière, Masafumi Oyamada, Mengdi Wang 0001, Yisong Yue, Santiago Paternain |
NeurIPS | 5 |
| 2025 | LLM-based Query Expansion Fails for Unfamiliar and Ambiguous QueriesabstractQuery expansion (QE) enhances retrieval by incorporating relevant terms, with large language models (LLMs) offering an effective alternative to traditional rule-based and statistical methods. However, LLM-based QE suffers from a fundamental limitation: it often fails to generate relevant knowledge, degrading search performance. Prior studies have focused on hallucination, yet its underlying cause-LLM knowledge deficiencies-remains underexplored. This paper systematically examines two failure cases in LLM-based QE: (1) when the LLM lacks query knowledge, leading to incorrect expansions, and (2) when the query is ambiguous, causing biased refinements that narrow search coverage. We conduct controlled experiments across multiple datasets, evaluating the effects of knowledge and query ambiguity on retrieval performance using sparse and dense retrieval models. Our results reveal that LLM-based QE can significantly degrade the retrieval effectiveness when knowledge in the LLM is insufficient or query ambiguity is high. We introduce a framework for evaluating QE under these conditions, providing insights into the limitations of LLM-based retrieval augmentation. Kenya Abe, Kunihiro Takeoka, Makoto P. Kato, Masafumi Oyamada |
SIGIR | 4 |
| 2024 | On the Use of Large Language Models for Table TasksabstractThe proliferation of large language models (LLMs) has catalyzed a diverse array of applications. This tutorial delves into the application of LLMs for tabular data and targets a variety of table-related tasks, such as table understanding, text-to-SQL conversion, and tabular data preprocessing. It surveys LLM solutions to these tasks in five classes, categorized by their underpinning techniques: prompting, fine-tuning, RAG, agents, and multimodal methods. It discusses how LLMs offer innovative ways to interpret, augment, query, and cleanse tabular data, featuring academic contributions and their practical use in the industrial sector. It emphasizes the versatility and effectiveness of LLMs in handling complex table tasks, showcasing their ability to improve data quality, enhance analytical capabilities, and facilitate more intuitive data interactions. By surveying different approaches, this tutorial highlights the strengths of LLMs in enriching table tasks with more accuracy and usability, setting a foundation for future research and application in data science and AI-driven analytics. Presentation slides for this tutorial will be available at: https://dongyuyang.github.io/tableLLM-tutorial/ . Yuyang Dong, Masafumi Oyamada, Chuan Xiao 0001 |
CIKM | 2 |
| 2024 | Jellyfish: Instruction-Tuning Local Large Language Models for Data PreprocessingabstractThis paper explores the utilization of LLMs for data preprocessing (DP), a crucial step in the data mining pipeline that transforms raw data into a clean format conducive to easy processing.Whereas the use of LLMs has sparked interest in devising universal solutions to DP, recent initiatives in this domain typically rely on GPT APIs, raising inevitable data breach concerns.Unlike these approaches, we consider instruction-tuning local LLMs (7 -13B models) as universal DP task solvers that operate on a local, single, and low-priced GPU, ensuring data security and enabling further customization.We select a collection of datasets across four representative DP tasks and construct instruction tuning data using data configuration, knowledge injection, and reasoning data distillation techniques tailored to DP.By tuning Mistral-7B, Llama 3-8B, and OpenOrca-Platypus2-13B, our models, namely, Jellyfish-7B/8B/13B, deliver competitiveness compared to GPT-3.5/4 models and strong generalizability to unseen tasks while barely compromising the base models' abilities in NLP tasks.Meanwhile, Jellyfish offers enhanced reasoning capabilities compared to GPT-3.5. Yuyang Dong, Chuan Xiao 0001, Masafumi Oyamada |
EMNLP | 4 |
| 2023 | Towards Large Language Model Organization: A Case Study on Abstractive SummarizationabstractIn this work we propose ”LLM organization”, an organizational structure-based LLM workflow for improving the performance of standard abstractive summarization techniques and mitigate unfaithful summary generation. We formulated the organizational structure-based LLM workflow as a directed acyclic graph (DAG), where each node corresponds to an LLM and each edge to a communication protocol. Our workflow is benchmarked on 5 datasets from various domains, using 7 evaluation metrics. The results indicate that LLM organization could mitigate unfaithfulness and increase the overall performance of abstractive summarization methods. Krisztián Boros, Masafumi Oyamada |
IEEE Big Data | 2 |
| 2023 | QA-Matcher: Unsupervised Entity Matching Using a Question Answering Model
Shogo Hayashi, Yuyang Dong, Masafumi Oyamada |
PAKDD (4) | 3 |
| 2023 | DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsabstractDue to the usefulness in data enrichment for data analysis tasks, joinable table discovery has become an important operation in data lake management. Existing approaches target equi-joins, the most common way of combining tables for creating a unified view, or semantic joins, which tolerate misspellings and different formats to deliver more join results. They are either exact solutions whose running time is linear in the sizes of query column and target table repository, or approximate solutions lacking precision. In this paper, we propose DeepJoin, a deep learning model for accurate and efficient joinable table discovery. Our solution is an embedding-based retrieval, which employs a pre-trained language model (PLM) and is designed as one framework serving both equi- and semantic (with a similarity condition on word embeddings) joins for textual attributes with fairly small cardinalities. We propose a set of contextualization options to transform column contents to a text sequence. The PLM reads the sequence and is fine-tuned to embed columns to vectors such that columns are expected to be joinable if they are close to each other in the vector space. Since the output of the PLM is fixed in length, the subsequent search procedure becomes independent of the column size. With a state-of-the-art approximate nearest neighbor search algorithm, the search time is sublinear in the repository size. To train the model, we devise the techniques for preparing training data as well as data augmentation. The experiments on real datasets demonstrate that by training on a small subset of a corpus, DeepJoin generalizes to large datasets and its precision consistently outperforms other approximate solutions'. DeepJoin is even more accurate than an exact solution to semantic joins when evaluated with labels from experts. Moreover, when equipped with a GPU, DeepJoin is up to two orders of magnitude faster than existing solutions. Yuyang Dong, Chuan Xiao 0001, Takuma Nozawa, Masafumi Enomoto, Masafumi Oyamada |
Proc. VLDB Endow. | 5 |
| 2022 | Table Enrichment System for Machine LearningabstractData scientists are constantly facing the problem of how to improve prediction accuracy with insufficient tabular data. We propose a table enrichment system that enriches a query table by adding external attributes (columns) from data lakes and improves the accuracy of machine learning predictive models. Our system has four stages, join row search, task-related table selection, row and column alignment, and feature selection and evaluation, to efficiently create an enriched table for a given query table and a specified machine learning task. We demonstrate our system with a web UI to show the use cases of table enrichment. Yuyang Dong, Masafumi Oyamada |
SIGIR | 2 |
| 2021 | Low-resource Taxonomy Enrichment with Pretrained Language ModelsabstractTaxonomies are symbolic representations of hierarchical relationships between terms or entities.While taxonomies are useful in broad applications, manually updating or maintaining them is labor-intensive and difficult to scale in practice.Conventional supervised methods for this enrichment task fail to find optimal parents of new terms in low-resource settings where only small taxonomies are available because of overfitting to hierarchical relationships in the taxonomies.To tackle the problem of low-resource taxonomy enrichment, we propose Musubu, an efficient framework for taxonomy enrichment in low-resource settings with pretrained language models (LMs) as knowledge bases to compensate for the shortage of information.Musubu leverages an LM-based classifier to determine whether or not inputted term pairs have hierarchical relationships.Musubu also utilizes Hearst patterns to generate queries to leverage implicit knowledge from the LM efficiently for more accurate prediction.We empirically demonstrate the effectiveness of our method in extensive experiments on taxonomies from both a SemEval task and real-world retailer datasets. Kunihiro Takeoka, Kosuke Akimoto, Masafumi Oyamada |
EMNLP (1) | 3 |
| 2021 | Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based ApproachabstractFinding joinable tables in data lakes is key procedure in many applications such as data integration, data augmentation, data analysis, and data market. Traditional approaches that find equi-joinable tables are unable to deal with misspellings and different formats, nor do they capture any semantic joins. In this paper, we propose PEXESO, a framework for joinable table discovery in data lakes. We target the case when textual values are embedded as high-dimensional vectors and columns are joined upon similarity predicates on high-dimensional vectors, hence to address the limitations of equi-join approaches and identify more meaningful results. To efficiently find joinable tables with similarity, we propose a block-and-verify method that utilizes pivot-based filtering. A partitioning technique is developed to cope with the case when the data lake is large and cannot fit in main memory. An experimental evaluation on real datasets shows that our solution identifies substantially more tables than equi-joins and outperforms other similarity-based options, and the join results are useful in data enrichment for machine learning tasks. The experiments also demonstrate the efficiency of the proposed method. Yuyang Dong, Kunihiro Takeoka, Chuan Xiao 0001, Masafumi Oyamada |
ICDE | 4 |
| 2021 | User Identity Linkage for Different Behavioral Patterns across Domains
Genki Kusano, Masafumi Oyamada |
ICWSM | 2 |
| 2021 | Quality Control for Hierarchical Classification with Incomplete Annotations
Masafumi Enomoto, Kunihiro Takeoka, Yuyang Dong, Masafumi Oyamada, Takeshi Okadome |
PAKDD (3) | 4 |
| 2021 | Continuous top-k spatial-keyword search on dynamic objects
Yuyang Dong, Chuan Xiao 0001, Hanxiong Chen, Jeffrey Xu Yu, Kunihiro Takeoka, Masafumi Oyamada, Hiroyuki Kitagawa |
VLDB J. | 6 |
| 2020 | Learning with Unsure ResponsesabstractMany annotation systems provide to add an unsure option in the labels, because the annotators have different expertise, and they may not have enough confidence to choose a label for some assigned instances. However, all the existing approaches only learn the labels with a clear class name and ignore the unsure responses. Due to the unsure response also account for a proportion of the dataset (e.g., about 10-30% in real datasets), existing approaches lead to high costs such as paying more money or taking more time to collect enough size of labeled data. Therefore, it is a significant issue to make use of these unsure.In this paper, we make the unsure responses contribute to training classifiers. We found a property that the instances corresponding to the unsure responses always appear close to the decision boundary of classification. We design a loss function called unsure loss based on this property. We extend the conventional methods for classification and learning from crowds with this unsure loss. Experimental results on realworld and synthetic data demonstrate the performance of our method and its superiority over baseline methods. Kunihiro Takeoka, Yuyang Dong, Masafumi Oyamada |
AAAI | 3 |
| 2019 | Meimei: An Efficient Probabilistic Approach for Semantically Annotating TablesabstractGiven a large amount of table data, how can we find the tables that contain the contents we want? A naive search fails when the column names are ambiguous, such as if columns containing stock price information are named “Close” in one table and named “P” in another table.One way of dealing with this problem that has been gaining attention is the semantic annotation of table data columns by using canonical knowledge. While previous studies successfully dealt with this problem for specific types of table data such as web tables, it still remains for various other types of table data: (1) most approaches do not handle table data with numerical values, and (2) their predictive performance is not satisfactory.This paper presents a novel approach for table data annotation that combines a latent probabilistic model with multilabel classifiers. It features three advantages over previous approaches due to using highly predictive multi-label classifiers in the probabilistic computation of semantic annotation. (1) It is more versatile due to using multi-label classifiers in the probabilistic model, which enables various types of data such as numerical values to be supported. (2) It is more accurate due to the multi-label classifiers and probabilistic model working together to improve predictive performance. (3) It is more efficient due to potential functions based on multi-label classifiers reducing the computational cost for annotation.Extensive experiments demonstrated the superiority of the proposed approach over state-of-the-art approaches for semantic annotation of real data (183 human-annotated tables obtained from the UCI Machine Learning Repository). Kunihiro Takeoka, Masafumi Oyamada, Shinji Nakadai, Takeshi Okadome |
AAAI | 2 |
| 2019 | Extracting Feature Engineering Knowledge from Data Science NotebooksabstractDesigning good features for machine learning models, which is called feature-engineering, is one of the most important tasks in data analysis. Well-designed features, which capture the characteristics of data, improve the predictive performance and explainability of the model. Since good features generally reflect the deep knowledge on business domains of the data and the analysis task, feature engineering is considered as one of the most difficult phases in data analysis. Nowadays, AutoML is trying to automate the data science process by producing good features by autonomous algorithms such as feature-synthesis and feature-selection. While AutoML is making success in some extent, it cannot reproduce all the features crafted by expert data scientists in reality because of its huge search space. In this paper, we take different approach for assisting feature engineering process: transfer expert data scientists knowledge as much as possible. Proposed approach extracts frequently used feature engineering operations from source codes or notebooks by pattern discovery. Since naive textual pattern discovery performs poor for the source code, our approach converts source codes into abstract syntax trees and discovers important feature engineering operations as subgraphs by performing frequent subgraph mining. Masafumi Oyamada |
IEEE BigData | 1 |
| 2018 | Accelerating Feature Engineering with Adaptive Partial Aggregation TreeabstractRange aggregation query is a fundamental operation in the feature engineering phase of the machine learning tasks, which computes statistics, such as the maximum and the standard deviation of a subset of records. Since the feature-engineering process is a trial-and-error process, data analysts repeatedly conduct tons of the range aggregation queries by changing the range conditions, which results in a heavy workload. To accelerate such repetitive range aggregation queries, we propose Adaptive Partial Aggregation Tree (APA-tree), which drastically reduces the amount of I/Os that happen in executing the range aggregation queries. The APA-tree partitions the data into several groups, executes range aggregations on each subgroups to obtain partial results, and caches the results in an imbalanced binary-tree. The APA-tree executes subsequent queries by reusing the cached partial query results as much as possible on the basis of the divide-and-conquer characteristic of the range aggregations. Experimental results confirm that APA-tree outperforms conventional partial aggregation methods regarding the amount of I/Os, especially in a skewed workload. Masafumi Oyamada |
IEEE BigData | 1 |
| 2017 | Relational Mixture of Experts: Explainable Demographics Prediction with Behavioral DataabstractGiven a collection of basic customer demographics (e.g., age and gender) andtheir behavioral data (e.g., item purchase histories), how can we predictsensitive demographics (e.g., income and occupation) that not every customermakes available?This demographics prediction problem is modeled as a classification task inwhich a customer's sensitive demographic y is predicted from his featurevector x. So far, two lines of work have tried to produce a"good" feature vector x from the customer's behavioraldata: (1) application-specific feature engineering using behavioral data and (2) representation learning (such as singular value decomposition or neuralembedding) on behavioral data. Although these approaches successfullyimprove the predictive performance, (1) designing a good feature requiresdomain experts to make a great effort and (2) features obtained fromrepresentation learning are hard to interpret. To overcome these problems, we present a Relational Infinite SupportVector Machine (R-iSVM), a mixture-of-experts model that can leveragebehavioral data. Instead of augmenting the feature vectors of customers, R-iSVM uses behavioral data to find out behaviorally similar customerclusters and constructs a local prediction model at each customer cluster. In doing so, R-iSVM successfully improves the predictive performance withoutrequiring application-specific feature designing and hard-to-interpretrepresentations. Experimental results on three real-world datasets demonstrate the predictiveperformance and interpretability of R-iSVM. Furthermore, R-iSVM can co-existwith previous demographics prediction methods to further improve theirpredictive performance. Masafumi Oyamada, Shinji Nakadai |
ICDM | 1 |
| 2017 | Link Prediction for Isolated Nodes in Heterogeneous Network by Topic-Based Co-clustering
Katsufumi Tomobe, Masafumi Oyamada, Shinji Nakadai |
PAKDD (1) | 2 |
| 2014 | MOARLE: Matrix Operation Accelerator Based on Run-Length Encoding
Masafumi Oyamada, Jianquan Liu, Kazuyo Narita, Takuya Araki |
APWeb | 1 |