VLDB 2026 Research / reviewers in the wild / expert
Satoshi Masuda
dblp:137/6889
· DBLP profile ↗
9ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-5633-7249ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Datetime Feature Recommendation by Word Embedding Methods Using Data Column NamesabstractThe analysis of large volumes of data to derive new insights is commonly referred to as data science, and its widespread adoption is increasingly necessary. One critical step in data science workflows is feature extraction from data, known as feature engineering. This process heavily depends on expert experience, which has led to current research efforts aimed at its automation. In this paper, we introduce a novel approach to automate feature extraction from the textual information found in data column names. Specifically, we employ natural language processing and source code analysis techniques on existing source code and column names, with a particular focus on datetime features, to construct a knowledge database. Utilizing this knowledge database, we propose a system that recommends datetime features based on newly provided textual information. Departing from conventional methods such as one-hot encoding and Word2Vec word embeddings, our approach is informed by prior research and shifts toward Doc2Vec, which vectorizes at the document level. In our experiments, we validate the classification accuracy of the knowledge database and demonstrate its application in predictive tasks, such as housing price prediction, where it shows improved prediction accuracy. Our proposed approach, which utilizes a Doc2Vec model pre-trained on Wikipedia, results in enhanced prediction accuracy during vectorization. Satoshi Masuda, Tomohiro Takeda |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2023 | Datetime Feature Recommendation Using Textual InformationabstractAnalysis to gain new knowledge from huge amounts of data is called data science, and its widespread use is now socially important. Feature engineering, the process of extracting features from data, is one of the main tasks in data science, and since this task relies on the experience of experts, research is being conducted to automate it. In this paper, we propose a novel approach to automate feature identification from textual information in data column names. Specifically, we use techniques of natural language processing and source code analysis, for data descriptions and source codes in Python notebooks to create a knowledge database with a particular focus on datetime features. We develop a recommendation system of datetime features for newly given text information based on that knowledge database. In experiments, we confirmed the classification accuracy of the knowledge database, applied the database to actual forecasting tasks such as home price forecasting, and achieved 5.76% on average for accuracy gain. Satoshi Masuda, Takaaki Tateishi, Toshihiro Takahashi |
KES | 1 |
| 2021 | Online Adaptation of Parameters using GRU-based Neural Network with BO for Accurate Driving ModelabstractTesting self-driving cars in different areas requires surrounding cars with accordingly different driving styles such as aggressive or conservative styles. Calibrating a driving model (DM) makes the simulated driving behavior closer to human-driving behavior, and enable the simulation of human-driving cars. Conventional DM-calibrating methods do not take into account that the parameters in a DM vary while driving. These "fixed" calibrating methods cannot reflect an actual interactive driving scenario. In this paper, we propose a DM-calibration method for measuring human driving styles to reproduce real car-following behavior more accurately. The method includes 1) an objective entropy weight method for measuring and clustering human driving styles, and 2) online adaption of DM parameters based on deep learning by combining Bayesian optimization and a gated recurrent unit neural network. We conducted experiments to evaluate the proposed method, and the results indicate that it can be easily used to measure human driver styles. The experiments also showed that we can calibrate a corresponding DM in a virtual testing environment with up to 26% more accuracy than with fixed calibration methods. Zhanhong Yang, Satoshi Masuda, Michiaki Tatsubori |
SIGSPATIAL/GIS | 2 |
| 2021 | Data Quality for Machine Learning TasksabstractThe quality of training data has a huge impact on the efficiency, accuracy and complexity of machine learning tasks. Data remains susceptible to errors or irregularities that may be introduced during collection, aggregation or annotation stage. This necessitates profiling and assessment of data to understand its suitability for machine learning tasks and failure to do so can result in inaccurate analytics and unreliable decisions. While researchers and practitioners have focused on improving the quality of models, there are limited efforts towards improving the data quality. Nitin Gupta 0005, Shashank Mujumdar, Hima Patel, Satoshi Masuda, Naveen Panwar, Sambaran Bandyopadhyay, Sameep Mehta, Shanmukha C. Guttula, Shazia Afzal, Ruhi Sharma Mittal, Vitobha Munigala |
KDD | 4 |
| 2021 | 2nd International Workshop on Data Quality Assessment for Machine LearningabstractThe 2nd International Workshop on Data Quality Assessment for Machine Learning (DQAML'21) is organized in conjunction with the Special Interest Group on Knowledge Discovery and Data Mining (SIGKDD). This workshop aims to serve as a forum for the presentation of research related to data quality assessment and remediation in AI/ML pipeline. Data quality is a critical issue in the data preparation phase and involves numerous challenging problems related to detection, remediation, visualization and evaluation of data issues. The workshop aims to provide a platform to researchers and practitioners to discuss such challenges across different modalities of data like structured, time series, text and graphical. The aim is to attract perspectives from both industrial and academic circles. Hima Patel, Fuyuki Ishikawa, Laure Berti-Équille, Nitin Gupta 0005, Sameep Mehta, Satoshi Masuda, Shashank Mujumdar, Shazia Afzal, Srikanta J. Bedathur, Yasuharu Nishi |
KDD | 6 |
| 2020 | Guidelines for Quality Assurance of Machine Learning-based Artificial Intelligence
Koichi Hamada, Fuyuki Ishikawa, Satoshi Masuda, Tomoyuki Myojin, Yasuharu Nishi, Hideto Ogawa, Takahiro Toku, Susumu Tokumoto, Kazunori Tsuchiya, Yasuhiro Ujita, Mineo Matsuya |
SEKE | 3 |
| 2020 | Guidelines for Quality Assurance of Machine Learning-Based Artificial IntelligenceabstractSignificant effort is being put into developing industrial applications for artificial intelligence (AI), especially those using machine learning (ML) techniques. Despite the intensive support for building ML applications, there are still challenges when it comes to evaluating, assuring, and improving the quality or dependability. The difficulty stems from the unique nature of ML, namely, system behavior is derived from training data not from logical design by human engineers. This leads to black-box and intrinsically imperfect implementations that invalidate many principles and techniques in traditional software engineering. In light of this situation, the Japanese industry has jointly worked on a set of guidelines for the quality assurance of AI systems (in the Consortium of Quality Assurance for AI-based Products and Services) from the viewpoint of traditional quality-assurance engineers and test engineers. We report on the second version of these guidelines, which cover a list of quality evaluation aspects, catalogue of current state-of-the-art techniques, and domain-specific discussions in five representative domains. The guidelines provide significant insights for engineers in terms of methodologies and designs for tests driven by application-specific requirements. Gaku Fujii, Koichi Hamada, Fuyuki Ishikawa, Satoshi Masuda, Mineo Matsuya, Tomoyuki Myojin, Yasuharu Nishi, Hideto Ogawa, Takahiro Toku, Susumu Tokumoto, Kazunori Tsuchiya, Yasuhiro Ujita |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2018 | Obtaining Exhaustive Answer Set for Q&A-based Inquiry System using Customer Behavior and Service Function ModelingabstractWhen customers are interested in a service or intend to buy it, they sometimes have questions on that service. In this study, we considered an inquiry system in which customers ask questions on a specific service and obtain correct information on the service. For such an inquiry system, a question-answering (Q&A) technology is needed. Many programming modules for such a technology have been developed and can be easily used for system development. In many Q&A technologies, machine-learning techniques are involved, and we need to prepare training data consisting of pairs of an answer and assumed questions. For training-data preparation, an answer set for a service should be defined as the first step and the answer set should cover all the information on the service that customers may ask about. By using a customer-behavior model and introducing a service-function model, we propose a method of effectively collecting knowledge information for an answer set on a service. Through a case study, we show that we can collect exhaustive knowledge information for an answer set with our method compared to the case in which domain experts collect knowledge information in their own way. For an actual project, we also considered an actual inquiry-system-development project, with training data obtained with the proposed method, and showed that the system covers almost all the information on the service that customers may ask after a user test. Hironori Takeuchi, Satoshi Masuda, Kohtaroh Miyamoto, Shiki Akihara |
KES | 2 |
| 2013 | A Method of Creating Testing Pattern for Pair-wise Method by Using Knowledge of Parameter ValuesabstractIt is important for software testing to create high test case coverage. Test cases are almost created by manually in current situation, so test case coverage is depend on the individual skills. We discuss a method of creating testing pattern for Pair- wise method by using knowledge of parameter values. The method targets functional testing from screens for Web application systems. The method uses knowledge base for identifying pair-wise parameter values by using document analysis to specification documents, boundary analysis and defects analysis, so that it gets rid of dependencies of individual skills. We also discuss case studies which demonstrate the method of creating high test case coverage by using pair-wise parameter values. Satoshi Masuda, Tohru Matsuodani, Kazuhiko Tsuda |
KES | 1 |