EDBT 2026 Demo / reviewers in the wild / expert
Yukihiko Okada
dblp:123/5407
· DBLP profile ↗
10ranked-venue papers in the field
0as first author
10since 2021 · last 2025
0000-0003-4903-4191ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 10
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cash-Flow Prediction Model Using Graph Neural Networks and Spatiotemporal Information from Double-Entry Bookkeeping Data
Ryoki Motai, Masato Kamebuchi, Sora Watanabe, Ryo Matsumoto, Keiichi Azuma, Yukihiko Okada |
IEEE Big Data | 6 |
| 2025 | DC-CP: Data Collaboration Conformal Prediction from Model Training to Conformal Prediction in Single-Round Communication While Preserving Privacy
Tomoru Nakayama, Yuji Kawamata, Akira Imakura, Yukihiko Okada |
IEEE Big Data | 4 |
| 2025 | A Reliable Decision Support Framework for SME Default Prediction Using Uncertainty-Aware Bayesian GEV Regression
Kaito Terada, Keiji Abe, Toshihiko Takeda, Yuji Kawamata, Ryoki Motai, Yukihiko Okada |
IEEE Big Data | 6 |
| 2024 | An explainable framework based on counterfactual explanations for multi-class financial distress prediction of small and medium enterprisesabstractSmall and medium enterprises (SMEs) play a crucial role in supporting the global economy by contributing significantly to employment and value creation. Therefore, accurately predicting early signs of financial distress in SMEs and taking timely management improvement actions is of paramount importance. However, there is a lack of research on developing multi-class financial distress prediction (MFDP) models specifically for SMEs. Moreover, traditional MFDP models often lack interpretability. To address this, this study proposes definitions for multi-class financial distress in SMEs and constructs a MFDP model using machine learning. Additionally, it introduces an explainable framework to interpret the constructed model. Specifically, the study classifies SMEs' financial conditions into three categories: healthy, mild financial distress, and severe financial distress. It conducts a comparative analysis of six machine learning models. Furthermore, it proposes new interpretation methods for the MFDP model using SHapley Additive exPlanations (SHAP) and counterfactual explanations (CE), both of which are explainable artificial intelligence techniques. The empirical results reveal that the MFDP model, which adjusts data balance using random oversampling and integrates LightGBM with a one-versus-rest decomposition method, demonstrates the highest performance. Moreover, the proposed explainable framework demonstrates that it can provide practical and concrete improvement strategies to the model users, such as managers and financial institutions. Renon Ando, Yuji Kawamata, Toshihiko Takeda, Yukihiko Okada |
IEEE Big Data | 4 |
| 2024 | Proposing a Low-Rank Approximation Method with Mathematical Guarantees for High-Dimensional Tensor DataabstractThe analysis of high-dimensional tensor data presents significant challenges for traditional methods such as CANDECOMP/PARAFAC (CP) decomposition and Higher-Order Singular Value Decomposition (HOSVD) due to the size, complexity, and noise inherent in the data. Recent studies have emphasized the critical need for noise-robust approaches capable of reducing dimensionality while preserving data integrity, particularly for matrix-form data, underscoring the necessity of theoretical advancements in noise reduction. In this study, we propose a novel method, Noise-Reduced Higher-Order Singular Value Decomposition (NR-HOSVD), which integrates noise reduction techniques into the HOSVD framework to enhance dimensionality reduction for matrix data. In comparative simulations with CP decomposition and HOSVD, NR-HOSVD consistently demonstrated lower relative errors as the rank increased, particularly for high-dimensional tensors, surpassing the performance of conventional methods beyond a certain rank threshold. Furthermore, NR-HOSVD exhibited superior stability and faster convergence on large-scale datasets compared to existing approaches. These findings suggest that NR-HOSVD provides a more accurate and efficient approach to high-dimensional tensor data analysis, making it particularly well-suited for applications in big data and machine learning. In the future, NR-HOSVD is expected to serve as a pivotal method for enhancing the accuracy and efficiency of high-dimensional tensor data analysis, with broad applicability anticipated in fields such as large-scale data processing and precision medicine. Hiroki Hasegawa, Kazuyoshi Yata, Yukihiko Okada, Jun Kunimatsu |
IEEE Big Data | 3 |
| 2024 | Practical Experiment of Predicting Cash Flows with LSTM and Double-entry Bookkeeping DataabstractDouble-entry bookkeeping data (DBD) systematically record the daily transactions of an organization and are the foundation of accounting information. Typically, they are more detailed and larger in scale than financial statements and feature cross-references between accounts. These characteristics make them beneficial for management. However, studies on quantitative evaluations of their usefulness in organizational management are scarce. This study investigates whether features derived from DBD by leveraging their characteristics can enhance future cash-balance predictions, with the primary focus on enhancing the cash management of organizations. To achieve this, DBD obtained from seven medical clinics were represented as graphs using their cross-references, and the graphs were then mapped into vector features based on graph similarity. The experimental results revealed that a long short-term memory model incorporating features unique to DBD obtained higher predictive accuracy than a simple long short-term memory model that only relies on past cash-balance time-series data. Additionally, an auto regressive integrated moving average model was more suitable for some clinics than the long short-term memory model. This can be attributed to the specific business characteristics and structural changes over time within each clinic. The main contribution of this study is the proposal and demonstration of a method that utilizes DBD, which are stored within organizations but often underutilized, to enhance cash management, thereby illustrating their significance. Based on our findings, business intelligence that provides predictive information using DBD could be developed by management professionals and companies offering accounting services. Ryoki Motai, Sota Mashiko, Masato Kamebuchi, Sora Watanabe, Ryo Matsumoto, Yukihiko Okada |
IEEE Big Data | 6 |
| 2023 | Data-Driven Approaches to Detecting Misdeliveries in Truck Logistics using GPS DataabstractMisdelivery in logistic services leads to increased costs and degradation of packages. For deliveries using trucks, erroneous deliveries are prevented by checking identification numbers on packages as the packages pass through delivery points. In recent years, it has become possible to monitor real-time location by attaching GPS receivers to packages, but few concrete efforts have been made to detect anomalies using truck transport data. This study introduces a basic framework for the detection of misdeliveries and a summary of the problems in actual truck misdelivery using special medical supply delivery data provided by a major Japanese logistics company. It is shown that the system can detect erroneous deliveries with high accuracy, even for actual delivery data with coarse resolution owing to the cost of installing the equipment. This study has the potential not only to improve logistics and reduce costs, but also to solve various social problems such as driver shortages. Ayumu Hidaka, Ryota Shin, Atsushi Tsuchiya, Norihiko Nakabayashi, Yukihiko Okada, Yohei Shida |
IEEE Big Data | 5 |
| 2023 | Does Double-entry Bookkeeping Information Generated Using node2vec Contribute to Forecasting Future Performance?abstractIn recent years, an increasing number of studies have attempted to extract information from journals that cannot be conveyed in financial statements. Journals are valuable assets that enhance corporate value and companies’ competitiveness because they contain abundant information about past transactions. However, only a few studies have quantitatively clarified the value of double-entry bookkeeping and journals. Therefore, finding useful business knowledge in journals is currently difficult for managers and administrators. This study focuses on cross-reference information, which is unique to double-entry bookkeeping and not included in financial statements and trial balances. We then created new features to enhance the explanatory power of future business performance from journals. To achieve this objective, we constructed a graph based on journals, in which cross-reference is represented as edges. We create features unique to journal entry using two methods: network metrics and feature generation by node embedding. Finally, we tested for Granger causality using a vector autoregressive model constructed from created feature and performance. Our experiment showed that we successfully identified Granger causality within the unique features of double-entry journals for all five performance indicators of seven companies. Practitioners can use the identified features to build an alert system for sudden performance declines and facilitate understanding of their business. Future research will be necessary to verify the significance of double-entry bookkeeping and journal entries. Ryoki Motai, Masato Kamebuchi, Sora Watanabe, Ryo Matsumoto, Yukihiko Okada |
IEEE Big Data | 5 |
| 2023 | Method for creating privacy-preserving information using Probabilistic Latent Semantic AnalysisabstractIn recent years, rapid digitization and technological advances have enabled governments and companies to utilize various types of personal information, including environment, medical, and lifestyle. Privacy protection is an important issue in the utilization of personal information, and anonymization methods such as k-anonymization and 1-diversity have been proposed in the past. However, to utilize data effectively, it is also important not to lose the usefulness of the information. Privacy protection and preserving the usefulness of information are in a trade-off relationship, and balancing these two aspects has been a difficult problem. Recently, microaggregation using clustering algorithms has been attracting attention as an anonymization method that ensures the usefulness of the data while protecting privacy. In this study, we proposed a new microaggregation method based on iterative processing using probabilistic latent semantic analysis and evaluated the proposed method based on the use case. The use case is a classification of social isolation and an analysis of the relationship with physical frailty risk using the classified groups. Compared to microaggregation using quasi-identifiers and microaggregation based on a conventional clustering algorithm, the proposed method was able to find social isolation characteristics obtained in the raw data with higher accuracy. In addition, in a more detailed analysis focusing on a population with certain characteristics, the proposed method was able to replicate the analysis results obtained using the raw data and to make appropriate assessments of the health risks. Yuki Sugawara, Eiichi Sakurai, Yoichi Motomura, Yukihiko Okada, Kai Tanabe, Akiko Tsukao, Shinya Kuno |
IEEE Big Data | 4 |
| 2022 | Predictive model of frailty onset using Bayesian networkabstractIn this study, we constructed a model for predicting the future frailty status of individuals based on information regarding their current lifestyle habits. In recent years, as the global population ages, the number of people certified for long-term care has increased. In an effort to prevent the need for long-term care, prediction of the onset of frailty, which is positioned as a preliminary stage of the caregiving state, is gaining attention. However, the onset of frailty takes a long time, and the incidence of frailty differs according to one’s health literacy. Therefore, a prediction is extremely difficult to achieve. In this study, we constructed a prediction model using a probabilistic latent semantic analysis and a Bayesian network. First, the probabilistic latent semantic analysis was conducted to classify individuals according to their health literacy. Second, we constructed a future frailty prediction model using a Bayesian network for each health literacy segment. As a result, we clarified the following: It is more appropriate to predict a future frailty by predicting the future lifestyle using information on the current lifestyle. In addition, more accurate prediction models can be constructed when individuals are divided based on their health literacy. Yujiro Kawai, Eiichi Sakurai, Yuki Sugawara, Yukihiko Okada, Kai Tanabe, Akiko Tsukao, Shinya Kuno |
IEEE Big Data | 4 |