EDBT 2026 Demo / reviewers in the wild / expert
Ryuta Matsuno
dblp:218/0669
· DBLP profile ↗
10ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0002-4543-2128ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Source Component Shift Adaptation via Offline Decomposition and Online Mixing ApproachabstractThis paper addresses source component shift adaptation, aiming to update predictions adapting to source component shifts for incoming data streams based on past training data. Existing online learning methods often fail to utilize recurring shifts effectively, while model-pool-based methods struggle to capture individual source components, leading to poor adaptation. In this paper, we propose a source component shift adaptation method via an offline decomposition and online mixing approach. We theoretically identify that the problem can be divided into two subproblems: offline source component decomposition and online mixing weight adaptation. Based on this, our method first determines prediction models, each of which learns a source component solely based on past training data offline through the EM algorithm. Then, it updates the mixing weight of the prediction models for precise prediction through online convex optimization. Thanks to our theoretical derivation, our method fully leverages the characteristics of the shifts, achieving superior adaptation performance over existing methods. Experiments conducted on various real-world regression datasets demonstrate that our method outperforms baselines, reducing the cumulative test loss by up to 67.4%. Ryuta Matsuno |
ECAI | 1 |
| 2025 | Learning from Two-Sample-Averaged DataabstractThis paper presents a novel method for training nonlinear regression models using two-sample-averaged data, where each sample is the average of input-output pairs from two i.i.d. samples. Our method leverages maximum likelihood estimation under the additive Gaussian noise and normally distributed inputs assumptions. For theoretical analysis on the likelihood, we utilize Taylor approximation and eigenvalue decomposition. We then apply Fourier transformation for better numerical calculation of the approximated likelihood. Experiments on synthetic datasets demonstrate that our method outperforms the baselines, reducing the test MSE by up to 97.07 %, which is even comparable with that of models trained with original data. Furthermore, on real-world datasets using neural networks, we confirm that our method achieves an average rank of 1.6, confirming its overall superiority. This paper establishes the foundation of averaged-data learning, which is promising for privacy-preserving machine learning. Ryuta Matsuno, Akira Kitaoka |
ICDM | 1 |
| 2025 | CDST-Viz: Tree-Based Segmentation and Visual Analytics of Concept DriftabstractConcept drift, defined as a change in the conditional distribution of the target given the features, can seriously degrade the predictive performance of machine-learning models. Existing analysis methods often lack intuitive visual summaries, making it difficult to understand and address drift. We introduce CDST-Viz, a framework for the segmentation and visualization of concept drift. CDST-Viz first trains a Concept Drift Segmentation Tree (CDST) that partitions the feature space into regions with similar drift patterns across two time periods. The trained tree is rendered into an interactive treemap and accompanied by time-series plots of the target in each region to reveal where and how the distribution shifts. Case studies on six public regression datasets show that CDST-Viz clearly reveals fine-grained drift patterns and supports informed responses to concept drift. Keita Sakuma, Ryuta Matsuno, Masakazu Hirokawa |
IV | 2 |
| 2025 | GBCE: Enhanced Training Loss to Estimate Accuracy of Models in Production
Ryuta Matsuno |
PAKDD (3) | 1 |
| 2024 | Backward Compatibility in Attributive Explanation and Enhanced Model Training MethodabstractModel update is a crucial process in the operation of ML/AI systems. While updating a model generally enhances the average prediction performance, it also significantly impacts the explanations of predictions. In real-world applications, even minor changes in explanations can have detrimental consequences. To tackle this issue, this paper introduces BCX, a quantitative metric that evaluates the backward compatibility of attributive explanations between pre- and post-update models. BCX utilizes practical agreement metrics to calculate the average agreement between the explanations of pre- and post-update models, specifically among samples on which both models accurately predict. In addition, we propose BCXR, a BCX-aware model training method by designing surrogate losses which theoretically lower bounds agreement scores. Furthermore, we present a universal variant of BCXR that improves all agreement metrics, utilizing L2 distance among the explanations of the models. To validate our approach, we conducted experiments on eight real-world data sets, demonstrating that BCXR achieves superior trade-offs between predictive performances and BCX scores, showcasing the effectiveness of our BCXR methods. Ryuta Matsuno |
ECAI | 1 |
| 2024 | Interactive Visualization of Ensemble Decision Trees Based on the Relations Among Weak LearnersabstractEnsemble learning that combines multiple weak learners for enhanced performance, is widely used but suffers from low interpretability/explainability. This leads challenges not only in operational aspects like model maintenance and quality assurance but also in addressing societal needs such as fairness and privacy. To tackle this, we propose a new visualization method focusing on the relationship among weak learners in ensemble models to improve understanding of the model structure and its learning processes. In this paper, we defined the relation between weak learners based on a “common sample” in gradient-boosting decision trees, and a visualization method as a three-dimensional graph structure was proposed. Ensemble models trained with synthetic data sets that include typical distribution shifts and real-world open data sets were visualized. As a result, we demonstrated that this approach enables a more accessible understanding of the behavior and structure of ensemble models comprising multiple weak learners, facilitating the identification of overfitting and underfitting through visualization of changes during the training and validation processes. Miyu Kashiyama, Masakazu Hirokawa, Ryuta Matsuno, Keita Sakuma, Takayuki Itoh |
IV | 3 |
| 2024 | Model Accuracy-Oriented Data Sets Visualization for Understanding Temporal Changes in DataabstractUnderstanding temporal changes in data is crucial for successful MLOps, however, this understanding is generally challenging. This paper presents an algorithm that visualizes multiple data sets and a prediction model in a single plot. The proposed algorithm efficiently captures the changes in characteristics of the data sets based on accuracies of models trained on each data set. Our visualization enables data scientists to easily identify trends in the direction of change, periodicity, and anomalies in the data sets, as well as the predictive performance of the prediction model, all at a glance. Case studies using six real-world open data demonstrates that our visualization effectively provides valuable insights into the changes of the data sets, facilitating more informed decision-making in MLOps. Ryuta Matsuno, Keita Sakuma, Masakazu Hirokawa |
IV | 1 |
| 2023 | A Robust Backward Compatibility Metric for Model RetrainingabstractModel retraining and updating are essential processes in AI applications. However, during updates, there is a potential for performance degradation, in which the overall performance improves, but local performance deteriorates. This study proposes a backward compatibility metric that focuses on the compatibility of local predictive performance. The score of the proposed metric increases if the accuracy over the conditional distribution for each input is higher than before. Furthermore, we propose a model retraining method based on the proposed metric. Due to the use of the conditional distribution, our metric and retraining method are robust against label noises, while existing sample-based backward compatibility metrics are often affected by noise. We perform a theoretical analysis of our method and derive an upper bound for the generalization error. Numerical experiments demonstrate that our retraining method enhances compatibility while achieving equal or better trade-offs in overall performance compared to existing methods. Ryuta Matsuno, Keita Sakuma |
CIKM | 1 |
| 2023 | Quantitative Decomposition of Prediction Errors Revealing Multi-Cause Impacts: An Insightful Framework for MLOpsabstractAs machine learning applications expand in various industries, MLOps, which enables continuous model operation and improvement, becomes increasingly significant. Identifying causes of prediction errors, such as low model performance or anomalous samples, and implementing appropriate countermeasures are essential for effective MLOps. Furthermore, quantitatively evaluating each cause's impact is necessary to determine the effectiveness of countermeasures. In this study, we propose a method to quantitatively decompose a single sample's prediction error into contributions from multiple causes. Our method involves four steps: calculating the prediction error, computing metrics related to error causes, using a regression model to learn the relationship between the error and metrics, and applying SHAP to interpret the model's predictions and calculate the contribution of each cause to the prediction error. Numerical experiments with open data show that our method offers valuable insights for model improvement, confirming the effectiveness of our approach. Keita Sakuma, Ryuta Matsuno, Yoshio Kameda |
CIKM | 2 |
| 2020 | Improved mixing time for k-subgraph samplingabstractUnderstanding the local structure of a graph provides valuable insights about the underlying phenomena from which the graph has originated. Sampling and examining k-subgraphs is a widely used approach to understand the local structure of a graph. In this paper, we study the problem of sampling uniformly k-subgraphs from a given graph. We analyse a few different Markov chain Monte Carlo (MCMC) approaches, and obtain analytical results on their mixing times, which improve significantly the state of the art. In particular, we improve the bound on the mixing times of the standard MCMC approach, and the state-of-the-art MCMC sampling method PSRW, using the canonical-paths argument. In addition, we propose a novel sampling method, which we call recursive subgraph sampling RSS, and its optimized variant RSS'. The proposed methods, RSS and RSS', are provably faster than the existing approaches. We conduct experiments and verify the uniformity of samples and the efficiency of RSS and RSS'. Ryuta Matsuno, Aristides Gionis |
SDM | 1 |