Yuta Nakahara

dblp:194/7906 · DBLP profile ↗
← Back
17ranked-venue papers
12as first author
12since 2021 · last 2026
0000-0002-0553-7910ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 8 · 6 first-author · 5 since 2021Security and privacy · 6 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Soft Bayesian Context Tree Models for Real-Valued Time Series
abstract
This paper proposes the soft Bayesian context tree model (Soft-BCT), which is a novel BCT model for real-valued time series. The Soft-BCT considers soft (probabilistic) splits of the context space, instead of hard (deterministic) splits of the context space as in the previous BCT for real-valued time series. A learning algorithm of the Soft-BCT is proposed based on the variational inference. The results of experiments demonstrate the superiority of the Soft-BCT compared to the previous BCT for some datasets.
Shota Saito, Yuta Nakahara, Toshiyasu Matsushima
ISIT2
2025 Bayesian Decision Theory on Decision Trees: Uncertainty Evaluation and Interpretability
abstract
Deterministic decision trees have difficulty in evaluating uncertainty especially for small samples. To solve this problem, we interpret the decision trees as stochastic models and consider prediction problems in the framework of Bayesian decision theory. Our models have three kinds of parameters: a tree shape, leaf parameters, and inner parameters. To make Bayesian optimal decisions, we have to calculate the posterior distribution of these parameters. Previously, two types of methods have been proposed. One marginalizes out the leaf parameters and samples the tree shape and the inner parameters by Metropolis-Hastings (MH) algorithms. The other marginalizes out both the leaf parameters and the tree shape based on a concept called meta-trees and approximates the posterior distribution for the inner parameters by a bagging-like method. In this paper, we propose a novel MH algorithm where the leaf parameters and the tree shape are marginalized out by using the meta-trees and only the inner parameters are sampled. Moreover, we update all the inner parameters simultaneously in each MH step. This algorithm accelerates the convergence and mixing of the Markov chain. We evaluate our algorithm on various benchmark datasets with other state-of-the-art methods. Further, our model provides a novel statistical evaluation of feature importance.
Yuta Nakahara, Shota Saito, Naoki Ichijo, Koki Kazama, Toshiyasu Matsushima
AISTATS1
2024 Bayesian Decision-Theoretic Prediction with Ensemble of Meta-Trees for Classification Problems
abstract
Decision tree algorithms are one of the most popular methods in machine learning. However, most decision tree algorithms do not assume a stochastic model behind data. On the other hand, a meta-tree was recently proposed as a stochastic model with a tree structure. The prediction under the assumption of the meta-tree is decided using Bayesian decision theory. Although the optimal prediction can be calculated with an assumption of a known meta-tree, an approximation is necessary to obtain a prediction under the problem setting of an unknown meta-tree because of the marginalization of all possible meta-trees. In this paper, we propose an approximation method, where a subset of meta-trees is sequentially constructed, and the prediction is made by weighting the meta-trees. For the approximation, we clarify the approaches of constructing the subset and making predictions with weighted meta-trees. We also examine the effectiveness of the approaches in an experiment using synthetic data. In addition, we conduct an experiment on benchmark data to confirm the performance of the proposal.
Naoki Ichijo, Ryota Maniwa, Yuta Nakahara, Koshi Shimada, Toshiyasu Matsushima
ISITA3
2024 Variational Bayesian Methods for a Tree-Structured Stick-Breaking Process Mixture of Gaussians by Application of the Bayes Codes for Context Tree Models
abstract
The tree-structured stick-breaking process (TS-SBP) mixture model is a non-parametric Bayesian model that can represent tree-like hierarchical structures among the mixture components. For TS-SBP mixture models, only a Markov chain Monte Carlo (MCMC) method has been proposed and any variational Bayesian (VB) methods has not been proposed. In general, MCMC methods are computationally more expensive than VB methods. Therefore, we require a large computational cost to learn the TS-SBP mixture model. In this paper, we propose a learning algorithm with less computational cost for the TS-SBP mixture of Gaussians by using the VB method under an assumption of finite tree width and depth. When constructing such VB method, the main challenge is efficient calculation of a sum over all possible trees. To solve this challenge, we utilizes a subroutine in the Bayes coding algorithm for context tree models. We confirm the computational efficiency of our VB method through an experiments on a benchmark dataset.
Yuta Nakahara
ISITA1
2024 An Algorithmic Framework for Constructing Multiple Decision Trees by Evaluating Their Combination Performance Throughout the Construction Process
abstract
Predictions using a combination of decision trees are known to be effective in machine learning. Typical ideas for constructing a combination of decision trees for prediction are bagging and boosting. Bagging independently constructs decision trees without evaluating their combination performance and averages them afterward. Boosting constructs decision trees sequentially, only evaluating a combination performance of a new decision tree and the fixed past decision trees at each step. Therefore, neither method directly constructs nor evaluates a combination of decision trees for the final prediction. When the final prediction is based on a combination of decision trees, it is natural to evaluate the appropriateness of the combination when constructing them. In this paper, we propose a new algorithmic framework that constructs decision trees simultaneously and evaluates their combination performance throughout the construction process. Our framework repeats two procedures. In the first procedure, we construct new candidates of combinations of decision trees to find a proper combination of decision trees. In the second procedure, we evaluate each combination performance of decision trees under some criteria and select a better combination. To confirm the performance of the proposed framework, we experiment with synthetic and benchmark data.
Keito Tajima, Naoki Ichijo, Yuta Nakahara, Koshi Shimada, Toshiyasu Matsushima
SMC3
2023 Hyperparameter Learning of Bayesian Context Tree Models
abstract
In recent years, Bayesian counterparts of the context tree weighting method are studied for many tasks. All these tasks require a hyperparameter setting of the prior distribution for context tree models. Therefore, we provide a framework for statistically learning these hyperparameters from data. Specifically, we consider a hierarchical Bayesian model that assumes hyperprior distributions behind the hyperparameters and learn them using an empirical variational Bayesian (EVB) method. This is the first study to propose an EVB method on the Bayesian context trees. The derived algorithm has a suggestive form that consists of subroutines partially optimal to each local probabilistic model.
Yuta Nakahara, Shota Saito, Koshi Shimada, Toshiyasu Matsushima
ISIT1
2023 Tree-Structured Gaussian Mixture Models and Their Variational Inference
abstract
In this paper, Gaussian mixture models with tree structure and their variational inference methods are proposed for non-parametric Bayesian clustering. In this model, the tree structure is included as an unobservable random variable. The number of leaf nodes corresponds to the number of mixture components. This model is expected to capture not only the number of clusters but also tree structure from data.
Yuta Nakahara
SMC1
2022 Stochastic Model of Block Segmentation Based on Improper Quadtree and Optimal Code under the Bayes Criterion
abstract
Most previous studies on lossless image compression have focused on improving preprocessing functions to reduce the redundancy of pixel values in real images. However, we assumed stochastic generative models directly on pixel values and focused on achieving the theoretical limit of the assumed models. In this study, we proposed a stochastic model based on improper quadtrees. We theoretically derive the optimal code for the proposed model under the Bayes criterion. In general, Bayes-optimal codes require an exponential order of calculation with respect to the data lengths. However, we propose an efficient algorithm that takes a polynomial order of calculation without losing optimality by assuming a novel prior distribution.
Yuta Nakahara, Toshiyasu Matsushima
DCC1
2022 Probability Distribution on Rooted Trees
abstract
The hierarchical and recursive expressive capability of rooted trees is applicable to represent statistical models in various areas, such as data compression, image processing, and machine learning. On the other hand, such hierarchical expressive capability causes a problem in tree selection to avoid overfitting. One unified approach to solve this is a Bayesian approach, on which the rooted tree is regarded as a random variable and a direct loss function can be assumed on the selected model or the predicted value for a new data point. However, all the previous studies on this approach are based on the probability distribution on full trees, to the best of our knowledge. In this paper, we propose a generalized probability distribution for any rooted trees in which only the maximum number of child nodes and the maximum depth are fixed. Furthermore, we derive recursive methods to evaluate the characteristics of the probability distribution without any approximations.
Yuta Nakahara, Shota Saito, Akira Kamatsuka, Toshiyasu Matsushima
ISIT1
2022 Bayes Optimal Estimation and Its Approximation Algorithm for Difference with and without Treatment under URLC Model
Taisuke Ishiwatari, Shota Saito, Yuta Nakahara, Yuji Iikubo, Toshiyasu Matsushima
ISITA3
2022 Two-dimensional Autoregressive Model with Time-varying Parameters and the Bayes Codes
Yuta Nakahara, Toshiyasu Matsushima
ISITA1
2021 Hyperparameter Learning of Stochastic Image Generative Models with Bayesian Hierarchical Modeling and Its Effect on Lossless Image Coding
abstract
Explicit assumption of stochastic data generative models is a remarkable feature of lossless compression of general data in information theory. However, current lossless image coding mostly focus on coding procedures without explicit assumption of the stochastic generative model. Therefore, we have difficulty discussing the theoretical optimality of the coding procedure to the stochastic generative model. In this paper, we solve this difficulty by constructing a stochastic generative model by interpreting the previous coding procedure from another perspective. An important problem of our approach is how to learn the hyperparameters of the stochastic generative model because the optimality of our coding algorithm is guaranteed only asymptotically and the hyperparameter setting still affects the expected code length for finite length data. For this problem, we use Bayesian hierarchical modeling and confirm its effect by numerical experiments. In lossless image coding, this is the first study assuming such an explicit stochastic generative model and learning its hyperparameters, to the best of our knowledge.
Yuta Nakahara, Toshiyasu Matsushima
ITW1
2020 A Stochastic Model of Block Segmentation Based on the Quadtree and the Bayes Code for It
abstract
In this paper, we propose a novel stochastic model based on the quadtree, so that our model effectively represents the variable block size segmentation of images. Then, we construct the Bayes code for the proposed stochastic model. In general, the computational cost to calculate the posterior distribution required in the Bayes code increases exponentially with respect to the data size. However, we introduce an efficient algorithm to calculate it in the polynomial order of the data size without loss of the optimality. Some experiments are performed to confirm the flexibility of the proposed stochastic model and the efficiency of the introduced algorithm.
Yuta Nakahara, Toshiyasu Matsushima
DCC1
2020 Theoretical Analysis of the Advantage of Deepening Neural Networks
abstract
We propose two new criteria to understand the advantage of deepening neural networks. It is important to know the expressivity of functions computable by deep neural networks in order to understand the advantage of deepening neural networks. Unless deep neural networks have enough expressivity, they cannot have good performance even though learning is successful. In this situation, the proposed criteria contribute to understanding the advantage of deepening neural networks since they can evaluate the expressivity independently from the efficiency of learning. The first criterion shows the approximation accuracy of deep neural networks to the target function. This criterion has the background that the goal of deep learning is approximating the target function by deep neural networks. The second criterion shows the property of linear regions of functions computable by deep neural networks. This criterion has the background that deep neural networks whose activation functions are piecewise linear are also piecewise linear. Furthermore, by the two criteria, we show that to increase layers is more effective than to increase units at each layer on improving the expressivity of deep neural networks.
Yasushi Esaki, Yuta Nakahara, Toshiyasu Matsushima
ICMLA2
2020 Autoregressive Image Generative Models with Normal and t-distributed Noise and the Bayes Codes for Them
Yuta Nakahara, Toshiyasu Matsushima
ISITA1
2019 Covariance Evolution for Spatially "Mt. Fuji" Coupled LDPC Codes
abstract
A spatially “Mt. Fuji” coupled low-density parity check (LDPC) ensemble is a modified version of the original spatially coupled (SC) LDPC ensemble. Its desirable properties are first observed in experimentally. The decoding error probability in the error floor region over the binary erasure channel (BEC) is theoretically analyzed later. In this paper, as the last piece of the theoretical analysis over the BEC, we analyze the decoding error probability in the waterfall region by modifying the covariance evolution which has been used to analyze the original SC-LDPC ensemble.
Yuta Nakahara, Toshiyasu Matsushima
ITW1
2016 Spatially "Mt. Fuji" coupled LDPC codes
Yuta Nakahara, Shota Saito, Toshiyasu Matsushima
ISITA1