VLDB 2026 Research / reviewers in the wild / expert
Tomoyuki Obuchi
dblp:77/8772
· DBLP profile ↗
9ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-1216-489XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Learning theory · 53% Trustworthy machine learning · 12% Deep learning architectures and training · 12% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 74% Algorithmic game theory and mechanism design · 13% Distributed computing theory · 4% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 50% Computational science and engineering · 50% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory
statistical learning theory |
1.6 | 3 | 2025 | Neural Collapse in Cumulative Link Models for Ordinal Regression: An Analysis with Unconstrained Feature Model · NeurIPS 2025 Semi-Analytic Resampling in Lasso · J. Mach. Learn. Res. 2019 Accelerating Cross-Validation in Multinomial Logistic Regression with $\ell_1$-Regularization · J. Mach. Learn. Res. 2018 |
Machine learning › Learning theory
generalization bounds |
0.9 | 1 | 2025 | Neural Collapse in Cumulative Link Models for Ordinal Regression: An Analysis with Unconstrained Feature Model · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Neural Collapse in Cumulative Link Models for Ordinal Regression: An Analysis with Unconstrained Feature Model · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
neural collapse |
0.9 | 1 | 2025 | Neural Collapse in Cumulative Link Models for Ordinal Regression: An Analysis with Unconstrained Feature Model · NeurIPS 2025 |
Machine learning › Learning theory
model selection |
0.7 | 2 | 2019 | Semi-Analytic Resampling in Lasso · J. Mach. Learn. Res. 2019 Accelerating Cross-Validation in Multinomial Logistic Regression with $\ell_1$-Regularization · J. Mach. Learn. Res. 2018 |
Mathematical optimization
continuous optimization |
0.7 | 2 | 2019 | Semi-Analytic Resampling in Lasso · J. Mach. Learn. Res. 2019 Accelerating Cross-Validation in Multinomial Logistic Regression with $\ell_1$-Regularization · J. Mach. Learn. Res. 2018 |
Mathematical optimization › statistical estimation › regression
regularized regression |
0.7 | 2 | 2019 | Semi-Analytic Resampling in Lasso · J. Mach. Learn. Res. 2019 Accelerating Cross-Validation in Multinomial Logistic Regression with $\ell_1$-Regularization · J. Mach. Learn. Res. 2018 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning › graphical model structure learning
ising model structure learning |
0.5 | 1 | 2021 | Ising Model Selection Using $\ell_{1}$-Regularized Linear Regression: A Statistical Mechanics Analysis · NeurIPS 2021 |
Machine learning › Reinforcement learning › sample efficiency
sample complexity analysis |
0.5 | 1 | 2021 | Ising Model Selection Using $\ell_{1}$-Regularized Linear Regression: A Statistical Mechanics Analysis · NeurIPS 2021 |
Machine learning › Learning theory › model selection
stability selection |
0.4 | 1 | 2019 | Semi-Analytic Resampling in Lasso · J. Mach. Learn. Res. 2019 |
Mathematical optimization › statistical estimation › regression › sparse regression
lasso |
0.4 | 1 | 2019 | Semi-Analytic Resampling in Lasso · J. Mach. Learn. Res. 2019 |
Machine learning › Learning theory › model selection
cross-validation |
0.3 | 1 | 2018 | Accelerating Cross-Validation in Multinomial Logistic Regression with $\ell_1$-Regularization · J. Mach. Learn. Res. 2018 |
Machine learning › Graph learning
graph neural network |
0.3 | 1 | 2018 | Mean-field theory of graph neural networks in graph partitioning · NeurIPS 2018 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
mean-field approximation |
0.3 | 1 | 2018 | Mean-field theory of graph neural networks in graph partitioning · NeurIPS 2018 |
Bioinformatics and computational biology › neuroscience › neuroinformatics
neural data analysis |
0.3 | 1 | 2018 | Objective and efficient inference for couplings in neuronal networks · NeurIPS 2018 |
Algorithmic game theory and mechanism design › decision theory
multinomial logit model |
0.3 | 1 | 2018 | Accelerating Cross-Validation in Multinomial Logistic Regression with $\ell_1$-Regularization · J. Mach. Learn. Res. 2018 |
Distributed computing theory
message passing |
0.1 | 1 | 2019 | Semi-Analytic Resampling in Lasso · J. Mach. Learn. Res. 2019 |
Information theory › signal processing › compressed sensing › approximate message passing
state evolution |
0.1 | 1 | 2019 | Semi-Analytic Resampling in Lasso · J. Mach. Learn. Res. 2019 |
Graph algorithms and graph theory
graph partitioning |
0.1 | 1 | 2018 | Mean-field theory of graph neural networks in graph partitioning · NeurIPS 2018 |
Mathematical optimization
regularization |
0.1 | 1 | 2018 | Accelerating Cross-Validation in Multinomial Logistic Regression with $\ell_1$-Regularization · J. Mach. Learn. Res. 2018 |
Methods — techniques the papers use, named apart from their topics
unconstrained feature model · 0.9cumulative link model · 0.9state evolution · 0.8stability selection · 0.8message passing · 0.8lasso · 0.8bootstrapped lasso · 0.8statistical mechanics analysis · 0.5replica method · 0.5l1-regularized linear regression · 0.5screening method · 0.3objective bayesian inference · 0.3mean-field theory · 0.3hodgkin-huxley model · 0.3elastic net · 0.3cross-validation · 0.3backpropagation · 0.3$\ell_1$-regularization · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural Collapse in Cumulative Link Models for Ordinal Regression: An Analysis with Unconstrained Feature ModelabstractA phenomenon known as ``Neural Collapse (NC)'' in deep classification tasks, in which the penultimate-layer features and the final classifiers exhibit an extremely simple geometric structure, has recently attracted considerable attention, with the expectation that it can deepen our understanding of how deep neural networks behave. The Unconstrained Feature Model (UFM) has been proposed to explain NC theoretically, and there emerges a growing body of work that extends NC to tasks other than classification and leverages it for practical applications. In this study, we investigate whether a similar phenomenon arises in deep Ordinal Regression (OR) tasks, via combining the cumulative link model for OR and UFM. We show that a phenomenon we call Ordinal Neural Collapse (ONC) indeed emerges and is characterized by the following three properties: (ONC1) all optimal features in the same class collapse to their within-class mean when regularization is applied; (ONC2) these class means align with the classifier, meaning that they collapse onto a one-dimensional subspace; (ONC3) the optimal latent variables (corresponding to logits or preactivations in classification tasks) are aligned according to the class order, and in particular, in the zero-regularization limit, a highly local and simple geometric relationship emerges between the latent variables and the threshold values. We prove these properties analytically within the UFM framework with fixed threshold values and corroborate them empirically across a variety of datasets. We also discuss how these insights can be leveraged in OR, highlighting the use of fixed thresholds. Tomoyuki Obuchi, Toshiyuki Tanaka 0003 |
NeurIPS | 2 |
| 2023 | On Model Selection Consistency of Lasso for High-Dimensional Ising ModelsabstractWe theoretically analyze the model selection consistency of least absolute shrinkage and selection operator (Lasso), both with and without post-thresholding, for high-dimensional Ising models. For random regular (RR) graphs of size $p$ with regular node degree $d$ and uniform couplings $\theta_0$, it is rigorously proved that Lasso without post-thresholding is model selection consistent in the whole paramagnetic phase with the same order of sample complexity $n=\Omega{(d^3\log{p})}$ as that of $\ell_1$-regularized logistic regression ($\ell_1$-LogR). This result is consistent with the conjecture in Meng, Obuchi, and Kabashima 2021 using the non-rigorous replica method from statistical physics and thus complements it with a rigorous proof. For general tree-like graphs, it is demonstrated that the same result as RR graphs can be obtained under mild assumptions of the dependency condition and incoherence condition. Moreover, we provide a rigorous proof of the model selection consistency of Lasso with post-thresholding for general tree-like graphs in the paramagnetic phase without further assumptions on the dependency and incoherence conditions. Experimental results agree well with our theoretical analysis. Xiangming Meng, Tomoyuki Obuchi, Yoshiyuki Kabashima |
AISTATS | 2 |
| 2021 | Sharp Asymptotics of Matrix Sketching for a Rank-One Spiked ModelabstractWe consider matrix sketching for principal component analysis (PCA) with the input data matrices generated by the rank-one spiked model. In the high-dimensional limit, we evaluate the estimation performance of matrix sketching via the replica method. Numerical studies confirm the validity of our results. The obtained result shows that the performance of the estimator undergoes a phase transition at a certain value of the signal strength. A similar asymptotic behavior is well-known for PCA. We demonstrate that our result is a one-parameter generalization of the existing results for PCA. On the basis of our performance evaluation, we also derive the condition for matrix sketching to recover the underlying signal. Fumito Tagashira, Tomoyuki Obuchi, Toshiyuki Tanaka 0003 |
ISIT | 2 |
| 2021 | Ising Model Selection Using $\ell_{1}$-Regularized Linear Regression: A Statistical Mechanics AnalysisabstractWe theoretically analyze the typical learning performance of $\ell_{1}$-regularized linear regression ($\ell_1$-LinR) for Ising model selection using the replica method from statistical mechanics. For typical random regular graphs in the paramagnetic phase, an accurate estimate of the typical sample complexity of $\ell_1$-LinR is obtained. Remarkably, despite the model misspecification, $\ell_1$-LinR is model selection consistent with the same order of sample complexity as $\ell_{1}$-regularized logistic regression ($\ell_1$-LogR), i.e., $M=\mathcal{O}\left(\log N\right)$, where $N$ is the number of variables of the Ising model. Moreover, we provide an efficient method to accurately predict the non-asymptotic behavior of $\ell_1$-LinR for moderate $M, N$, such as precision and recall. Simulations show a fairly good agreement between theoretical predictions and experimental results, even for graphs with many loops, which supports our findings. Although this paper mainly focuses on $\ell_1$-LinR, our method is readily applicable for precisely characterizing the typical learning performances of a wide class of $\ell_{1}$-regularized $M$-estimators including $\ell_1$-LogR and interaction screening. Xiangming Meng, Tomoyuki Obuchi, Yoshiyuki Kabashima |
NeurIPS | 2 |
| 2020 | Inferring Neuronal Couplings From Spiking Data Using a Systematic Procedure With a Statistical CriterionabstractRecent remarkable advances in experimental techniques have provided a background for inferring neuronal couplings from point process data that include a great number of neurons. Here, we propose a systematic procedure for pre- and postprocessing generic point process data in an objective manner to handle data in the framework of a binary simple statistical model, the Ising or generalized McCulloch-Pitts model. The procedure has two steps: (1) determining time bin size for transforming the point process data into discrete-time binary data and (2) screening relevant couplings from the estimated couplings. For the first step, we decide the optimal time bin size by introducing the null hypothesis that all neurons would fire independently, then choosing a time bin size so that the null hypothesis is rejected with the strict criteria. The likelihood associated with the null hypothesis is analytically evaluated and used for the rejection process. For the second postprocessing step, after a certain estimator of coupling is obtained based on the preprocessed data set (any estimator can be used with the proposed procedure), the estimate is compared with many other estimates derived from data sets obtained by randomizing the original data set in the time direction. We accept the original estimate as relevant only if its absolute value is sufficiently larger than those of randomized data sets. These manipulations suppress false positive couplings induced by statistical noise. We apply this inference procedure to spiking data from synthetic and in vitro neuronal networks. The results show that the proposed procedure identifies the presence or absence of synaptic couplings fairly well, including their signs, for the synthetic and experimental data. In particular, the results support that we can infer the physical connections of underlying systems in favorable situations, even when using a simple statistical model. Yu Terada, Tomoyuki Obuchi, Takuya Isomura, Yoshiyuki Kabashima |
Neural Comput. | 2 |
| 2019 | Semi-Analytic Resampling in LassoabstractAn approximate method for conducting resampling in Lasso, the $\ell_1$ penalized linear regression, in a semi-analytic manner is developed, whereby the average over the resampled datasets is directly computed without repeated numerical sampling, thus enabling an inference free of the statistical fluctuations due to sampling finiteness, as well as a significant reduction of computational time. The proposed method is based on a message passing type algorithm, and its fast convergence is guaranteed by the state evolution analysis, when covariates are provided as zero-mean independently and identically distributed Gaussian random variables. It is employed to implement bootstrapped Lasso (Bolasso) and stability selection, both of which are variable selection methods using resampling in conjunction with Lasso, and resolves their disadvantage regarding computational cost. To examine approximation accuracy and efficiency, numerical experiments were carried out using simulated datasets. Moreover, an application to a real-world dataset, the wine quality dataset, is presented. To process such real-world datasets, an objective criterion for determining the relevance of selected variables is also introduced by the addition of noise variables and resampling. MATLAB codes implementing the proposed method are distributed in (Obuchi, 2018). Tomoyuki Obuchi, Yoshiyuki Kabashima |
J. Mach. Learn. Res. | 1 |
| 2018 | Mean-field theory of graph neural networks in graph partitioningabstractA theoretical performance analysis of the graph neural network (GNN) is presented. For classification tasks, the neural network approach has the advantage in terms of flexibility that it can be employed in a data-driven manner, whereas Bayesian inference requires the assumption of a specific model. A fundamental question is then whether GNN has a high accuracy in addition to this flexibility. Moreover, whether the achieved performance is predominately a result of the backpropagation or the architecture itself is a matter of considerable interest. To gain a better insight into these questions, a mean-field theory of a minimal GNN architecture is developed for the graph partitioning problem. This demonstrates a good agreement with numerical experiments. Tatsuro Kawamoto, Masashi Tsubaki, Tomoyuki Obuchi |
NeurIPS | 3 |
| 2018 | Objective and efficient inference for couplings in neuronal networksabstractInferring directional couplings from the spike data of networks is desired in various scientific fields such as neuroscience. Here, we apply a recently proposed objective procedure to the spike data obtained from the Hodgkin-Huxley type models and in vitro neuronal networks cultured in a circular structure. As a result, we succeed in reconstructing synaptic connections accurately from the evoked activity as well as the spontaneous one. To obtain the results, we invent an analytic formula approximately implementing a method of screening relevant couplings. This significantly reduces the computational cost of the screening method employed in the proposed objective procedure, making it possible to treat large-size systems as in this study. Yu Terada, Tomoyuki Obuchi, Takuya Isomura, Yoshiyuki Kabashima |
NeurIPS | 2 |
| 2018 | Accelerating Cross-Validation in Multinomial Logistic Regression with $\ell_1$-RegularizationabstractWe develop an approximate formula for evaluating a cross-validation estimator of predictive likelihood for multinomial logistic regression regularized by an $\ell_1$-norm. This allows us to avoid repeated optimizations required for literally conducting cross-validation; hence, the computational time can be significantly reduced. The formula is derived through a perturbative approach employing the largeness of the data size and the model dimensionality. An extension to the elastic net regularization is also addressed. The usefulness of the approximate formula is demonstrated on simulated data and the ISOLET dataset from the UCI machine learning repository. MATLAB and python codes implementing the approximate formula are distributed in (Obuchi, 2017; Takahashi and Obuchi, 2017). Tomoyuki Obuchi, Yoshiyuki Kabashima |
J. Mach. Learn. Res. | 1 |