Kohei Miyaguchi

dblp:172/7749 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0002-6702-7780ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Reinforcement learning · 43% Learning theory · 16% Optimization for machine learning · 8%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning
0.912025
A Provable Approach for End-to-End Safe Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning › safe reinforcement learning
safe deployment
0.912025
A Provable Approach for End-to-End Safe Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
safe reinforcement learning
0.912025
A Provable Approach for End-to-End Safe Reinforcement Learning · NeurIPS 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Worst-Case Offline Reinforcement Learning with Arbitrary Data Support · NeurIPS 2024
Machine learning › Learning theory
sample complexity
0.812024
Worst-Case Offline Reinforcement Learning with Arbitrary Data Support · NeurIPS 2024
Machine learning › Generative modeling › molecular generation
molecular optimization
0.712023
Biases in Evaluation of Molecular Optimization Methods and Bias Reduction Strategies · ICML 2023
Machine learning › Trustworthy machine learning › interpretability › explainable AI › interpretable neural network
monotonic neural networks
0.612022
Hierarchical Lattice Layer for Partially Monotone Neural Networks · NeurIPS 2022
Machine learning › Deep learning architectures and training
neural network layer design
0.612022
Hierarchical Lattice Layer for Partially Monotone Neural Networks · NeurIPS 2022
Machine learning › Representation and self-supervised learning › representation learning
sequence representation learning
0.612022
Cumulative Stay-time Representation for Electronic Health Records in Medical Event Time Prediction · IJCAI 2022
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.612022
Variational Inference for Discriminative Learning with Generative Modeling of Feature Incompletion · ICLR 2022
Medical and health informatics
electronic health records
0.612022
Cumulative Stay-time Representation for Electronic Health Records in Medical Event Time Prediction · IJCAI 2022
Machine learning › Reinforcement learning › off-policy evaluation
doubly robust estimation
0.512021
Asymptotically Exact Error Characterization of Offline Policy Evaluation with Misspecified Linear Models · NeurIPS 2021
Machine learning › Reinforcement learning
off-policy evaluation
0.512021
Asymptotically Exact Error Characterization of Offline Policy Evaluation with Misspecified Linear Models · NeurIPS 2021
Machine learning › Optimization for machine learning › adaptive optimization
adaptive learning rate
0.412019
Cogra: Concept-Drift-Aware Stochastic Gradient Descent for Time-Series Forecasting · AAAI 2019
Machine learning › Optimization for machine learning
stochastic gradient descent
0.412019
Cogra: Concept-Drift-Aware Stochastic Gradient Descent for Time-Series Forecasting · AAAI 2019
Machine learning › Learning theory
model selection
0.212016
Structure Selection for Convolutive Non-negative Matrix Factorization Using Normalized Maximum Likelihood Coding · ICDM 2016
Recommender systems › collaborative filtering
matrix factorization
0.212016
Structure Selection for Convolutive Non-negative Matrix Factorization Using Normalized Maximum Likelihood Coding · ICDM 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression
0.212022
Hierarchical Lattice Layer for Partially Monotone Neural Networks · NeurIPS 2022
Machine learning › Time series and sequential data › non-stationary environments
concept drift
0.112019
Cogra: Concept-Drift-Aware Stochastic Gradient Descent for Time-Series Forecasting · AAAI 2019
Machine learning › Time series and sequential data › time series analysis
time series forecasting
0.112019
Cogra: Concept-Drift-Aware Stochastic Gradient Descent for Time-Series Forecasting · AAAI 2019

Methods — techniques the papers use, named apart from their topics

neural network · 1.1stochastic gradient descent · 1.0return-conditioned supervised learning · 0.9gaussian process · 0.9worst-case policy optimization · 0.8predictor-based evaluation · 0.7bias reduction · 0.7variational inference · 0.6lattice layer · 0.6generative modeling · 0.6normalized maximum likelihood coding · 0.2minimum description length · 0.2latent variable completion · 0.2
YearPublicationVenuePosition
2026 Detection of unobserved common causes under additive noise models based on NML code for discrete, mixed, and continuous variables
abstract
Abstract Causal discovery from observational data alone in the presence of unobserved common causes is crucial yet challenging. We categorize the causal relationship between two random variables $$X$$ and $$Y$$ into the following four categories: two direct-case cases ( $$X \rightarrow Y$$ or $$X \leftarrow Y$$ ), a case in which $$X$$ and $$Y$$ are independent, and a case in which $$X$$ and $$Y$$ are confounded by unobserved common causes $$C$$ , and aim to decide among them from observational data alone. We call this the Reichenbach problem, since this categorization is valid if the Reichenbach’s common cause principle is assumed to be true. Although many methods have been proposed to address causal structure learning in the presence of $$C$$ , they typically impose assumptions on the nature of $$C$$ that require knowledge of how unobserved confounders behave, and it is difficult to guarantee in many practical cases compared to assumptions about observed variables. In our previous study (Kobayashi et al. in: 2022 IEEE international conference on big data (Big Data). IEEE, pp 45–54, 2022), we proposed a causal discovery method without such assumptions regarding the nature of unmeasured confounders, named $${\mathbf{\mathsf{CLOUD}}}$$ , for discrete data. Building on the Reichenbach’s common cause principle, if neither the two direct-cause models nor the independence model can explain the data well, we can infer the involvement of unobserved common causes in generating $$X$$ and $$Y$$ . To realize this idea in practice, $$\mathbf{\mathsf{CLOUD}}$$ imposes structural assumptions on the direct-cause models such as additive noise models ( $$\mathbf{\mathsf{ANM}}$$ s) (Peters et al. in: Proceedings of the thirteenth international conference on artificial intelligence and statistics. JMLR workshop and conference proceedings, pp 597–604, 2010), while allowing the confounded model to represent any joint distribution on ( $$X$$ , $$Y$$ ) by placing no assumptions on $$C$$ . $$\mathbf{\mathsf{CLOUD}}$$ then employs the normalized maximum likelihood (NML) codelength (Shtar’kov in Problemy Peredachi Informatsii 23(3):3–17, 1987) as an information criterion and compares models of varying capacities. This study extends $$\mathbf{\mathsf{CLOUD}}$$ to include all data types. We provide NML formulations and theoretical guarantees for consistency in model selection for mixed and continuous cases. Through extensive experiments on both synthetic and real-world data, we demonstrated that $$\mathbf{\mathsf{CLOUD}}$$ is more effective than existing methods in inferring causal relationships.
Masatoshi Kobayashi, Kohei Miyaguchi, Shin Matsushima
Data Min. Knowl. Discov.2
2025 A Provable Approach for End-to-End Safe Reinforcement Learning
abstract
A longstanding goal in safe reinforcement learning (RL) is a method to ensure the safety of a policy throughout the entire process, from learning to operation. However, existing safe RL paradigms inherently struggle to achieve this objective. We propose a method, called Provably Lifetime Safe RL (PLS), that integrates offline safe RL with safe policy deployment to address this challenge. Our proposed method learns a policy offline using return-conditioned supervised learning and then deploys the resulting policy while cautiously optimizing a limited set of parameters, known as target returns, using Gaussian processes (GPs). Theoretically, we justify the use of GPs by analyzing the mathematical relationship between target and actual returns. We then prove that PLS finds near-optimal target returns while guaranteeing safety with high probability. Empirically, we demonstrate that PLS outperforms baselines both in safety and reward performance, thereby achieving the longstanding goal to obtain high rewards while ensuring the safety of a policy throughout the lifetime from learning to operation.
Akifumi Wachi, Kohei Miyaguchi, Takumi Tanabe, Rei Sato, Youhei Akimoto
NeurIPS2
2024 Worst-Case Offline Reinforcement Learning with Arbitrary Data Support
abstract
We propose a method of offline reinforcement learning (RL) featuring the performance guarantee without any assumptions on the data support. Under such conditions, estimating or optimizing the conventional performance metric is generally infeasible due to the distributional discrepancy between data and target policy distributions. To address this issue, we employ a worst-case policy value as a new metric and constructively show that the sample complexity bound of $O(\epsilon^{−2})$ is attainable without any data-support conditions, where $\epsilon>0$ is the policy suboptimality in the new metric. Moreover, as the new metric generalizes the conventional one, the algorithm can address standard offline RL tasks without modification. In this context, our sample complexity bound can be seen as a strict improvement on the previous bounds under the single-policy concentrability and the single-policy realizability.
Kohei Miyaguchi
NeurIPS1
2023 Biases in Evaluation of Molecular Optimization Methods and Bias Reduction Strategies
abstract
We are interested in an evaluation methodology for molecular optimization. Given a sample of molecules and their properties of our interest, we wish not only to train a generator of molecules optimized with respect to a target property but also to evaluate its performance accurately. A common practice is to train a predictor of the target property using the sample and apply it to both training and evaluating the generator. However, little is known about its statistical properties, and thus, we are not certain about whether this performance estimate is reliable or not. We theoretically investigate this evaluation methodology and show that it potentially suffers from two biases; one is due to misspecification of the predictor and the other to reusing the same finite sample for training and evaluation. We discuss bias reduction methods for each of the biases, and empirically investigate their effectiveness.
Hiroshi Kajino, Kohei Miyaguchi, Takayuki Osogami
ICML2
2022 Detection of Unobserved Common Cause in Discrete Data Based on the MDL Principle
abstract
Inference of causal structure in random variables from observed data only is a crucial science problem. This study categorizes the causal relationship between two discrete random variables into four categories, which include the case that variables have a common cause and that variables are independent in addition to two directions of the direct causality. Although several existing methods have been proposed for causal inference from a joint distribution of two discrete random variables that can select either direction of the direct causality, methods that can directly infer a causal relationship from the four categories mentioned are limited. We proposed the first method to infer a causal relationship without any assumption on the unobserved confounder in discrete data. Based on the minimum description length (MDL) principle, the proposed method calculates the code length of the observed data for each causal model and then selects a model which yields the minimum code length. We demonstrated that the proposed method is effective in detecting common causes and outperforms the existing methods in terms of the accuracy of inference using synthetic data. We further showed that the proposed method effectively detects an unobserved confounder in real-world data.
Masatoshi Kobayashi, Kohei Miyaguchi, Shin Matsushima
IEEE Big Data2
2022 Variational Inference for Discriminative Learning with Generative Modeling of Feature Incompletion
Kohei Miyaguchi, Takayuki Katsuki, Akira Koseki, Toshiya Iwamori
ICLR1
2022 Cumulative Stay-time Representation for Electronic Health Records in Medical Event Time Prediction
abstract
We address the problem of predicting when a disease will develop, i.e., medical event time (MET), from a patient's electronic health record (EHR). The MET of non-communicable diseases like diabetes is highly correlated to cumulative health conditions, more specifically, how much time the patient spent with specific health conditions in the past. The common time-series representation is indirect in extracting such information from EHR because it focuses on detailed dependencies between values in successive observations, not cumulative information. We propose a novel data representation for EHR called cumulative stay-time representation (CTR), which directly models such cumulative health conditions. We derive a trainable construction of CTR based on neural networks that has the flexibility to fit the target data and scalability to handle high-dimensional EHR. Numerical experiments using synthetic and real-world datasets demonstrate that CTR alone achieves a high prediction performance, and it enhances the performance of existing models when combined with them.
Takayuki Katsuki, Kohei Miyaguchi, Akira Koseki, Toshiya Iwamori, Ryosuke Yanagiya
IJCAI2
2022 Hierarchical Lattice Layer for Partially Monotone Neural Networks
abstract
Partially monotone regression is a regression analysis in which the target values are monotonically increasing with respect to a subset of input features. The TensorFlow Lattice library is one of the standard machine learning libraries for partially monotone regression. It consists of several neural network layers, and its core component is the lattice layer. One of the problems of the lattice layer is that it requires the projected gradient descent algorithm with many constraints to train it. Another problem is that it cannot receive a high-dimensional input vector due to the memory consumption. We propose a novel neural network layer, the hierarchical lattice layer (HLL), as an extension of the lattice layer so that we can use a standard stochastic gradient descent algorithm to train HLL while satisfying monotonicity constraints and so that it can receive a high-dimensional input vector. Our experiments demonstrate that HLL did not sacrifice its prediction performance on real datasets compared with the lattice layer.
Hiroki Yanagisawa, Kohei Miyaguchi, Takayuki Katsuki
NeurIPS2
2021 Asymptotically Exact Error Characterization of Offline Policy Evaluation with Misspecified Linear Models
abstract
We consider the problem of offline policy evaluation~(OPE) with Markov decision processes~(MDPs), where the goal is to estimate the utility of given decision-making policies based on static datasets. Recently, theoretical understanding of OPE has been rapidly advanced under (approximate) realizability assumptions, i.e., where the environments of interest are well approximated with the given hypothetical models. On the other hand, the OPE under unrealizability has not been well understood as much as in the realizable setting despite its importance in real-world applications.To address this issue, we study the behavior of a simple existing OPE method called the linear direct method~(DM) under the unrealizability. Consequently, we obtain an asymptotically exact characterization of the OPE error in a doubly robust form. Leveraging this result, we also establish the nonparametric consistency of the tile-coding estimators under quite mild assumptions.
Kohei Miyaguchi
NeurIPS1
2019 Cogra: Concept-Drift-Aware Stochastic Gradient Descent for Time-Series Forecasting
abstract
We approach the time-series forecasting problem in the presence of concept drift by automatic learning rate tuning of stochastic gradient descent (SGD). The SGD-based approach is preferable to other concept drift algorithms in that it can be applied to any model and it can keep learning efficiently whilst predicting online. Among a number of SGD algorithms, the variance-based SGD (vSGD) can successfully handle concept drift by automatic learning rate tuning, which is reduced to an adaptive mean estimation problem. However, its performance is still limited because of its heuristic mean estimator. In this paper, we present a concept-drift-aware stochastic gradient descent (Cogra), equipped with more theoretically-sound mean estimator called sequential mean tracker (SMT). Our key contribution is that we define a goodness criterion for the mean estimators; SMT is designed to be optimal according to this criterion. As a result of comprehensive experiments, we find that (i) our SMT can estimate the mean better than vSGD’s estimator in the presence of concept drift, and (ii) in terms of predictive performance, Cogra reduces the predictive loss by 16–67% for real-world datasets, indicating that SMT improves the prediction accuracy significantly.
Kohei Miyaguchi, Hiroshi Kajino
AAAI1
2019 Adaptive Minimax Regret against Smooth Logarithmic Losses over High-Dimensional l1-Balls via Envelope Complexity
abstract
We develop a new theoretical framework, the envelope complexity, to analyze the minimax regret with logarithmic loss functions. Within the framework, we derive a Bayesian predictor that adaptively achieves the minimax regret over high-dimensional l1-balls within a factor of two. The prior is newly derived for achieving the minimax regret and called the spike-and-tails (ST) prior as it looks like. The resulting regret bound is so simple that it is completely determined with the smoothness of the loss function and the radius of the balls except with logarithmic factors, and it has a generalized form of existing regret/risk bounds.
Kohei Miyaguchi, Kenji Yamanishi
AISTATS1
2018 High-dimensional penalty selection via minimum description length principle
Kohei Miyaguchi, Kenji Yamanishi
Mach. Learn.1
2017 Detecting changes in streaming data with information-theoretic windowing
abstract
This study addresses the problem of detecting change points in streaming data. Herein, we focus on detecting changes with low complexity and high accuracy to manage large-volume, high-velocity data. To this end, we propose a novel method of detecting changes via data compression, by introducing minimum description length (MDL) change statistics to an adaptive windowing regime. The time complexity of the resulting algorithm, Sequential Compression with Adaptive Windowing (SCAW), is optimal except for a logarithmic factor, and is theoretically justified through exponential upper bounds on error probabilities. Moreover, SCAW can be used to detect arbitrary types of changes by choosing an appropriate compressor. We also introduce the notion of asymptotic reliability as a criterion of change point detection algorithms, and determine SCAW's parameter using this criterion. Finally, we demonstrate the effectiveness of the proposed method in experiments using synthetic and real-world data from markets and industrial machinery.
Ryoya Kaneko, Kohei Miyaguchi, Kenji Yamanishi
IEEE BigData2
2017 Sparse Graphical Modeling via Stochastic Complexity
abstract
Discovering a true sparse model capable of generating data is a challenging yet important problem for understanding the nature of the source of the data. A major part of the challenge arises from the fact that the number of possible sparse models grows exponentially as the dimensionality of the models increases. In this study, we consider a method for estimating the true model over an exponentially large number of sparse models based on the minimum description length principle. We show that a novel criterion derived by continuous relaxation of the stochastic complexity induces selection of the true model by solving the l1-regularization problem for which the hyperparameters are appropriately chosen. Moreover, we provide an efficient optimization algorithm for finding the appropriate hyperparameters and select the sparse model accordingly. The experimental results we obtained for the problem of sparse graphical modeling indicate that the proposed method estimates the true model effectively in comparison to existing methods for choosing hyperparameters to solve the l1-regularization problem.
Kohei Miyaguchi, Shin Matsushima, Kenji Yamanishi
SDM1
2016 Detecting gradual changes from data stream using MDL-change statistics
abstract
In this paper we propose a novel methodology of sequential change detection using the minimum description length (MDL)-change statistics. We first introduce the MDL-change statistics as the difference between the code-lengths with change and that without change. We give a theoretical justification for its use in the scenario of hypothesis testing. In it we evaluate the error probabilities for the MDL-change detection to relate them to the information-theoretic complexities of the probabilistic models and their discrepancy measure. We then convert the MDL-change statistics into the sequential change detection algorithm. It is designed to detect gradual changes as well as abrupt changes from big stream data. We empirically demonstrate the effectiveness of the proposed method by showing that it performs better than existing algorithms for synthetic data. We also show its validity through real problems such as SQL injection detection and failure symptom detection.
Kenji Yamanishi, Kohei Miyaguchi
IEEE BigData2
2016 Structure Selection for Convolutive Non-negative Matrix Factorization Using Normalized Maximum Likelihood Coding
abstract
Convolutive non-negative matrix factorization (CNMF) is a promising method for extracting features from sequential multivariate data. Conventional algorithms for CNMF require that the structure, or the number of bases for expressing the data, be specified in advance. We are concerned with the issue of how we can select the best structure of CNMF from given data. We first introduce a framework of probabilistic modeling of CNMF and reduce this issue to statistical model selection. The problem is here that conventional model selection criteria such as AIC, BIC, MDL cannot straightforwardly be applied since the probabilistic model for CNMF is irregular in the sense that parameters are not uniquely identifiable. We overcome this problem to propose a novel criterion for best structure selection for CNMF. The key idea is to apply the technique of latent variable completion in combination with normalized maximum likelihood coding criterion under the minimum description length principle. We empirically demonstrate the effectiveness of our method using artificial and real data sets.
Atsushi Suzuki 0002, Kohei Miyaguchi, Kenji Yamanishi
ICDM2
2015 On-line detection of continuous changes in stochastic processes
abstract
This paper addresses the issue of detecting changes in stochastic processes. In conventional studies on change detection, it has been explored how to detect discrete changes for which the statistical models of data suddenly change. We are rather concerned with how to detect continuous changes which occurs incrementally over some successive periods. This paper gives a novel methodology for detecting continuous changes. We first define the information-theoretic measure of continuous change and prove that it is invariant with respect to the parametrization of statistical model. We then propose an efficient algorithm to detect continuous changes according to the proposed measure. We demonstrate the effectiveness of our method through the experiments using synthetic data and the applications to security and economic event detection.
Kohei Miyaguchi, Kenji Yamanishi
DSAA1