EDBT 2026 Demo / reviewers in the wild / expert
Masanori Koyama
dblp:151/6113
· DBLP profile ↗
18ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Generative modeling · 30% Representation and self-supervised learning · 18% Optimization for machine learning · 15% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 100% |
Topics — the 30 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
generative adversarial network |
1.5 | 4 | 2020 | Train Sparsely, Generate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN · Int. J. Comput. Vis. 2020 Distributional Concavity Regularization for GANs · ICLR (Poster) 2019 Spectral Normalization for Generative Adversarial Networks · ICLR 2018 |
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning |
1.3 | 2 | 2024 | Neural Fourier Transform: A General Approach to Equivariant Representation Learning · ICLR 2024 Unsupervised Learning of Equivariant Structure from Sequences · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
1.0 | 2 | 2022 | Unsupervised Learning of Equivariant Structure from Sequences · NeurIPS 2022 Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic Perspective · ICML 2020 |
Machine learning › Generative modeling › flow matching
conditional flow matching |
0.9 | 1 | 2025 | Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer Model · NeurIPS 2025 |
Machine learning › Optimization for machine learning
convergence analysis |
0.9 | 1 | 2025 | Flow matching achieves almost minimax optimal convergence · ICLR 2025 |
Machine learning › Generative modeling
flow matching |
0.9 | 1 | 2025 | Flow matching achieves almost minimax optimal convergence · ICLR 2025 |
Machine learning › Learning theory
minimax optimality |
0.9 | 1 | 2025 | Flow matching achieves almost minimax optimal convergence · ICLR 2025 |
Machine learning › Generative modeling
normalizing flow |
0.9 | 1 | 2025 | Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer Model · NeurIPS 2025 |
Machine learning › Optimization for machine learning
optimal transport |
0.9 | 1 | 2025 | Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer Model · NeurIPS 2025 |
Machine learning › Optimization for machine learning › optimal transport
wasserstein distance |
0.9 | 1 | 2025 | Flow matching achieves almost minimax optimal convergence · ICLR 2025 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.8 | 2 | 2020 | Train Sparsely, Generate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN · Int. J. Comput. Vis. 2020 A Graph Theoretic Framework of Recomputation Algorithms for Memory-Efficient Backpropagation · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent factor learning |
0.4 | 1 | 2020 | Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic Perspective · ICML 2020 |
Machine learning › Representation and self-supervised learning › latent representation
latent space structure |
0.4 | 1 | 2020 | Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic Perspective · ICML 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.4 | 1 | 2020 | Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic Perspective · ICML 2020 |
Machine learning › Efficient and distributed learning › model compression
sparse training |
0.4 | 1 | 2020 | Train Sparsely, Generate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN · Int. J. Comput. Vis. 2020 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.4 | 1 | 2019 | Robustness to Adversarial Perturbations in Learning from Incomplete Data · NeurIPS 2019 |
Machine learning › Trustworthy machine learning › robustness
distributionally robust optimization |
0.4 | 1 | 2019 | Robustness to Adversarial Perturbations in Learning from Incomplete Data · NeurIPS 2019 |
Machine learning › Transfer learning and domain adaptation › test-time adaptation
entropy minimization |
0.4 | 1 | 2019 | Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2019 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning |
0.4 | 1 | 2019 | A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning · ICML 2019 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.4 | 1 | 2019 | Optuna: A Next-generation Hyperparameter Optimization Framework · KDD 2019 |
Machine learning › Generative modeling › generative model
probabilistic generative model |
0.4 | 1 | 2019 | A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning · ICML 2019 |
Machine learning › Deep learning architectures and training
regularization |
0.4 | 1 | 2019 | Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2019 |
Machine learning › Trustworthy machine learning
robustness |
0.4 | 1 | 2019 | Robustness to Adversarial Perturbations in Learning from Incomplete Data · NeurIPS 2019 |
Machine learning › Learning paradigms
semi-supervised learning |
0.4 | 1 | 2019 | Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2019 |
Machine learning › Learning theory › computational learning theory › learnability
semi-supervised learning theory |
0.4 | 1 | 2019 | Robustness to Adversarial Perturbations in Learning from Incomplete Data · NeurIPS 2019 |
Machine learning › Deep learning architectures and training › regularization
training regularization |
0.4 | 1 | 2019 | Distributional Concavity Regularization for GANs · ICLR (Poster) 2019 |
Machine learning › Generative modeling
variational autoencoder |
0.4 | 1 | 2019 | A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning · ICML 2019 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.4 | 1 | 2019 | A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning · ICML 2019 |
Graph algorithms and graph theory › graph algorithms
graph-theoretic optimization |
0.4 | 1 | 2019 | A Graph Theoretic Framework of Recomputation Algorithms for Memory-Efficient Backpropagation · NeurIPS 2019 |
Graph algorithms and graph theory
graph theory |
0.4 | 1 | 2019 | A Graph Theoretic Framework of Recomputation Algorithms for Memory-Efficient Backpropagation · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
wasserstein distance analysis · 0.9optimal transport · 0.9flow matching · 0.9adversarial training · 0.7simultaneous block-diagonalization · 0.6representation theory · 0.6encoder-decoder model · 0.6sparse training · 0.4mask variables · 0.4information bottleneck · 0.4graph theory · 0.4dynamic programming · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Flow matching achieves almost minimax optimal convergenceabstractFlow matching (FM) has gained significant attention as a simulation-free generative model. Unlike diffusion models, which are based on stochastic differential equations, FM employs a simpler approach by solving an ordinary differential equation with an initial condition from a normal distribution, thus streamlining the sample generation process. This paper discusses the convergence properties of FM in terms of the $p$-Wasserstein distance, a measure of distributional discrepancy. We establish that FM can achieve an almost minimax optimal convergence rate for $1 \leq p \leq 2$, presenting the first theoretical evidence that FM can reach convergence rates comparable to those of diffusion models. Our analysis extends existing frameworks by examining a broader class of mean and variance functions for the vector fields and identifies specific conditions necessary to attain these optimal rates. Kenji Fukumizu, Taiji Suzuki, Noboru Isobe, Kazusato Oko, Masanori Koyama |
ICLR | 5 |
| 2025 | Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer ModelabstractIn this paper, we propose a flow-based method for learning all-to-all transfer maps among conditional distributions that approximates pairwise optimal transport. The proposed method addresses the challenge of handling the case of continuous conditions, which often involve a large set of conditions with sparse empirical observations per condition. We introduce a novel cost function that enables simultaneous learning of optimal transports for all pairs of conditional distributions. Our method is supported by a theoretical guarantee that, in the limit, it converges to the pairwise optimal transports among infinite pairs of conditional distributions. The learned transport maps are subsequently used to couple data points in conditional flow matching. We demonstrate the effectiveness of this method on synthetic and benchmark datasets, as well as on chemical datasets in which continuous physical properties are defined as conditions. Kotaro Ikeda, Masanori Koyama, Jinzhe Zhang, Kohei Hayashi, Kenji Fukumizu |
NeurIPS | 2 |
| 2024 | Neural Fourier Transform: A General Approach to Equivariant Representation LearningabstractSymmetry learning has proven to be an effective approach for extracting the hidden structure of data, with the concept of equivariance relation playing the central role.
However, most of the current studies are built on architectural theory and corresponding assumptions on the form of data.
We propose Neural Fourier Transform (NFT), a general framework of learning the latent linear action of the group without assuming explicit knowledge of how the group acts on data.
We present the theoretical foundations of NFT and show that
the existence of a linear equivariant feature, which has been assumed ubiquitously in equivariance learning, is equivalent to the existence of a group invariant kernel on the dataspace.
We also provide experimental results to demonstrate the application of NFT in typical scenarios with varying levels of knowledge about the acting group. Masanori Koyama, Kenji Fukumizu, Kohei Hayashi, Takeru Miyato |
ICLR | 1 |
| 2022 | Unsupervised Learning of Equivariant Structure from SequencesabstractIn this study, we present \textit{meta-sequential prediction} (MSP), an unsupervised framework to learn the symmetry from the time sequence of length at least three. Our method leverages the stationary property~(e.g. constant velocity, constant acceleration) of the time sequence to learn the underlying equivariant structure of the dataset by simply training the encoder-decoder model to be able to predict the future observations. We will demonstrate that, with our framework, the hidden disentangled structure of the dataset naturally emerges as a by-product by applying \textit{simultaneous block-diagonalization} to the transition operators in the latent space, the procedure which is commonly used in representation theory to decompose the feature-space based on the type of response to group actions.We will showcase our method from both empirical and theoretical perspectives.Our result suggests that finding a simple structured relation and learning a model with extrapolation capability are two sides of the same coin. The code is available at https://github.com/takerum/metasequentialprediction. Takeru Miyato, Masanori Koyama, Kenji Fukumizu |
NeurIPS | 2 |
| 2021 | Reconnaissance for Reinforcement Learning with Safety Constraints
Shin-ichi Maeda, Hayato Watahiki, Shintarou Okada, Masanori Koyama, Prabhat Nagarajan |
ECML/PKDD (2) | 5 |
| 2020 | Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic PerspectiveabstractLearning controllable and generalizable representation of multivariate data with desired structural properties remains a fundamental problem in machine learning. In this paper, we present a novel framework for learning generative models with various underlying structures in the latent space. Learning controllable and generalizable representation of multivariate data with desired structural properties remains a fundamental problem in machine learning. In this paper, we present a novel framework for learning generative models with various underlying structures in the latent space. We represent the inductive bias in the form of mask variables to model the dependency structure in the graphical model and extend the theory of multivariate information bottleneck (Friedman et al., 2001) to enforce it. Our model provides a principled approach to learn a set of semantically meaningful latent factors that reflect various types of desired structures like capturing correlation or encoding invariance, while also offering the flexibility to automatically estimate the dependency structure from data. We show that our framework unifies many existing generative models and can be applied to a variety of tasks, including multimodal data modeling, algorithmic fairness, and out-of-distribution generalization. Ruixiang Zhang, Masanori Koyama, Katsuhiko Ishiguro |
ICML | 2 |
| 2020 | Train Sparsely, Generate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN
Masaki Saito, Shunta Saito, Masanori Koyama, Sosuke Kobayashi |
Int. J. Comput. Vis. | 3 |
| 2019 | Distributional Concavity Regularization for GANs
Shoichiro Yamaguchi, Masanori Koyama |
ICLR (Poster) | 2 |
| 2019 | A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based LearningabstractHyperbolic space is a geometry that is known to be well-suited for representation learning of data with an underlying hierarchical structure. In this paper, we present a novel hyperbolic distribution called hyperbolic wrapped distribution, a wrapped normal distribution on hyperbolic space whose density can be evaluated analytically and differentiated with respect to the parameters. Our distribution enables the gradient-based learning of the probabilistic models on hyperbolic space that could never have been considered before. Also, we can sample from this hyperbolic probability distribution without resorting to auxiliary means like rejection sampling. As applications of our distribution, we develop a hyperbolic-analog of variational autoencoder and a method of probabilistic word embedding on hyperbolic space. We demonstrate the efficacy of our distribution on various datasets including MNIST, Atari 2600 Breakout, and WordNet. Yoshihiro Nagano, Shoichiro Yamaguchi, Yasuhiro Fujita 0001, Masanori Koyama |
ICML | 4 |
| 2019 | Optuna: A Next-generation Hyperparameter Optimization FrameworkabstractThe purpose of this study is to introduce new design-criteria for next-generation hyperparameter optimization software. The criteria we propose include (1) define-by-run API that allows users to construct the parameter search space dynamically, (2) efficient implementation of both searching and pruning strategies, and (3) easy-to-setup, versatile architecture that can be deployed for various purposes, ranging from scalable distributed computing to light-weight experiment conducted via interactive interface. In order to prove our point, we will introduce Optuna, an optimization software which is a culmination of our effort in the development of a next generation optimization software. As an optimization software designed with define-by-run principle, Optuna is particularly the first of its kind. We will present the design-techniques that became necessary in the development of the software that meets the above criteria, and demonstrate the power of our new design through experimental results and real world applications. Our software is available under the MIT license (https://github.com/pfnet/optuna/). Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, Masanori Koyama |
KDD | 5 |
| 2019 | A Graph Theoretic Framework of Recomputation Algorithms for Memory-Efficient BackpropagationabstractRecomputation algorithms collectively refer to a family of methods that aims to reduce the memory consumption of the backpropagation by selectively discarding the intermediate results of the forward propagation and recomputing the discarded results as needed. In this paper, we will propose a novel and efficient recomputation method that can be applied to a wider range of neural nets than previous methods. We use the language of graph theory to formalize the general recomputation problem of minimizing the computational overhead under a fixed memory budget constraint, and provide a dynamic programming solution to the problem. Our method can reduce the peak memory consumption on various benchmark networks by $36\%\sim81\%$, which outperforms the reduction achieved by other methods. Mitsuru Kusumoto, Takuya Inoue, Gentaro Watanabe, Takuya Akiba, Masanori Koyama |
NeurIPS | 5 |
| 2019 | Robustness to Adversarial Perturbations in Learning from Incomplete DataabstractWhat is the role of unlabeled data in an inference problem, when the presumed underlying distribution is adversarially perturbed? To provide a concrete answer to this question, this paper unifies two major learning frameworks: Semi-Supervised Learning (SSL) and Distributionally Robust Learning (DRL). We develop a generalization theory for our framework based on a number of novel complexity measures, such as an adversarial extension of Rademacher complexity and its semi-supervised analogue. Moreover, our analysis is able to quantify the role of unlabeled data in the generalization under a more general condition compared to the existing theoretical works in SSL. Based on our framework, we also present a hybrid of DRL and EM algorithms that has a guaranteed convergence rate. When implemented with deep neural networks, our method shows a comparable performance to those of the state-of-the-art on a number of real-world benchmark datasets. Amir Najafi 0002, Shin-ichi Maeda, Masanori Koyama, Takeru Miyato |
NeurIPS | 3 |
| 2019 | Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised LearningabstractWe propose a new regularization method based on virtual adversarial loss: a new measure of local smoothness of the conditional label distribution given input. Virtual adversarial loss is defined as the robustness of the conditional label distribution around each input data point against local perturbation. Unlike adversarial training, our method defines the adversarial direction without label information and is hence applicable to semi-supervised learning. Because the directions in which we smooth the model are only "virtually" adversarial, we call our method virtual adversarial training (VAT). The computational cost of VAT is relatively low. For neural networks, the approximated gradient of virtual adversarial loss can be computed with no more than two pairs of forward- and back-propagations. In our experiments, we applied VAT to supervised and semi-supervised learning tasks on multiple benchmark datasets. With a simple enhancement of the algorithm based on the entropy minimization principle, our VAT achieves state-of-the-art performance for semi-supervised learning tasks on SVHN and CIFAR-10. Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, Shin Ishii |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | cGANs with Projection Discriminator
Takeru Miyato, Masanori Koyama |
ICLR (Poster) | 2 |
| 2018 | Spectral Normalization for Generative Adversarial Networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, Yuichi Yoshida |
ICLR | 3 |
| 2015 | Principal Sensitivity Analysis
Sotetsu Koyamada, Masanori Koyama, Ken Nakae, Shin Ishii |
PAKDD (1) | 2 |
| 2015 | Efficient Monte Carlo Image Analysis for the Location of Vascular EntityabstractTubular shaped networks appear not only in medical images like X-ray-, time-of-flight MRI- or CT-angiograms but also in microscopic images of neuronal networks. We present EMILOVE (Efficient Monte-carlo Image-analysis for the Location Of Vascular Entity), a novel modeling algorithm for tubular networks in biomedical images. The model is constructed using tablet shaped particles and edges connecting them. The particles encode the intrinsic information of tubular structure, including position, scale and orientation. The edges connecting the particles determine the topology of the networks. For simulated data, EMILOVE was able to accurately extract the tubular network. EMILOVE showed high performance in real data as well; it successfully modeled vascular networks in real cerebral X-ray and time-of-flight MRI angiograms. We also show some promising, preliminary results on microscopic images of neurons. Henrik Skibbe, Marco Reisert, Shin-ichi Maeda, Masanori Koyama, Shigeyuki Oba, Kei Ito, Shin Ishii |
IEEE Trans. Medical Imaging | 4 |
| 2014 | A Statistical Method of Identifying Interactions in Neuron-Glia Systems Based on Functional Multicell Ca2+ ImagingabstractCrosstalk between neurons and glia may constitute a significant part of information processing in the brain. We present a novel method of statistically identifying interactions in a neuron-glia network. We attempted to identify neuron-glia interactions from neuronal and glial activities via maximum-a-posteriori (MAP)-based parameter estimation by developing a generalized linear model (GLM) of a neuron-glia network. The interactions in our interest included functional connectivity and response functions. We evaluated the cross-validated likelihood of GLMs that resulted from the addition or removal of connections to confirm the existence of specific neuron-to-glia or glia-to-neuron connections. We only accepted addition or removal when the modification improved the cross-validated likelihood. We applied the method to a high-throughput, multicellular in vitro Ca2+ imaging dataset obtained from the CA3 region of a rat hippocampus, and then evaluated the reliability of connectivity estimates using a statistical test based on a surrogate method. Our findings based on the estimated connectivity were in good agreement with currently available physiological knowledge, suggesting our method can elucidate undiscovered functions of neuron-glia systems. Ken Nakae, Yuji Ikegaya, Tomoe Ishikawa, Shigeyuki Oba, Hidetoshi Urakubo, Masanori Koyama, Shin Ishii |
PLoS Comput. Biol. | 6 |