Masanori Koyama

dblp:151/6113 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Generative modeling · 30% Representation and self-supervised learning · 18% Optimization for machine learning · 15%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative adversarial network
1.542020
Train Sparsely, Generate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN · Int. J. Comput. Vis. 2020
Distributional Concavity Regularization for GANs · ICLR (Poster) 2019
Spectral Normalization for Generative Adversarial Networks · ICLR 2018
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning
1.322024
Neural Fourier Transform: A General Approach to Equivariant Representation Learning · ICLR 2024
Unsupervised Learning of Equivariant Structure from Sequences · NeurIPS 2022
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
1.022022
Unsupervised Learning of Equivariant Structure from Sequences · NeurIPS 2022
Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic Perspective · ICML 2020
Machine learning › Generative modeling › flow matching
conditional flow matching
0.912025
Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer Model · NeurIPS 2025
Machine learning › Optimization for machine learning
convergence analysis
0.912025
Flow matching achieves almost minimax optimal convergence · ICLR 2025
Machine learning › Generative modeling
flow matching
0.912025
Flow matching achieves almost minimax optimal convergence · ICLR 2025
Machine learning › Learning theory
minimax optimality
0.912025
Flow matching achieves almost minimax optimal convergence · ICLR 2025
Machine learning › Generative modeling
normalizing flow
0.912025
Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer Model · NeurIPS 2025
Machine learning › Optimization for machine learning
optimal transport
0.912025
Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer Model · NeurIPS 2025
Machine learning › Optimization for machine learning › optimal transport
wasserstein distance
0.912025
Flow matching achieves almost minimax optimal convergence · ICLR 2025
Machine learning › Efficient and distributed learning
memory-efficient training
0.822020
Train Sparsely, Generate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN · Int. J. Comput. Vis. 2020
A Graph Theoretic Framework of Recomputation Algorithms for Memory-Efficient Backpropagation · NeurIPS 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent factor learning
0.412020
Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic Perspective · ICML 2020
Machine learning › Representation and self-supervised learning › latent representation
latent space structure
0.412020
Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic Perspective · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.412020
Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic Perspective · ICML 2020
Machine learning › Efficient and distributed learning › model compression
sparse training
0.412020
Train Sparsely, Generate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN · Int. J. Comput. Vis. 2020
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412019
Robustness to Adversarial Perturbations in Learning from Incomplete Data · NeurIPS 2019
Machine learning › Trustworthy machine learning › robustness
distributionally robust optimization
0.412019
Robustness to Adversarial Perturbations in Learning from Incomplete Data · NeurIPS 2019
Machine learning › Transfer learning and domain adaptation › test-time adaptation
entropy minimization
0.412019
Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › manifold learning › geometric representation learning
hyperbolic representation learning
0.412019
A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning · ICML 2019
Machine learning › Optimization for machine learning
hyperparameter optimization
0.412019
Optuna: A Next-generation Hyperparameter Optimization Framework · KDD 2019
Machine learning › Generative modeling › generative model
probabilistic generative model
0.412019
A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning · ICML 2019
Machine learning › Deep learning architectures and training
regularization
0.412019
Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Machine learning › Trustworthy machine learning
robustness
0.412019
Robustness to Adversarial Perturbations in Learning from Incomplete Data · NeurIPS 2019
Machine learning › Learning paradigms
semi-supervised learning
0.412019
Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Machine learning › Learning theory › computational learning theory › learnability
semi-supervised learning theory
0.412019
Robustness to Adversarial Perturbations in Learning from Incomplete Data · NeurIPS 2019
Machine learning › Deep learning architectures and training › regularization
training regularization
0.412019
Distributional Concavity Regularization for GANs · ICLR (Poster) 2019
Machine learning › Generative modeling
variational autoencoder
0.412019
A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning · ICML 2019
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.412019
A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning · ICML 2019
Graph algorithms and graph theory › graph algorithms
graph-theoretic optimization
0.412019
A Graph Theoretic Framework of Recomputation Algorithms for Memory-Efficient Backpropagation · NeurIPS 2019
Graph algorithms and graph theory
graph theory
0.412019
A Graph Theoretic Framework of Recomputation Algorithms for Memory-Efficient Backpropagation · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

wasserstein distance analysis · 0.9optimal transport · 0.9flow matching · 0.9adversarial training · 0.7simultaneous block-diagonalization · 0.6representation theory · 0.6encoder-decoder model · 0.6sparse training · 0.4mask variables · 0.4information bottleneck · 0.4graph theory · 0.4dynamic programming · 0.4
YearPublicationVenuePosition
2025 Flow matching achieves almost minimax optimal convergence
abstract
Flow matching (FM) has gained significant attention as a simulation-free generative model. Unlike diffusion models, which are based on stochastic differential equations, FM employs a simpler approach by solving an ordinary differential equation with an initial condition from a normal distribution, thus streamlining the sample generation process. This paper discusses the convergence properties of FM in terms of the $p$-Wasserstein distance, a measure of distributional discrepancy. We establish that FM can achieve an almost minimax optimal convergence rate for $1 \leq p \leq 2$, presenting the first theoretical evidence that FM can reach convergence rates comparable to those of diffusion models. Our analysis extends existing frameworks by examining a broader class of mean and variance functions for the vector fields and identifies specific conditions necessary to attain these optimal rates.
Kenji Fukumizu, Taiji Suzuki, Noboru Isobe, Kazusato Oko, Masanori Koyama
ICLR5
2025 Pairwise Optimal Transports for Training All-to-All Flow-Based Condition Transfer Model
abstract
In this paper, we propose a flow-based method for learning all-to-all transfer maps among conditional distributions that approximates pairwise optimal transport. The proposed method addresses the challenge of handling the case of continuous conditions, which often involve a large set of conditions with sparse empirical observations per condition. We introduce a novel cost function that enables simultaneous learning of optimal transports for all pairs of conditional distributions. Our method is supported by a theoretical guarantee that, in the limit, it converges to the pairwise optimal transports among infinite pairs of conditional distributions. The learned transport maps are subsequently used to couple data points in conditional flow matching. We demonstrate the effectiveness of this method on synthetic and benchmark datasets, as well as on chemical datasets in which continuous physical properties are defined as conditions.
Kotaro Ikeda, Masanori Koyama, Jinzhe Zhang, Kohei Hayashi, Kenji Fukumizu
NeurIPS2
2024 Neural Fourier Transform: A General Approach to Equivariant Representation Learning
abstract
Symmetry learning has proven to be an effective approach for extracting the hidden structure of data, with the concept of equivariance relation playing the central role. However, most of the current studies are built on architectural theory and corresponding assumptions on the form of data. We propose Neural Fourier Transform (NFT), a general framework of learning the latent linear action of the group without assuming explicit knowledge of how the group acts on data. We present the theoretical foundations of NFT and show that the existence of a linear equivariant feature, which has been assumed ubiquitously in equivariance learning, is equivalent to the existence of a group invariant kernel on the dataspace. We also provide experimental results to demonstrate the application of NFT in typical scenarios with varying levels of knowledge about the acting group.
Masanori Koyama, Kenji Fukumizu, Kohei Hayashi, Takeru Miyato
ICLR1
2022 Unsupervised Learning of Equivariant Structure from Sequences
abstract
In this study, we present \textit{meta-sequential prediction} (MSP), an unsupervised framework to learn the symmetry from the time sequence of length at least three. Our method leverages the stationary property~(e.g. constant velocity, constant acceleration) of the time sequence to learn the underlying equivariant structure of the dataset by simply training the encoder-decoder model to be able to predict the future observations. We will demonstrate that, with our framework, the hidden disentangled structure of the dataset naturally emerges as a by-product by applying \textit{simultaneous block-diagonalization} to the transition operators in the latent space, the procedure which is commonly used in representation theory to decompose the feature-space based on the type of response to group actions.We will showcase our method from both empirical and theoretical perspectives.Our result suggests that finding a simple structured relation and learning a model with extrapolation capability are two sides of the same coin. The code is available at https://github.com/takerum/metasequentialprediction.
Takeru Miyato, Masanori Koyama, Kenji Fukumizu
NeurIPS2
2021 Reconnaissance for Reinforcement Learning with Safety Constraints
Shin-ichi Maeda, Hayato Watahiki, Shintarou Okada, Masanori Koyama, Prabhat Nagarajan
ECML/PKDD (2)5
2020 Learning Structured Latent Factors from Dependent Data:A Generative Model Framework from Information-Theoretic Perspective
abstract
Learning controllable and generalizable representation of multivariate data with desired structural properties remains a fundamental problem in machine learning. In this paper, we present a novel framework for learning generative models with various underlying structures in the latent space. Learning controllable and generalizable representation of multivariate data with desired structural properties remains a fundamental problem in machine learning. In this paper, we present a novel framework for learning generative models with various underlying structures in the latent space. We represent the inductive bias in the form of mask variables to model the dependency structure in the graphical model and extend the theory of multivariate information bottleneck (Friedman et al., 2001) to enforce it. Our model provides a principled approach to learn a set of semantically meaningful latent factors that reflect various types of desired structures like capturing correlation or encoding invariance, while also offering the flexibility to automatically estimate the dependency structure from data. We show that our framework unifies many existing generative models and can be applied to a variety of tasks, including multimodal data modeling, algorithmic fairness, and out-of-distribution generalization.
Ruixiang Zhang, Masanori Koyama, Katsuhiko Ishiguro
ICML2
2020 Train Sparsely, Generate Densely: Memory-Efficient Unsupervised Training of High-Resolution Temporal GAN
Masaki Saito, Shunta Saito, Masanori Koyama, Sosuke Kobayashi
Int. J. Comput. Vis.3
2019 Distributional Concavity Regularization for GANs
Shoichiro Yamaguchi, Masanori Koyama
ICLR (Poster)2
2019 A Wrapped Normal Distribution on Hyperbolic Space for Gradient-Based Learning
abstract
Hyperbolic space is a geometry that is known to be well-suited for representation learning of data with an underlying hierarchical structure. In this paper, we present a novel hyperbolic distribution called hyperbolic wrapped distribution, a wrapped normal distribution on hyperbolic space whose density can be evaluated analytically and differentiated with respect to the parameters. Our distribution enables the gradient-based learning of the probabilistic models on hyperbolic space that could never have been considered before. Also, we can sample from this hyperbolic probability distribution without resorting to auxiliary means like rejection sampling. As applications of our distribution, we develop a hyperbolic-analog of variational autoencoder and a method of probabilistic word embedding on hyperbolic space. We demonstrate the efficacy of our distribution on various datasets including MNIST, Atari 2600 Breakout, and WordNet.
Yoshihiro Nagano, Shoichiro Yamaguchi, Yasuhiro Fujita 0001, Masanori Koyama
ICML4
2019 Optuna: A Next-generation Hyperparameter Optimization Framework
abstract
The purpose of this study is to introduce new design-criteria for next-generation hyperparameter optimization software. The criteria we propose include (1) define-by-run API that allows users to construct the parameter search space dynamically, (2) efficient implementation of both searching and pruning strategies, and (3) easy-to-setup, versatile architecture that can be deployed for various purposes, ranging from scalable distributed computing to light-weight experiment conducted via interactive interface. In order to prove our point, we will introduce Optuna, an optimization software which is a culmination of our effort in the development of a next generation optimization software. As an optimization software designed with define-by-run principle, Optuna is particularly the first of its kind. We will present the design-techniques that became necessary in the development of the software that meets the above criteria, and demonstrate the power of our new design through experimental results and real world applications. Our software is available under the MIT license (https://github.com/pfnet/optuna/).
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, Masanori Koyama
KDD5
2019 A Graph Theoretic Framework of Recomputation Algorithms for Memory-Efficient Backpropagation
abstract
Recomputation algorithms collectively refer to a family of methods that aims to reduce the memory consumption of the backpropagation by selectively discarding the intermediate results of the forward propagation and recomputing the discarded results as needed. In this paper, we will propose a novel and efficient recomputation method that can be applied to a wider range of neural nets than previous methods. We use the language of graph theory to formalize the general recomputation problem of minimizing the computational overhead under a fixed memory budget constraint, and provide a dynamic programming solution to the problem. Our method can reduce the peak memory consumption on various benchmark networks by $36\%\sim81\%$, which outperforms the reduction achieved by other methods.
Mitsuru Kusumoto, Takuya Inoue, Gentaro Watanabe, Takuya Akiba, Masanori Koyama
NeurIPS5
2019 Robustness to Adversarial Perturbations in Learning from Incomplete Data
abstract
What is the role of unlabeled data in an inference problem, when the presumed underlying distribution is adversarially perturbed? To provide a concrete answer to this question, this paper unifies two major learning frameworks: Semi-Supervised Learning (SSL) and Distributionally Robust Learning (DRL). We develop a generalization theory for our framework based on a number of novel complexity measures, such as an adversarial extension of Rademacher complexity and its semi-supervised analogue. Moreover, our analysis is able to quantify the role of unlabeled data in the generalization under a more general condition compared to the existing theoretical works in SSL. Based on our framework, we also present a hybrid of DRL and EM algorithms that has a guaranteed convergence rate. When implemented with deep neural networks, our method shows a comparable performance to those of the state-of-the-art on a number of real-world benchmark datasets.
Amir Najafi 0002, Shin-ichi Maeda, Masanori Koyama, Takeru Miyato
NeurIPS3
2019 Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning
abstract
We propose a new regularization method based on virtual adversarial loss: a new measure of local smoothness of the conditional label distribution given input. Virtual adversarial loss is defined as the robustness of the conditional label distribution around each input data point against local perturbation. Unlike adversarial training, our method defines the adversarial direction without label information and is hence applicable to semi-supervised learning. Because the directions in which we smooth the model are only "virtually" adversarial, we call our method virtual adversarial training (VAT). The computational cost of VAT is relatively low. For neural networks, the approximated gradient of virtual adversarial loss can be computed with no more than two pairs of forward- and back-propagations. In our experiments, we applied VAT to supervised and semi-supervised learning tasks on multiple benchmark datasets. With a simple enhancement of the algorithm based on the entropy minimization principle, our VAT achieves state-of-the-art performance for semi-supervised learning tasks on SVHN and CIFAR-10.
Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, Shin Ishii
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 cGANs with Projection Discriminator
Takeru Miyato, Masanori Koyama
ICLR (Poster)2
2018 Spectral Normalization for Generative Adversarial Networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, Yuichi Yoshida
ICLR3
2015 Principal Sensitivity Analysis
Sotetsu Koyamada, Masanori Koyama, Ken Nakae, Shin Ishii
PAKDD (1)2
2015 Efficient Monte Carlo Image Analysis for the Location of Vascular Entity
abstract
Tubular shaped networks appear not only in medical images like X-ray-, time-of-flight MRI- or CT-angiograms but also in microscopic images of neuronal networks. We present EMILOVE (Efficient Monte-carlo Image-analysis for the Location Of Vascular Entity), a novel modeling algorithm for tubular networks in biomedical images. The model is constructed using tablet shaped particles and edges connecting them. The particles encode the intrinsic information of tubular structure, including position, scale and orientation. The edges connecting the particles determine the topology of the networks. For simulated data, EMILOVE was able to accurately extract the tubular network. EMILOVE showed high performance in real data as well; it successfully modeled vascular networks in real cerebral X-ray and time-of-flight MRI angiograms. We also show some promising, preliminary results on microscopic images of neurons.
Henrik Skibbe, Marco Reisert, Shin-ichi Maeda, Masanori Koyama, Shigeyuki Oba, Kei Ito, Shin Ishii
IEEE Trans. Medical Imaging4
2014 A Statistical Method of Identifying Interactions in Neuron-Glia Systems Based on Functional Multicell Ca2+ Imaging
abstract
Crosstalk between neurons and glia may constitute a significant part of information processing in the brain. We present a novel method of statistically identifying interactions in a neuron-glia network. We attempted to identify neuron-glia interactions from neuronal and glial activities via maximum-a-posteriori (MAP)-based parameter estimation by developing a generalized linear model (GLM) of a neuron-glia network. The interactions in our interest included functional connectivity and response functions. We evaluated the cross-validated likelihood of GLMs that resulted from the addition or removal of connections to confirm the existence of specific neuron-to-glia or glia-to-neuron connections. We only accepted addition or removal when the modification improved the cross-validated likelihood. We applied the method to a high-throughput, multicellular in vitro Ca2+ imaging dataset obtained from the CA3 region of a rat hippocampus, and then evaluated the reliability of connectivity estimates using a statistical test based on a surrogate method. Our findings based on the estimated connectivity were in good agreement with currently available physiological knowledge, suggesting our method can elucidate undiscovered functions of neuron-glia systems.
Ken Nakae, Yuji Ikegaya, Tomoe Ishikawa, Shigeyuki Oba, Hidetoshi Urakubo, Masanori Koyama, Shin Ishii
PLoS Comput. Biol.6