VLDB 2026 Research / reviewers in the wild / expert
Isao Ishikawa
dblp:220/5361
· DBLP profile ↗
10ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Learning theory · 38% Deep learning architectures and training · 30% Kernel, tree and ensemble methods · 11% | |
| Theoretical computer science
3 papers |
Algorithms and data structures · 57% Mathematical optimization · 42% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation |
2.5 | 4 | 2025 | Deep Ridgelet Transform and Unified Universality Theorem for Deep and Shallow Joint-Group-Equivariant Machines · ICML 2025 Universal Approximation Property of Invertible Neural Networks · J. Mach. Learn. Res. 2023 Fully-Connected Network on Noncompact Symmetric Space and Ridgelet Transform based on Helgason-Fourier Analysis · ICML 2022 |
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network |
1.1 | 2 | 2023 | Universal Approximation Property of Invertible Neural Networks · J. Mach. Learn. Res. 2023 Coupling-based Invertible Neural Networks Are Universal Diffeomorphism Approximators · NeurIPS 2020 |
Machine learning › Deep learning architectures and training
equivariant neural network |
0.9 | 1 | 2025 | Deep Ridgelet Transform and Unified Universality Theorem for Deep and Shallow Joint-Group-Equivariant Machines · ICML 2025 |
Machine learning › Kernel, tree and ensemble methods
kernel methods |
0.8 | 2 | 2021 | Reproducing kernel Hilbert C*-module and kernel mean embeddings · J. Mach. Learn. Res. 2021 Metric on Nonlinear Dynamical Systems with Perron-Frobenius Operators · NeurIPS 2018 |
Machine learning › Learning theory
generalization bounds |
0.8 | 1 | 2024 | Koopman-based generalization bound: New aspect for full-rank weights · ICLR 2024 |
Machine learning › Learning theory › neural network theory
neural network generalization |
0.8 | 1 | 2024 | Koopman-based generalization bound: New aspect for full-rank weights · ICLR 2024 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.6 | 1 | 2022 | Universality of Group Convolutional Neural Networks Based on Ridgelet Analysis on Groups · NeurIPS 2022 |
Computer vision › 3D vision
geometric deep learning |
0.6 | 1 | 2022 | Fully-Connected Network on Noncompact Symmetric Space and Ridgelet Transform based on Helgason-Fourier Analysis · ICML 2022 |
Machine learning › Deep learning architectures and training › equivariant neural network
group convolutional network |
0.6 | 1 | 2022 | Universality of Group Convolutional Neural Networks Based on Ridgelet Analysis on Groups · NeurIPS 2022 |
Machine learning › Deep learning architectures and training › equivariant neural network
symmetric neural networks |
0.6 | 1 | 2022 | Fully-Connected Network on Noncompact Symmetric Space and Ridgelet Transform based on Helgason-Fourier Analysis · ICML 2022 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel mean embedding |
0.5 | 1 | 2021 | Reproducing kernel Hilbert C*-module and kernel mean embeddings · J. Mach. Learn. Res. 2021 |
Machine learning › Probabilistic and Bayesian machine learning
dynamical system |
0.4 | 1 | 2020 | Krylov Subspace Method for Nonlinear Dynamical Systems with Random Noise · J. Mach. Learn. Res. 2020 |
Machine learning › Generative modeling
normalizing flow |
0.4 | 1 | 2020 | Coupling-based Invertible Neural Networks Are Universal Diffeomorphism Approximators · NeurIPS 2020 |
Mathematical optimization › iterative methods
krylov subspace methods |
0.4 | 1 | 2020 | Krylov Subspace Method for Nonlinear Dynamical Systems with Random Noise · J. Mach. Learn. Res. 2020 |
Algorithms and data structures
numerical linear algebra |
0.4 | 1 | 2020 | Krylov Subspace Method for Nonlinear Dynamical Systems with Random Noise · J. Mach. Learn. Res. 2020 |
Machine learning › Representation and self-supervised learning
group representations |
0.3 | 1 | 2025 | Deep Ridgelet Transform and Unified Universality Theorem for Deep and Shallow Joint-Group-Equivariant Machines · ICML 2025 |
Computer vision › 3D vision › geometric deep learning
hyperbolic neural networks |
0.2 | 1 | 2022 | Fully-Connected Network on Noncompact Symmetric Space and Ridgelet Transform based on Helgason-Fourier Analysis · ICML 2022 |
Machine learning › Time series and sequential data
anomaly detection |
0.1 | 1 | 2020 | Krylov Subspace Method for Nonlinear Dynamical Systems with Random Noise · J. Mach. Learn. Res. 2020 |
Mathematical optimization
approximation theory |
0.1 | 1 | 2020 | Coupling-based Invertible Neural Networks Are Universal Diffeomorphism Approximators · NeurIPS 2020 |
Methods — techniques the papers use, named apart from their topics
ridgelet transform · 1.4group representation theory · 0.9operator-theoretic analysis · 0.8koopman operator · 0.8differential geometry · 0.7wavelet function · 0.6helgason-fourier analysis · 0.6group representation · 0.6universality analysis · 0.5representer theorem · 0.5shift-invert arnoldi · 0.4maximum mean discrepancy · 0.4kernel mean embedding · 0.4diffeomorphism approximation · 0.4coupling flows · 0.4arnoldi method · 0.4affine coupling · 0.4reproducing kernel hilbert space · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Deep Ridgelet Transform and Unified Universality Theorem for Deep and Shallow Joint-Group-Equivariant MachinesabstractWe present a constructive universal approximation theorem for learning machines equipped with joint-group-equivariant feature maps, called the joint-equivariant machines, based on the group representation theory. ``Constructive'' here indicates that the distribution of parameters is given in a closed-form expression known as the ridgelet transform. Joint-group-equivariance encompasses a broad class of feature maps that generalize classical group-equivariance. Particularly, fully-connected networks are *not* group-equivariant *but* are joint-group-equivariant. Our main theorem also unifies the universal approximation theorems for both shallow and deep networks. Until this study, the universality of deep networks has been shown in a different manner from the universality of shallow networks, but our results discuss them on common ground. Now we can understand the approximation schemes of various learning machines in a unified manner. As applications, we show the constructive universal approximation properties of four examples: depth-$n$ joint-equivariant machine, depth-$n$ fully-connected network, depth-$n$ group-convolutional network, and a new depth-$2$ network with quadratic forms whose universality has not been known. Sho Sonoda, Yuka Hashimoto, Isao Ishikawa, Masahiro Ikeda |
ICML | 3 |
| 2024 | Koopman-based generalization bound: New aspect for full-rank weightsabstractWe propose a new bound for generalization of neural networks using Koopman operators. Whereas most of existing works focus on low-rank weight matrices, we focus on full-rank weight matrices. Our bound is tighter than existing norm-based bounds when the condition numbers of weight matrices are small. Especially, it is completely independent of the width of the network if the weight matrices are orthogonal. Our bound does not contradict to the existing bounds but is a complement to the existing bounds. As supported by several existing empirical results, low-rankness is not the only reason for generalization. Furthermore, our bound can be combined with the existing bounds to obtain a tighter bound. Our result sheds new light on understanding generalization of neural networks with full-rank weight matrices, and it provides a connection between operator-theoretic analysis and generalization of neural networks. Yuka Hashimoto, Sho Sonoda, Isao Ishikawa, Atsushi Nitanda, Taiji Suzuki |
ICLR | 3 |
| 2023 | Universal Approximation Property of Invertible Neural NetworksabstractInvertible neural networks (INNs) are neural network architectures with invertibility by design. Thanks to their invertibility and the tractability of their Jacobians, INNs have various machine learning applications such as probabilistic modeling, generative modeling, and representation learning. However, their attractive properties often come at the cost of restricting the layer design, which poses a question on their representation power: can we use these models to approximate sufficiently diverse functions? To answer this question, we have developed a general theoretical framework to investigate the representation power of INNs, building on a structure theorem of differential geometry. The framework simplifies the approximation problem of diffeomorphisms, which enables us to show the universal approximation properties of INNs. We apply the framework to two representative classes of INNs, namely Coupling-Flow-based INNs (CF-INNs) and Neural Ordinary Differential Equations (NODEs), and elucidate their high representation power despite the restrictions on their architectures. Isao Ishikawa, Takeshi Teshima, Koichi Tojo, Kenta Oono, Masahiro Ikeda, Masashi Sugiyama |
J. Mach. Learn. Res. | 1 |
| 2022 | Fully-Connected Network on Noncompact Symmetric Space and Ridgelet Transform based on Helgason-Fourier AnalysisabstractNeural network on Riemannian symmetric space such as hyperbolic space and the manifold of symmetric positive definite (SPD) matrices is an emerging subject of research in geometric deep learning. Based on the well-established framework of the Helgason-Fourier transform on the noncompact symmetric space, we present a fully-connected network and its associated ridgelet transform on the noncompact symmetric space, covering the hyperbolic neural network (HNN) and the SPDNet as special cases. The ridgelet transform is an analysis operator of a depth-2 continuous network spanned by neurons, namely, it maps an arbitrary given function to the weights of a network. Thanks to the coordinate-free reformulation, the role of nonlinear activation functions is revealed to be a wavelet function. Moreover, the reconstruction formula is applied to present a constructive proof of the universality of finite networks on symmetric spaces. Sho Sonoda, Isao Ishikawa, Masahiro Ikeda |
ICML | 2 |
| 2022 | Universality of Group Convolutional Neural Networks Based on Ridgelet Analysis on GroupsabstractWe show the universality of depth-2 group convolutional neural networks (GCNNs) in a unified and constructive manner based on the ridgelet theory. Despite widespread use in applications, the approximation property of (G)CNNs has not been well investigated. The universality of (G)CNNs has been shown since the late 2010s. Yet, our understanding on how (G)CNNs represent functions is incomplete because the past universality theorems have been shown in a case-by-case manner by manually/carefully assigning the network parameters depending on the variety of convolution layers, and in an indirect manner by converting/modifying the (G)CNNs into other universal approximators such as invariant polynomials and fully-connected networks. In this study, we formulate a versatile depth-2 continuous GCNN $S[\gamma]$ as a nonlinear mapping between group representations, and directly obtain an analysis operator, called the ridgelet trasform, that maps a given function $f$ to the network parameter $\gamma$ so that $S[\gamma]=f$. The proposed GCNN covers typical GCNNs such as the cyclic convolution on multi-channel images, networks on permutation-invariant inputs (Deep Sets), and $\mathrm{E}(n)$-equivariant networks. The closed-form expression of the ridgelet transform can describe how the network parameters are organized to represent a function. While it has been known only for fully-connected networks, this study is the first to obtain the ridgelet transform for GCNNs. By discretizing the closed-form expression, we can systematically generate a constructive proof of the $cc$-universality of finite GCNNs. In other words, our universality proofs are more unified and constructive than previous proofs. Sho Sonoda, Isao Ishikawa, Masahiro Ikeda |
NeurIPS | 2 |
| 2021 | Ridge Regression with Over-parametrized Two-Layer Networks Converge to Ridgelet SpectrumabstractCharacterization of local minima draws much attention in theoretical studies of deep learning. In this study, we investigate the distribution of parameters in an over-parametrized finite neural network trained by ridge regularized empirical square risk minimization (RERM). We develop a new theory of ridgelet transform, a wavelet-like integral transform that provides a powerful and general framework for the theoretical study of neural networks involving not only the ReLU but general activation functions. We show that the distribution of the parameters converges to a spectrum of the ridgelet transform. This result provides a new insight into the characterization of the local minima of neural networks, and the theoretical background of an inductive bias theory based on lazy regimes. We confirm the visual resemblance between the parameter distribution trained by SGD, and the ridgelet spectrum calculated by numerical integration through numerical experiments with finite models. Sho Sonoda, Isao Ishikawa, Masahiro Ikeda |
AISTATS | 2 |
| 2021 | Reproducing kernel Hilbert C*-module and kernel mean embeddingsabstractKernel methods have been among the most popular techniques in machine learning, where learning tasks are solved using the property of reproducing kernel Hilbert space (RKHS). In this paper, we propose a novel data analysis framework with reproducing kernel Hilbert $C^*$-module (RKHM) and kernel mean embedding (KME) in RKHM. Since RKHM contains richer information than RKHS or vector-valued RKHS (vvRKHS), analysis with RKHM enables us to capture and extract structural properties in such as functional data. We show a branch of theories for RKHM to apply to data analysis, including the representer theorem, and the injectivity and universality of the proposed KME. We also show RKHM generalizes RKHS and vvRKHS. Then, we provide concrete procedures for employing RKHM and the proposed KME to data analysis. Yuka Hashimoto, Isao Ishikawa, Masahiro Ikeda, Fuyuta Komura, Takeshi Katsura, Yoshinobu Kawahara |
J. Mach. Learn. Res. | 2 |
| 2020 | Coupling-based Invertible Neural Networks Are Universal Diffeomorphism ApproximatorsabstractInvertible neural networks based on coupling flows (CF-INNs) have various machine learning applications such as image synthesis and representation learning. However, their desirable characteristics such as analytic invertibility come at the cost of restricting the functional forms. This poses a question on their representation power: are CF-INNs universal approximators for invertible functions? Without a universality, there could be a well-behaved invertible transformation that the CF-INN can never approximate, hence it would render the model class unreliable. We answer this question by showing a convenient criterion: a CF-INN is universal if its layers contain affine coupling and invertible linear functions as special cases. As its corollary, we can affirmatively resolve a previously unsolved problem: whether normalizing flow models based on affine coupling can be universal distributional approximators. In the course of proving the universality, we prove a general theorem to show the equivalence of the universality for certain diffeomorphism classes, a theoretical insight that is of interest by itself. Takeshi Teshima, Isao Ishikawa, Koichi Tojo, Kenta Oono, Masahiro Ikeda, Masashi Sugiyama |
NeurIPS | 2 |
| 2020 | Krylov Subspace Method for Nonlinear Dynamical Systems with Random NoiseabstractOperator-theoretic analysis of nonlinear dynamical systems has attracted much attention in a variety of engineering and scientific fields, endowed with practical estimation methods using data such as dynamic mode decomposition. In this paper, we address a lifted representation of nonlinear dynamical systems with random noise based on transfer operators, and develop a novel Krylov subspace method for estimating the operators using finite data, with consideration of the unboundedness of operators. For this purpose, we first consider Perron-Frobenius operators with kernel-mean embeddings for such systems. We then extend the Arnoldi method, which is the most classical type of Kryov subspace methods, so that it can be applied to the current case. Meanwhile, the Arnoldi method requires the assumption that the operator is bounded, which is not necessarily satisfied for transfer operators on nonlinear systems. We accordingly develop the shift-invert Arnoldi method for Perron-Frobenius operators to avoid this problem. Also, we describe an approach of evaluating predictive accuracy by estimated operators on the basis of the maximum mean discrepancy, which is applicable, for example, to anomaly detection in complex systems. The empirical performance of our methods is investigated using synthetic and real-world healthcare data. Yuka Hashimoto, Isao Ishikawa, Masahiro Ikeda, Yoichi Matsuo, Yoshinobu Kawahara |
J. Mach. Learn. Res. | 2 |
| 2018 | Metric on Nonlinear Dynamical Systems with Perron-Frobenius OperatorsabstractThe development of a metric for structural data is a long-term problem in pattern recognition and machine learning. In this paper, we develop a general metric for comparing nonlinear dynamical systems that is defined with Perron-Frobenius operators in reproducing kernel Hilbert spaces. Our metric includes the existing fundamental metrics for dynamical systems, which are basically defined with principal angles between some appropriately-chosen subspaces, as its special cases. We also describe the estimation of our metric from finite data. We empirically illustrate our metric with an example of rotation dynamics in a unit disk in a complex plane, and evaluate the performance with real-world time-series data. Isao Ishikawa, Keisuke Fujii 0001, Masahiro Ikeda, Yuka Hashimoto, Yoshinobu Kawahara |
NeurIPS | 1 |