Michael Puthawala

dblp:267/5586 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Learning theory · 34% Generative modeling · 31% Deep learning architectures and training · 27%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
neural operator
1.422024
Can neural operators always be continuously discretized? · NeurIPS 2024
Globally injective and bijective neural operators · NeurIPS 2023
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation
1.222023
Globally injective and bijective neural operators · NeurIPS 2023
Universal Joint Approximation of Manifolds and Densities by Simple Injective Flows · ICML 2022
Machine learning › Learning theory
approximation theory
0.812024
Can neural operators always be continuously discretized? · NeurIPS 2024
Machine learning › Generative modeling
generative prior
0.612022
Globally Injective ReLU Networks · J. Mach. Learn. Res. 2022
Machine learning › Generative modeling › normalizing flow
injective flow
0.612022
Universal Joint Approximation of Manifolds and Densities by Simple Injective Flows · ICML 2022
Machine learning › Generative modeling
inverse problem
0.612022
Globally Injective ReLU Networks · J. Mach. Learn. Res. 2022
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.612022
Universal Joint Approximation of Manifolds and Densities by Simple Injective Flows · ICML 2022
Machine learning › Generative modeling
normalizing flow
0.612022
Universal Joint Approximation of Manifolds and Densities by Simple Injective Flows · ICML 2022
Machine learning › Deep learning architectures and training
ReLU networks
0.612022
Globally Injective ReLU Networks · J. Mach. Learn. Res. 2022
Machine learning › Learning theory
well-posedness
0.612022
Globally Injective ReLU Networks · J. Mach. Learn. Res. 2022

Methods — techniques the papers use, named apart from their topics

fredholm theory · 1.4category theory · 0.8leray-schauder degree theory · 0.7random projection · 0.6lipschitz constant analysis · 0.6differential topology · 0.6clean trick · 0.6algebraic topology · 0.6
YearPublicationVenuePosition
2024 Can neural operators always be continuously discretized?
abstract
In this work we consider the problem of discretization of neural operators in a general setting. Using category theory, we give a no-go theorem that shows that diffeomorphisms between Hilbert spaces may not admit any continuous approximations by diffeomorphisms on finite spaces, even if the discretization is non-linear. This shows how infinite-dimensional Hilbert spaces and finite-dimensional vector spaces fundamentally differ. A key take-away is that to obtain discretization invariance, considerable effort is needed to ensure that finite-dimensional approximations of neural operator converge not only as sequences of functions, but that their representations converge in a suitable sense as well. With this perspective, we give several positive results. We first show that strongly monotone diffeomorphism operators always admit finite-dimensional strongly monotone diffeomorphisms. Next we show that bilipschitz neural operators may always be written via the repeated alternating composition of strongly monotone neural operators and invertible linear maps. We also show that such operators may be inverted locally via iteration provided that such inverse exists. Finally, we conclude by showing how our framework may be used `out of the box' to prove quantitative approximation results for discretization of neural operators.
Takashi Furuya, Michael Puthawala, Matti Lassas, Maarten V. de Hoop
NeurIPS2
2023 Globally injective and bijective neural operators
abstract
Recently there has been great interest in operator learning, where networks learn operators between function spaces from an essentially infinite-dimensional perspective. In this work we present results for when the operators learned by these networks are injective and surjective. As a warmup, we combine prior work in both the finite-dimensional ReLU and operator learning setting by giving sharp conditions under which ReLU layers with linear neural operators are injective. We then consider the case when the activation function is pointwise bijective and obtain sufficient conditions for the layer to be injective. We remark that this question, while trivial in the finite-rank setting, is subtler in the infinite-rank setting and is proven using tools from Fredholm theory. Next, we prove that our supplied injective neural operators are universal approximators and that their implementation, with finite-rank neural networks, are still injective. This ensures that injectivity is not 'lost' in the transcription from analytical operators to their finite-rank implementation with networks. Finally, we conclude with an increase in abstraction and consider general conditions when subnetworks, which may have many layers, are injective and surjective and provide an exact inversion from a 'linearization.’ This section uses general arguments from Fredholm theory and Leray-Schauder degree theory for non-linear integral equations to analyze the mapping properties of neural operators in function spaces. These results apply to subnetworks formed from the layers considered in this work, under natural conditions. We believe that our work has applications in Bayesian uncertainty quantification where injectivity enables likelihood estimation and in inverse problems where surjectivity and injectivity corresponds to existence and uniqueness of the solutions, respectively.
Takashi Furuya, Michael Puthawala, Matti Lassas, Maarten V. de Hoop
NeurIPS2
2022 Universal Joint Approximation of Manifolds and Densities by Simple Injective Flows
abstract
We study approximation of probability measures supported on n-dimensional manifolds embedded in R^m by injective flows—neural networks composed of invertible flows and injective layers. We show that in general, injective flows between R^n and R^m universally approximate measures supported on images of extendable embeddings, which are a subset of standard embeddings: when the embedding dimension m is small, topological obstructions may preclude certain manifolds as admissible targets. When the embedding dimension is sufficiently large, m >= 3n+1, we use an argument from algebraic topology known as the clean trick to prove that the topological obstructions vanish and injective flows universally approximate any differentiable embedding. Along the way we show that the studied injective flows admit efficient projections on the range, and that their optimality can be established "in reverse," resolving a conjecture made in Brehmer & Cranmer 2020.
Michael Puthawala, Matti Lassas, Ivan Dokmanic, Maarten V. de Hoop
ICML1
2022 Globally Injective ReLU Networks
abstract
Injectivity plays an important role in generative models where it enables inference; in inverse problems and compressed sensing with generative priors it is a precursor to well posedness. We establish sharp characterizations of injectivity of fully-connected and convolutional ReLU layers and networks. First, through a layerwise analysis, we show that an expansivity factor of two is necessary and sufficient for injectivity by constructing appropriate weight matrices. We show that global injectivity with iid Gaussian matrices, a commonly used tractable model, requires larger expansivity between 3.4 and 10.5. We also characterize the stability of inverting an injective network via worst-case Lipschitz constants of the inverse. We then use arguments from differential topology to study injectivity of deep networks and prove that any Lipschitz map can be approximated by an injective ReLU network. Finally, using an argument based on random projections, we show that an end-to-end---rather than layerwise---doubling of the dimension suffices for injectivity. Our results establish a theoretical basis for the study of nonlinear inverse and inference problems using neural networks.
Michael Puthawala, Konik Kothari, Matti Lassas, Ivan Dokmanic, Maarten V. de Hoop
J. Mach. Learn. Res.1