Poorya Mianjy

dblp:182/8944 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Deep learning architectures and training · 35% Trustworthy machine learning · 26% Learning theory · 16%
Theoretical computer science
5 papers
Algorithms and data structures · 66% Mathematical optimization · 34%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › regularization
dropout
1.642021
Dropout: Explicit Forms and Capacity Control · ICML 2021
On Convergence and Generalization of Dropout Training · NeurIPS 2020
On Dropout and Nuclear Norm Regularization · ICML 2019
Machine learning › Deep learning architectures and training
regularization
1.542021
Dropout: Explicit Forms and Capacity Control · ICML 2021
On Dropout and Nuclear Norm Regularization · ICML 2019
On the Implicit Bias of Dropout · ICML 2018
Machine learning › Learning theory
generalization bounds
1.432021
Robust Learning for Data Poisoning Attacks · ICML 2021
Dropout: Explicit Forms and Capacity Control · ICML 2021
On Convergence and Generalization of Dropout Training · NeurIPS 2020
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
1.222023
Robustness Guarantees for Adversarially Trained Neural Networks · NeurIPS 2023
Adversarial Robustness is at Odds with Lazy Training · NeurIPS 2022
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.712023
Robustness Guarantees for Adversarially Trained Neural Networks · NeurIPS 2023
Machine learning › Optimization for machine learning
bilevel optimization
0.712023
Robustness Guarantees for Adversarially Trained Neural Networks · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
certified robustness
0.712023
Robustness Guarantees for Adversarially Trained Neural Networks · NeurIPS 2023
Algorithms and data structures › numerical linear algebra
dimensionality reduction
0.622018
Streaming Kernel PCA with \tilde{O}(\sqrt{n}) Random Features · NeurIPS 2018
Stochastic Approximation for Canonical Correlation Analysis · NIPS 2017
Machine learning › Trustworthy machine learning › robustness
adversarial examples
0.612022
Adversarial Robustness is at Odds with Lazy Training · NeurIPS 2022
Machine learning › Deep learning architectures and training › training dynamics
lazy training regime
0.612022
Adversarial Robustness is at Odds with Lazy Training · NeurIPS 2022
Machine learning › Deep learning architectures and training
overparameterized neural network
0.612022
Adversarial Robustness is at Odds with Lazy Training · NeurIPS 2022
Mathematical optimization › stochastic optimization
stochastic approximation
0.522017
Stochastic Approximation for Canonical Correlation Analysis · NIPS 2017
Stochastic Optimization for Multiview Representation Learning using Partial Least Squares · ICML 2016
Mathematical optimization
stochastic optimization
0.522017
Stochastic Approximation for Canonical Correlation Analysis · NIPS 2017
Stochastic Optimization for Multiview Representation Learning using Partial Least Squares · ICML 2016
Machine learning › Learning theory › statistical learning theory
capacity control
0.512021
Dropout: Explicit Forms and Capacity Control · ICML 2021
Machine learning › Trustworthy machine learning › robustness
poisoning attack defense
0.512021
Robust Learning for Data Poisoning Attacks · ICML 2021
Machine learning › Trustworthy machine learning
robustness
0.512021
Robust Learning for Data Poisoning Attacks · ICML 2021
Machine learning › Optimization for machine learning
convergence analysis
0.412020
On Convergence and Generalization of Dropout Training · NeurIPS 2020
Machine learning › Deep learning architectures and training › regularization
dropout training
0.412020
On Convergence and Generalization of Dropout Training · NeurIPS 2020
Algorithms and data structures › numerical linear algebra › dimensionality reduction
principal component analysis
0.422018
Streaming Kernel PCA with \tilde{O}(\sqrt{n}) Random Features · NeurIPS 2018
Stochastic PCA with 𝓁2 and 𝓁1 Regularization · ICML 2018
Machine learning › Optimization for machine learning
implicit regularization
0.412019
On Dropout and Nuclear Norm Regularization · ICML 2019
Machine learning › Deep learning architectures and training › regularization › spectral regularization
nuclear norm regularization
0.412019
On Dropout and Nuclear Norm Regularization · ICML 2019
Machine learning › Learning theory
implicit bias
0.312018
On the Implicit Bias of Dropout · ICML 2018
Machine learning › Deep learning architectures and training
neural network expressivity
0.312018
Understanding Deep Neural Networks with Rectified Linear Units · ICLR (Poster) 2018
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
principal component analysis
0.312018
Streaming Principal Component Analysis in Noisy Settings · ICML 2018
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random features
0.312018
Streaming Kernel PCA with \tilde{O}(\sqrt{n}) Random Features · NeurIPS 2018
Machine learning › Optimization for machine learning
stochastic optimization
0.312018
Stochastic PCA with 𝓁2 and 𝓁1 Regularization · ICML 2018
Machine learning › Time series and sequential data
streaming data
0.312018
Streaming Kernel PCA with \tilde{O}(\sqrt{n}) Random Features · NeurIPS 2018
Algorithms and data structures › data streams
streaming algorithms
0.312018
Streaming Principal Component Analysis in Noisy Settings · ICML 2018
Algorithms and data structures › data streams › streaming algorithms
streaming PCA
0.312018
Streaming Principal Component Analysis in Noisy Settings · ICML 2018
Algorithms and data structures › numerical linear algebra › dimensionality reduction
canonical correlation analysis
0.312017
Stochastic Approximation for Canonical Correlation Analysis · NIPS 2017

Methods — techniques the papers use, named apart from their topics

trace norm regularization · 1.0stochastic gradient descent · 1.0rademacher complexity · 1.0surrogate loss · 0.7projected gradient descent · 0.7lazy training analysis · 0.6gradient ascent · 0.6RKHS analysis · 0.5overparametrization theory · 0.4neural tangent kernel · 0.4streaming PCA · 0.3stochastic optimization · 0.3regularization · 0.3random fourier features · 0.3perturbation analysis · 0.3oja's algorithm · 0.3matrix stochastic gradient · 0.3matrix exponentiated gradient · 0.3
YearPublicationVenuePosition
2023 Robustness Guarantees for Adversarially Trained Neural Networks
abstract
We study robust adversarial training of two-layer neural networks as a bi-level optimization problem. In particular, for the inner loop that implements the adversarial attack during training using projected gradient descent (PGD), we propose maximizing a \emph{lower bound} on the $0/1$-loss by reflecting a surrogate loss about the origin. This allows us to give a convergence guarantee for the inner-loop PGD attack. Furthermore, assuming the data is linearly separable, we provide precise iteration complexity results for end-to-end adversarial training, which holds for any width and initialization. We provide empirical evidence to support our theoretical results.
Poorya Mianjy, Raman Arora
NeurIPS1
2022 Adversarial Robustness is at Odds with Lazy Training
abstract
Recent works show that adversarial examples exist for random neural networks [Daniely and Schacham, 2020] and that these examples can be found using a single step of gradient ascent [Bubeck et al., 2021]. In this work, we extend this line of work to ``lazy training'' of neural networks -- a dominant model in deep learning theory in which neural networks are provably efficiently learnable. We show that over-parametrized neural networks that are guaranteed to generalize well and enjoy strong computational guarantees remain vulnerable to attacks generated using a single step of gradient ascent.
Yunjuan Wang, Enayat Ullah, Poorya Mianjy, Raman Arora
NeurIPS3
2021 Dropout: Explicit Forms and Capacity Control
abstract
We investigate the capacity control provided by dropout in various machine learning problems. First, we study dropout for matrix completion, where it induces a distribution-dependent regularizer that equals the weighted trace-norm of the product of the factors. In deep learning, we show that the distribution-dependent regularizer due to dropout directly controls the Rademacher complexity of the underlying class of deep neural networks. These developments enable us to give concrete generalization error bounds for the dropout algorithm in both matrix completion as well as training deep neural networks.
Raman Arora, Peter L. Bartlett, Poorya Mianjy, Nathan Srebro
ICML3
2021 Robust Learning for Data Poisoning Attacks
abstract
We investigate the robustness of stochastic approximation approaches against data poisoning attacks. We focus on two-layer neural networks with ReLU activation and show that under a specific notion of separability in the RKHS induced by the infinite-width network, training (finite-width) networks with stochastic gradient descent is robust against data poisoning attacks. Interestingly, we find that in addition to a lower bound on the width of the network, which is standard in the literature, we also require a distribution-dependent upper bound on the width for robust generalization. We provide extensive empirical evaluations that support and validate our theoretical results.
Yunjuan Wang, Poorya Mianjy, Raman Arora
ICML2
2020 On Convergence and Generalization of Dropout Training
abstract
We study dropout in two-layer neural networks with rectified linear unit (ReLU) activations. Under mild overparametrization and assuming that the limiting kernel can separate the data distribution with a positive margin, we show that the dropout training with logistic loss achieves $\epsilon$-suboptimality in the test error in $O(1/\epsilon)$ iterations.
Poorya Mianjy, Raman Arora
NeurIPS1
2019 On Dropout and Nuclear Norm Regularization
abstract
We give a formal and complete characterization of the explicit regularizer induced by dropout in deep linear networks with squared loss. We show that (a) the explicit regularizer is composed of an $\ell_2$-path regularizer and other terms that are also re-scaling invariant, (b) the convex envelope of the induced regularizer is the squared nuclear norm of the network map, and (c) for a sufficiently large dropout rate, we characterize the global optima of the dropout objective. We validate our theoretical findings with empirical results.
Poorya Mianjy, Raman Arora
ICML1
2018 Understanding Deep Neural Networks with Rectified Linear Units
Raman Arora, Amitabh Basu, Poorya Mianjy, Anirbit Mukherjee
ICLR (Poster)3
2018 Streaming Principal Component Analysis in Noisy Settings
Teodor V. Marinov, Poorya Mianjy, Raman Arora
ICML2
2018 Stochastic PCA with 𝓁2 and 𝓁1 Regularization
Poorya Mianjy, Raman Arora
ICML1
2018 On the Implicit Bias of Dropout
abstract
Algorithmic approaches endow deep learning systems with implicit bias that helps them generalize even in over-parametrized settings. In this paper, we focus on understanding such a bias induced in learning through dropout, a popular technique to avoid overfitting in deep learning. For single hidden-layer linear neural networks, we show that dropout tends to make the norm of incoming/outgoing weight vectors of all the hidden nodes equal. In addition, we provide a complete characterization of the optimization landscape induced by dropout.
Poorya Mianjy, Raman Arora, René Vidal
ICML1
2018 Streaming Kernel PCA with \tilde{O}(\sqrt{n}) Random Features
abstract
We study the statistical and computational aspects of kernel principal component analysis using random Fourier features and show that under mild assumptions, $O(\sqrt{n} \log n)$ features suffices to achieve $O(1/\epsilon^2)$ sample complexity. Furthermore, we give a memory efficient streaming algorithm based on classical Oja's algorithm that achieves this rate
Enayat Ullah, Poorya Mianjy, Teodor V. Marinov, Raman Arora
NeurIPS2
2017 Stochastic Approximation for Canonical Correlation Analysis
abstract
We propose novel first-order stochastic approximation algorithms for canonical correlation analysis (CCA). Algorithms presented are instances of inexact matrix stochastic gradient (MSG) and inexact matrix exponentiated gradient (MEG), and achieve $\epsilon$-suboptimality in the population objective in $\operatorname{poly}(\frac{1}{\epsilon})$ iterations. We also consider practical variants of the proposed algorithms and compare them with other methods for CCA both theoretically and empirically.
Raman Arora, Teodor V. Marinov, Poorya Mianjy, Nathan Srebro
NIPS3
2016 Stochastic Optimization for Multiview Representation Learning using Partial Least Squares
abstract
Partial Least Squares (PLS) is a ubiquitous statistical technique for bilinear factor analysis. It is used in many data analysis, machine learning, and information retrieval applications to model the covariance structure between a pair of data matrices. In this paper, we consider PLS for representation learning in a multiview setting where we have more than one view in data at training time. Furthermore, instead of framing PLS as a problem about a fixed given data set, we argue that PLS should be studied as a stochastic optimization problem, especially in a "big data" setting, with the goal of optimizing a population objective based on sample. This view suggests using Stochastic Approximation (SA) approaches, such as Stochastic Gradient Descent (SGD) and enables a rigorous analysis of their benefits. In this paper, we develop SA approaches to PLS and provide iteration complexity bounds for the proposed algorithms.
Raman Arora, Poorya Mianjy, Teodor V. Marinov
ICML2