Orestis Plevrakis

dblp:236/5636 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 20% Optimization for machine learning · 20% Reinforcement learning · 19%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 40% Approximation and online algorithms · 40% Computational complexity · 20%

Topics — the 21 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Approximation and online algorithms
approximation algorithms
0.812024
On the Cut-Query Complexity of Approximating Max-Cut · ICALP 2024
Graph algorithms and graph theory
graph algorithms
0.812024
On the Cut-Query Complexity of Approximating Max-Cut · ICALP 2024
Graph algorithms and graph theory › graph cut
max-cut
0.812024
On the Cut-Query Complexity of Approximating Max-Cut · ICALP 2024
Approximation and online algorithms › approximation algorithms › approximation algorithms for graph problems
max-cut approximation
0.812024
On the Cut-Query Complexity of Approximating Max-Cut · ICALP 2024
Computational complexity
query complexity
0.812024
On the Cut-Query Complexity of Approximating Max-Cut · ICALP 2024
Machine learning › Trustworthy machine learning › learning with incomplete data
censored data
0.512021
Learning from Censored and Dependent Data: The case of Linear Dynamics · COLT 2021
Machine learning › Time series and sequential data
linear dynamical systems
0.512021
Learning from Censored and Dependent Data: The case of Linear Dynamics · COLT 2021
Machine learning › Optimization for machine learning › online optimization
online newton method
0.512021
Learning from Censored and Dependent Data: The case of Linear Dynamics · COLT 2021
Robotics › Motion planning and robot control
system identification
0.512021
Learning from Censored and Dependent Data: The case of Linear Dynamics · COLT 2021
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412020
Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of Dimensionality · NeurIPS 2020
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.412020
Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of Dimensionality · NeurIPS 2020
Machine learning › Optimization for machine learning › online optimization
bandit convex optimization
0.412020
Geometric Exploration for Online Control · NeurIPS 2020
Machine learning › Optimization for machine learning
convergence analysis
0.412020
Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of Dimensionality · NeurIPS 2020
Machine learning › Reinforcement learning
online control
0.412020
Geometric Exploration for Online Control · NeurIPS 2020
Machine learning › Deep learning architectures and training
overparameterized neural network
0.412020
Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of Dimensionality · NeurIPS 2020
Machine learning › Reinforcement learning
policy optimization
0.412020
Geometric Exploration for Online Control · NeurIPS 2020
Machine learning › Reinforcement learning
regret minimization
0.412020
Geometric Exploration for Online Control · NeurIPS 2020
Machine learning › Representation and self-supervised learning
contrastive learning
0.412019
A Theoretical Analysis of Contrastive Unsupervised Representation Learning · ICML 2019
Machine learning › Learning theory
generalization bounds
0.412019
A Theoretical Analysis of Contrastive Unsupervised Representation Learning · ICML 2019
Machine learning › Representation and self-supervised learning › contrastive learning
theoretical analysis of contrastive learning
0.412019
A Theoretical Analysis of Contrastive Unsupervised Representation Learning · ICML 2019
Machine learning › Deep learning architectures and training
ReLU networks
0.112020
Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of Dimensionality · NeurIPS 2020

Methods — techniques the papers use, named apart from their topics

sparsifier · 0.8cut dimension · 0.8stochastic gradient · 0.5online newton step · 0.5error bounds · 0.5step function approximation · 0.4online learning tools · 0.4min-max optimization · 0.4geometric exploration · 0.4barycentric spanner · 0.4latent class analysis · 0.4contrastive learning · 0.4
YearPublicationVenuePosition
2024 On the Cut-Query Complexity of Approximating Max-Cut
abstract
We consider the problem of query-efficient global max-cut on a weighted undirected graph in the value oracle model examined by [RSW18]. Graph algorithms in this cut query model and other query models have recently been studied for various other problems such as min-cut, connectivity, bipartiteness, and triangle detection. Max-cut in the cut query model can also be viewed as a natural special case of submodular function maximization: on query $S \subseteq V$, the oracle returns the total weight of the cut between $S$ and $V \backslash S$. Our first main technical result is a lower bound stating that a deterministic algorithm achieving a $c$-approximation for any $c > 1/2$ requires $Ω(n)$ queries. This uses an extension of the cut dimension to rule out approximation (prior work of [GPRW20] introducing the cut dimension only rules out exact solutions). Secondly, we provide a randomized algorithm with $\tilde{O}(n)$ queries that finds a $c$-approximation for any $c < 1$. We achieve this using a query-efficient sparsifier for undirected weighted graphs (prior work of [RSW18] holds only for unweighted graphs). To complement these results, for most constants $c \in (0,1]$, we nail down the query complexity of achieving a $c$-approximation, for both deterministic and randomized algorithms (up to logarithmic factors). Analogously to general submodular function maximization in the same model, we observe a phase transition at $c = 1/2$: we design a deterministic algorithm for global $c$-approximate max-cut in $O(\log n)$ queries for any $c < 1/2$, and show that any randomized algorithm requires $Ω(n/\log n)$ queries to find a $c$-approximate max-cut for any $c > 1/2$. Additionally, we show that any deterministic algorithm requires $Ω(n^2)$ queries to find an exact max-cut (enough to learn the entire graph).
Orestis Plevrakis, Seyoon Ragavan, S. Matthew Weinberg
ICALP1
2021 Learning from Censored and Dependent Data: The case of Linear Dynamics
abstract
Observations from dynamical systems often exhibit irregularities, such as censoring, where values are recorded only if they fall within a certain range. Censoring is ubiquitous in practice, due to saturating sensors, limit-of-detection effects, image frame effects, and combined with the temporal dependencies within the data, makes the task of system identification particularly challenging. In light of recent developments on learning linear dynamical systems (LDSs), and on censored statistics with independent data, we revisit the decades-old problem of learning an LDS, from censored observations (Lee and Maddala (1985), Zeger and Brookmeyer (1986)). Here, the learner observes the state x_t \in R^d if and only if x_t belongs to some set S_t _x0012_\in R^d. We develop the first computationally and statistically efficient algorithm for learning the system, assuming only oracle to the sets St. Our algorithm, Stochastic Online Newton with Switching Gradients, is a novel second-order method that builds on the Online Newton Step (ONS) of Hazan et al. (2007). Our Switching-Gradient scheme does not always use (stochastic) gradients of the function we want to optimize, which we call censor-aware function. Instead, in each iteration, it performs a simple test to decide whether to use the censor-aware, or another censor-oblivious function, for getting a stochastic gradient. In our analysis, we consider a “generic” Online Newton method, which uses arbitrary vectors instead of gradients, and we prove an error-bound for it. This can be used to appropriately design these vectors, Leading to our Switching-Gradient scheme. This framework significantly deviates from the recent long line of works on censored statistics (e.g, Daskalakis et al. (2018); Kontonis et al. (2019); Daskalakis et al. (2019), which apply Stochastic Gradient Descent (SGD), and their analysis reduces to establishing conditions for off-the-shelf SGD-bounds. Our approach enables to relax these conditions, and gives rise to phenomena that might appear counterintuitive, given the previous works. Specifically, our method makes progress even when the current “survival probability” is exponentially small. We believe that our analysis framework will have applications in more settings where the data are subject to censoring.
Orestis Plevrakis
COLT1
2020 Geometric Exploration for Online Control
abstract
We study the control of an \emph{unknown} linear dynamical system under general convex costs. The objective is minimizing regret vs the class of strongly-stable linear policies. In this work, we first consider the case of known cost functions, for which we design the first polynomial-time algorithm with $n^3\sqrt{T}$-regret, where $n$ is the dimension of the state plus the dimension of control input. The $\sqrt{T}$-horizon dependence is optimal, and improves upon the previous best known bound of $T^{2/3}$. The main component of our algorithm is a novel geometric exploration strategy: we adaptively construct a sequence of barycentric spanners in an over-parameterized policy space. Second, we consider the case of bandit feedback, for which we give the first polynomial-time algorithm with $poly(n)\sqrt{T}$-regret, building on Stochastic Bandit Convex Optimization.
Orestis Plevrakis, Elad Hazan
NeurIPS1
2020 Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of Dimensionality
abstract
Adversarial training is a popular method to give neural nets robustness against adversarial perturbations. In practice adversarial training leads to low robust training loss. However, a rigorous explanation for why this happens under natural conditions is still missing. Recently a convergence theory of standard (non-adversarial) supervised training was developed by various groups for {\em very overparametrized} nets. It is unclear how to extend these results to adversarial training because of the min-max objective. Recently, a first step towards this direction was made by Gao et al. using tools from online learning, but they require the width of the net to be \emph{exponential} in input dimension $d$, and with an unnatural activation function. Our work proves convergence to low robust training loss for \emph{polynomial} width instead of exponential, under natural assumptions and with ReLU activations. A key element of our proof is showing that ReLU networks near initialization can approximate the step function, which may be of independent interest.
Yi Zhang 0074, Orestis Plevrakis, Simon S. Du, Xingguo Li, Zhao Song 0002, Sanjeev Arora
NeurIPS2
2019 A Theoretical Analysis of Contrastive Unsupervised Representation Learning
abstract
Recent empirical works have successfully used unlabeled data to learn feature representations that are broadly useful in downstream classification tasks. Several of these methods are reminiscent of the well-known word2vec embedding algorithm: leveraging availability of pairs of semantically “similar" data points and “negative samples," the learner forces the inner product of representations of similar pairs with each other to be higher on average than with negative samples. The current paper uses the term contrastive learning for such algorithms and presents a theoretical framework for analyzing them by introducing latent classes and hypothesizing that semantically similar points are sampled from the same latent class. This framework allows us to show provable guarantees on the performance of the learned representations on the average classification task that is comprised of a subset of the same set of latent classes. Our generalization bound also shows that learned representations can reduce (labeled) sample complexity on downstream tasks. We conduct controlled experiments in both the text and image domains to support the theory.
Nikunj Saunshi, Orestis Plevrakis, Sanjeev Arora, Mikhail Khodak, Hrishikesh Khandeparkar
ICML2