VLDB 2026 Research / reviewers in the wild / expert
Payam Delgosha
dblp:36/11107
· DBLP profile ↗
18ranked-venue papers
15as first author
8since 2021 · last 2025
0000-0003-0266-9237ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 7 first-author · 3 since 2021Theory of computation · 7 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CurricuLLM: Automatic Task Curricula Design for Learning Complex Robot Skills Using Large Language ModelsabstractCurriculum learning is a training mechanism in reinforcement learning (RL) that facilitates the achievement of complex policies by progressively increasing the task difficulty during training. However, designing effective curricula for a specific task often requires extensive domain knowledge and human intervention, which limits its applicability across various domains. Our core idea is that large language models (LLMs), with their extensive training on diverse language data and ability to encapsulate world knowledge, present significant potential for efficiently breaking down tasks and decomposing skills across various robotics environments. Additionally, the demonstrated success of LLMs in translating natural language into executable code for RL agents strengthens their role in generating task curricula. In this work, we propose CurricuLLM, which leverages the high-level planning and programming capabilities of LLMs for curriculum design, thereby enhancing the efficient learning of complex target tasks. CurricuLLM consists of: (Step 1) Generating a sequence of subtasks that aid target task learning in natural language form, (Step 2) Translating natural language description of subtasks in executable task code, including the reward code and goal distribution code, and (Step 3) Evaluating trained policies based on trajectory rollout and subtask description. We evaluate Cur-ricuLLM in various robotics simulation environments, ranging from manipulation, navigation, and locomotion, to show that CurricuLLM can aid learning complex robot control tasks. In addition, we validate humanoid locomotion policy learned through CurricuLLM in the real-world. Project website is https://iconlab.negarmehr.com/CurricuLLM/ Kanghyun Ryu, Qiayuan Liao, Zhongyu Li 0003, Payam Delgosha, Koushil Sreenath, Negar Mehr |
ICRA | 4 |
| 2024 | Binary Classification Under ℓ0 Attacks for General Noise DistributionabstractAdversarial examples have recently drawn considerable attention in the field of machine learning due to the fact that small perturbations in the data can result in major performance degradation. This phenomenon is usually modeled by a malicious adversary that can apply perturbations to the data in a constrained fashion, such as being bounded in a certain norm. In this paper, we study this problem when the adversary is constrained by the$\ell _{0}$norm; i.e., it can perturb a certain number of coordinates in the input, but has no limit on how much it can perturb those coordinates. Due to the combinatorial nature of this setting, we need to go beyond the standard techniques in robust machine learning to address this problem. We consider a binary classification scenario where$d$noisy data samples of the true label are provided to us after adversarial perturbations. We introduce a classification method which employs a nonlinear component called truncation, and show in an asymptotic scenario, as long as the adversary is restricted to perturb no more than$\sqrt {d}$data samples, we can almost achieve the optimal classification error in the absence of the adversary, i.e., we can completely neutralize adversary’s effect. Surprisingly, we observe a phase transition in the sense that using a converse argument, we show that if the adversary can perturb more than$\sqrt {d}$coordinates, no classifier can do better than a random guess. Payam Delgosha, Seyed Hamed Hassani, Ramtin Pedarsani |
IEEE Trans. Inf. Theory | 1 |
| 2023 | Generalization Properties of Adversarial Training for -ℓ0 Bounded Adversarial AttacksabstractWe have widely observed that neural networks are vulnerable to small additive perturbations to the input causing misclassification. In this paper, we focus on the ℓ0-bounded adversarial attacks, and aim to theoretically characterize the performance of adversarial training for an important class of truncated classifiers. Such classifiers are shown to have strong performance empirically, as well as theoretically in the Gaussian mixture model, in the ℓ0-adversarial setting. The main contribution of this paper is to prove a novel generalization bound for the binary classification setting with ℓ0-bounded adversarial perturbation that is distribution-independent. Deriving a generalization bound in this setting has two main challenges: (i) the truncated inner product which is highly non-linear; and (ii) maximization over the ℓ0ball due to adversarial training is non-convex and highly non-smooth. To tackle these challenges, we develop new coding techniques for bounding the combinatorial dimension of the truncated hypothesis class. Payam Delgosha, Seyed Hamed Hassani, Ramtin Pedarsani |
ITW | 1 |
| 2023 | A Universal Lossless Compression Method Applicable to Sparse Graphs and Heavy-Tailed Sparse GraphsabstractGraphical data arises naturally in several modern applications, including but not limited to internet graphs, social networks, genomics and proteomics. The typically large size of graphical data argues for the importance of designing universal compression methods for such data. In most applications, the graphical data is sparse, meaning that the number of edges in the graph scales more slowly than$n^{2}$, where$n$denotes the number of vertices. Although in some applications the number of edges scales linearly with$n$, in others the number of edges is much smaller than$n^{2}$but appears to scale superlinearly with$n$. We call the former sparse graphs and the latter heavy-tailed sparse graphs. In this paper we introduce a universal lossless compression method which is simultaneously applicable to both classes. We do this by employing the local weak convergence framework for sparse graphs and the sparse graphon framework for heavy-tailed sparse graphs. Payam Delgosha, Venkat Anantharam |
IEEE Trans. Inf. Theory | 1 |
| 2022 | Efficient and Robust Classification for Sparse AttacksabstractIn the past two decades we have seen the popularity of neural networks increase in conjunction with their classification accuracy. Parallel to this, we have also witnessed how fragile the very same prediction models are: tiny perturbations to the inputs can cause misclassification errors throughout entire datasets. In this paper, we consider perturbations bounded by the ℓ0–norm, which have been shown as effective attacks in the domains of image-recognition, natural language processing, and malware-detection. To this end, we propose a novel defense method that consists of "truncation" and "adversarial training". We then theoretically study the Gaussian mixture setting and prove the asymptotic optimality of our proposed classifier. Motivated by the insights we obtain, we extend these components to neural network classifiers. We conduct numerical experiments in the domain of computer vision using the MNIST and CIFAR datasets, demonstrating significant improvement for the robust classification error of neural networks. Mark Beliaev, Payam Delgosha, Seyed Hamed Hassani, Ramtin Pedarsani |
ISIT | 2 |
| 2022 | Binary Classification Under ℓ0 Attacks for General Noise DistributionabstractAdversarial examples have recently drawn considerable attention in the field of machine learning due to the fact that small perturbations in the data can result in major performance degradation. This phenomenon is usually modeled by a malicious adversary that can apply perturbations to the data in a constrained fashion, such as being bounded in a certain norm. In this paper, we study this problem when the adversary is constrained by the ℓ0norm; i.e., it can perturb a certain number of coordinates in the input, but has no limit on how much it can perturb those coordinates. Due to the combinatorial nature of this setting, we need to go beyond the standard techniques in robust machine learning to address this problem. We consider a binary classification scenario where d noisy data samples of the true label are provided to us after adversarial perturbations. We introduce a classification method which employs a nonlinear component called truncation, and show in an asymptotic scenario, as long as the adversary is restricted to perturb no more than $\sqrt d $ data samples, we can almost achieve the optimal classification error in the absence of the adversary, i.e. we can completely neutralize adversary’s effect. Surprisingly, we observe a phase transition in the sense that using a converse argument, we show that if the adversary can perturb more than $\sqrt d $ coordinates, no classifier can do better than a random guess. Payam Delgosha, Seyed Hamed Hassani, Ramtin Pedarsani |
ISIT | 1 |
| 2022 | Distributed Compression of Graphical Data
Payam Delgosha, Venkat Anantharam |
IEEE Trans. Inf. Theory | 1 |
| 2021 | A Universal Lossless Compression Method applicable to Sparse Graphs and heavy-tailed Sparse GraphsabstractGraphical data arises naturally in several modern applications, including but not limited to internet graphs, social networks, genomics and proteomics. The typically large size of graphical data argues for the importance of designing universal compression methods for such data. In most applications, the graphical data is sparse, meaning that the number of edges in the graph scales more slowly than$n^{2}$, where$n$denotes the number of vertices. Although in some applications the number of edges scales linearly with$n$, in others the number of edges is much smaller than$n^{2}$but appears to scale superlinearly with$n$. We call the former sparse graphs and the latter heavy-tailed sparse graphs. In this paper we introduce a universal lossless compression method which is simultaneously applicable to both classes. We do this by employing the local weak convergence framework for sparse graphs and the sparse graphon framework for heavy-tailed sparse graphs. Payam Delgosha, Venkat Anantharam |
ISIT | 1 |
| 2020 | A Universal Low Complexity Compression Algorithm for Sparse Marked GraphsabstractMany modern applications involve accessing and processing graphical data, i.e. data that is naturally indexed by graphs. Examples come from internet graphs, social networks, genomics and proteomics, and other sources. The typically large size of such data motivates seeking efficient ways for its compression and decompression. The current compression methods are usually tailored to specific models, or do not provide theoretical guarantees. In this paper, we introduce a low-complexity lossless compression algorithm for sparse marked graphs, i.e. graphical data indexed by sparse graphs, which is capable of universally achieving the optimal compression rate in a precisely defined sense. In order to define universality, we employ the framework of local weak convergence, which allows one to make sense of a notion of stochastic processes for graphs. Moreover, we investigate the performance of our algorithm through some experimental results on both synthetic and real-world data. Payam Delgosha, Venkat Anantharam |
ISIT | 1 |
| 2020 | Universal Lossless Compression of Graphical DataabstractGraphical data is comprised of a graph with marks on its edges and vertices. The mark indicates the value of some attribute associated to the respective edge or vertex. Examples of such data arise in social networks, molecular and systems biology, and web graphs, as well as in several other application areas. Our goal is to design schemes that can efficiently compress such graphical data without making assumptions about its stochastic properties. Namely, we wish to develop a universal compression algorithm for graphical data sources. To formalize this goal, we employ the framework of local weak convergence, also called the objective method, which provides a technique to think of a marked graph as a kind of stationary stochastic processes, stationary with respect to movement between vertices of the graph. In recent work, we have generalized a notion of entropy for unmarked graphs in this framework, due to Bordenave and Caputo, to the case of marked graphs. We use this notion to evaluate the efficiency of a compression scheme. The lossless compression scheme we propose in this paper is then proved to be universally optimal in a precise technical sense. It is also capable of performing local data queries in the compressed form. Payam Delgosha, Venkat Anantharam |
IEEE Trans. Inf. Theory | 1 |
| 2019 | Deep Switch Networks for Generating Discrete Data and LanguageabstractMultilayer switch networks are proposed as artificial generators of high-dimensional discrete data (e.g., binary vectors, categorical data, natural language, network log files, and discrete-valued time series). Unlike deconvolution networks which generate continuous-valued data and which consist of upsampling filters and reverse pooling layers, multilayer switch networks are composed of adaptive switches which model conditional distributions of discrete random variables. An interpretable, statistical framework is introduced for training these nonlinear networks based on a maximum-likelihood objective function. To learn network parameters, stochastic gradient descent is applied to the objective, and is stable until convergence. This direct optimization does not involve back-propagation over separate encoder and decoder networks, or adversarial training of dueling networks. While training remains tractable for moderately sized networks, Markov-chain Monte Carlo (MCMC) approximations of gradients are derived for deep networks which contain latent variables. The statistical framework is evaluated on synthetic data, high-dimensional binary data of handwritten digits, and web-crawled natural language data. Aspects of the model’s framework such as interpretability, computational complexity, and generalization ability are discussed. Payam Delgosha, Naveen Goela |
AISTATS | 1 |
| 2018 | Distributed Compression of Graphical DataabstractIn contrast to time series, graphical data is data indexed by the nodes and edges of a graph. Modern applications such as the internet, social networks, genomics and proteomics generate graphical data, often at large scale. The large scale argues for the need to compress such data for storage and subsequent processing. Since this data might have several components available in different locations, it is also important to study distributed compression of graphical data. In this paper, we derive a rate region for this problem which is a counterpart of the Slepian-Wolf Theorem. We characterize the rate region when the statistical description of the distributed graphical data is one of two types - a marked sparse Erdos-Renyi ensemble or a marked configuration model. Our results are in terms of a generalization of the notion of entropy introduced by Bordenave and Caputo in the study of local weak limits of sparse graphs. Payam Delgosha, Venkat Anantharam |
ISIT | 1 |
| 2017 | Universal lossless compression of graphical dataabstractConsider a data source comprised of a graph with marks on its edges and vertices. Examples of such data sources are social networks, biological data, web graphs, etc. Our goal is to design schemes that can efficiently compress and store such data. We aim for universal compression, i.e. without making assumptions about the stochastic properties of the data. To make sense of this, we employ the framework of local weak convergence, also called the objective method, which formalizes the notion of stationary stochastic processes indexed by graphs. We generalize a recently developed notion of entropy for such processes, due to Bordenave and Caputo, to the case of marked graphs, and argue that it is an appropriate way to evaluate the efficiency of a compression scheme. The lossless compression scheme we propose in this paper is then proved to be universally optimal. It is also capable of performing local data queries in the compressed form. Payam Delgosha, Venkat Anantharam |
ISIT | 1 |
| 2017 | High-Probability Guarantees in Repeated Games: Theory and Applications in Information Theory
Payam Delgosha, Amin Gohari, Mohammad Akbarpour |
Proc. IEEE | 1 |
| 2017 | Information Theoretic Cutting of a Cake
Payam Delgosha, Amin Gohari |
IEEE Trans. Inf. Theory | 1 |
| 2016 | High probability guarantees in repeated games: Theory and applications in information theoryabstractWe introduce a “high-probability” framework for repeated games with incomplete information. In our non-equilibrium setting, players aim to guarantee a certain payoff with high probability, rather than in expected value. We provide a high-probability counterpart of the classical result of Mertens and Zamir for the zero-sum repeated games. Any payoff that can be guaranteed with high probability can be guaranteed in expectation, but the reverse is not true. Hence, unlike the average payoff case where the payoff guaranteed by each player is the negative of the payoff by the other player, the two guaranteed payoffs would differ in the high-probability framework. One motivation for this framework comes from information transmission systems, where it is customary to formulate problems in terms of asymptotically vanishing probability of error. Finally, we introduce compound arbitrarily varying channels, and use the high-probability framework to study this problem. Payam Delgosha, Amin Gohari, Mohammad Akbarpour |
ISIT | 1 |
| 2012 | Information theoretic cutting of a cakeabstractCutting a cake is a metaphor for the problem of dividing a resource (cake) among several agents. The problem becomes non-trivial when the agents have different valuations for different parts of the cake (i.e. one agent may like chocolate while the other may like cream). A fair division of the cake is one that takes into account the individual valuations of agents and partitions the cake based on some fairness criterion. Fair division may be accomplished in a distributed or centralized way. Due to its natural and practical appeal, it has been a subject of study in economics under the topic of “Fair Division”. To best of our knowledge the role of partial information in fair division has not been studied so far from an information theoretic perspective. In this paper we study two important algorithms in fair division, namely “divide and choose” and “adjusted winner” for the case of two agents. We quantify the benefit of negotiation in the divide and choose algorithm, and its use in tricking the adjusted winner algorithm. Lastly we consider a centralized algorithm for maximizing the overall welfare of the agents under the Nash collective utility function (CUF). This corresponds to a clustering problem. Drawing a conceptual link between this problem and the portfolio selection problem in stock markets, we prove an upper bound on the increase of the Nash CUF for a clustering refinement. Payam Delgosha, Amin Gohari |
ITW | 1 |
| 2011 | Randomized Algorithms for Comparison-based SearchabstractThis paper addresses the problem of finding the nearest neighbor (or one of the $R$-nearest neighbors) of a query object $q$ in a database of $n$ objects, when we can only use a comparison oracle. The comparison oracle, given two reference objects and a query object, returns the reference object most similar to the query object. The main problem we study is how to search the database for the nearest neighbor (NN) of a query, while minimizing the questions. The difficulty of this problem depends on properties of the underlying database. We show the importance of a characterization: \emph{combinatorial disorder} $D$ which defines approximate triangle inequalities on ranks. We present a lower bound of $\Omega(D\log \frac{n}{D}+D^2)$ average number of questions in the search phase for any randomized algorithm, which demonstrates the fundamental role of $D$ for worst case behavior. We develop a randomized scheme for NN retrieval in $O(D^3\log^2 n+ D\log^2 n \log\log n^{D^3})$ questions. The learning requires asking $O(n D^3\log^2 n+ D \log^2 n \log\log n^{D^3})$ questions and $O(n\log^2n/\log(2D))$ bits to store. Dominique Tschopp, Suhas N. Diggavi, Payam Delgosha, Soheil Mohajer |
NIPS | 3 |