VLDB 2026 Research / reviewers in the wild / expert
Ashley Prater-Bennette
dblp:158/9018 · also Ashley Prater
· DBLP profile ↗
14ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-4272-423XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Revisiting Large-Scale Non-convex Distributionally Robust OptimizationabstractDistributionally robust optimization (DRO) is a powerful technique to train robust machine learning models that perform well under distribution shifts. Compared with empirical risk minimization (ERM), DRO optimizes the expected loss under the worst-case distribution in
an uncertainty set of distributions. This paper revisits the important problem of DRO with non-convex smooth loss functions. For this problem, Jin et al. (2021) showed that its dual problem is generalized $(L_0, L_1)$-smooth condition and gradient noise satisfies the affine variance condition, designed an algorithm of mini-batch normalized gradient descent with momentum, and proved its convergence and complexity. In this paper, we show that the dual problem and the gradient noise satisfy simpler yet more precise partially generalized smoothness condition and partially affine variance condition by studying the optimization variable and dual variable separately, which further yields much simpler algorithm design and convergence analysis. We develop a double stochastic gradient descent with clipping (D-SGD-C) algorithm that converges to an $\epsilon$-stationary point with $\mathcal O(\epsilon^{-4})$ gradient complexity, which matches with results in Jin et al. (2021). Our algorithm does not need to use momentum, and the proof is much simpler, thanks to the more precise characterization of partially generalized smoothness and partially affine variance noise. We further design a variance-reduced method that achieves a lower gradient complexity of $\mathcal O(\epsilon^{-3})$. Our theoretical results and insights are further verified numerically on a number of tasks, and our algorithms outperform the existing DRO method (Jin et al., 2021). Qi Zhang 0069, Yi Zhou 0017, Simon Khan, Ashley Prater-Bennette, Lixin Shen, Shaofeng Zou |
ICLR | 4 |
| 2025 | Robust Multimodal Learning With Missing Modalities via Parameter-Efficient AdaptationabstractMultimodal learning seeks to utilize data from multiple sources to improve the overall performance of downstream tasks. It is desirable for redundancies in the data to make multimodal systems robust to missing or corrupted observations in some correlated modalities. However, we observe that the performance of several existing multimodal networks significantly deteriorates if one or multiple modalities are absent at test time. To enable robustness to missing modalities, we propose a simple and parameter-efficient adaptation procedure for pretrained multimodal networks. In particular, we exploit modulation of intermediate features to compensate for the missing modalities. We demonstrate that such adaptation can partially bridge performance drop due to missing modalities and outperform independent, dedicated networks trained for the available modality combinations in some cases. The proposed adaptation requires extremely small number of parameters (e.g., fewer than 1% of the total parameters) and applicable to a wide range of modality combinations and tasks. We conduct a series of experiments to highlight the missing modality robustness of our proposed method on five different multimodal tasks across seven datasets. Our proposed method demonstrates versatility across various tasks and datasets, and outperforms existing methods for robust multimodal learning with missing modalities. Md Kaykobad Reza, Ashley Prater-Bennette, Muhammad Salman Asif |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Large-Scale Non-convex Stochastic Constrained Distributionally Robust OptimizationabstractDistributionally robust optimization (DRO) is a powerful framework for training robust models against data distribution shifts. This paper focuses on constrained DRO, which has an explicit characterization of the robustness level. Existing studies on constrained DRO mostly focus on convex loss function, and exclude the practical and challenging case with non-convex loss function, e.g., neural network. This paper develops a stochastic algorithm and its performance analysis for non-convex constrained DRO. The computational complexity of our stochastic algorithm at each iteration is independent of the overall dataset size, and thus is suitable for large-scale applications. We focus on the general Cressie-Read family divergence defined uncertainty set which includes chi^2-divergences as a special case. We prove that our algorithm finds an epsilon-stationary point with an improved computational complexity than existing methods. Our method also applies to the smoothed conditional value at risk (CVaR) DRO. Qi Zhang 0069, Yi Zhou 0017, Ashley Prater-Bennette, Lixin Shen, Shaofeng Zou |
AAAI | 3 |
| 2024 | Robust Average-Reward Reinforcement LearningabstractRobust Markov decision processes (MDPs) aim to find a policy that optimizes the worst-case performance over an uncertainty set of MDPs. Existing studies mostly have focused on the robust MDPs under the discounted reward criterion, leaving the ones under the average-reward criterion largely unexplored. In this paper, we develop the first comprehensive and systematic study of robust average-reward MDPs, where the goal is to optimize the long-term average performance under the worst case. Our contributions are four-folds: (1) we prove the uniform convergence of the robust discounted value function to the robust average-reward function as the discount factor γ goes to 1; (2) we derive the robust average-reward Bellman equation, characterize the structure of its solution set, and prove the equivalence between solving the robust Bellman equation and finding the optimal robust policy; (3) we design robust dynamic programming algorithms, and theoretically characterize their convergence to the optimal policy; and (4) we design two model-free algorithms unitizing the multi-level Monte-Carlo approach, and prove their asymptotic convergence Yue Wang 0068, Alvaro Velasquez, George Atia, Ashley Prater-Bennette, Shaofeng Zou |
J. Artif. Intell. Res. | 4 |
| 2023 | Robust Average-Reward Markov Decision ProcessesabstractIn robust Markov decision processes (MDPs), the uncertainty in the transition kernel is addressed by finding a policy that optimizes the worst-case performance over an uncertainty set of MDPs. While much of the literature has focused on discounted MDPs, robust average-reward MDPs remain largely unexplored. In this paper, we focus on robust average-reward MDPs, where the goal is to find a policy that optimizes the worst-case average reward over an uncertainty set. We first take an approach that approximates average-reward MDPs using discounted MDPs. We prove that the robust discounted value function converges to the robust average-reward as the discount factor goes to 1, and moreover when it is large, any optimal policy of the robust discounted MDP is also an optimal policy of the robust average-reward. We further design a robust dynamic programming approach, and theoretically characterize its convergence to the optimum. Then, we investigate robust average-reward MDPs directly without using discounted MDPs as an intermediate step. We derive the robust Bellman equation for robust average-reward MDPs, prove that the optimal policy can be derived from its solution, and further design a robust relative value iteration algorithm that provably finds its solution, or equivalently, the optimal robust policy. Yue Wang 0068, Alvaro Velasquez, George Atia, Ashley Prater-Bennette, Shaofeng Zou |
AAAI | 4 |
| 2023 | Model-Free Robust Average-Reward Reinforcement LearningabstractRobust Markov decision processes (MDPs) address the challenge of model uncertainty by optimizing the worst-case performance over an uncertainty set of MDPs. In this paper, we focus on the robust average-reward MDPs under the model-free setting. We first theoretically characterize the structure of solutions to the robust average-reward Bellman equation, which is essential for our later convergence analysis. We then design two model-free algorithms, robust relative value iteration (RVI) TD and robust RVI Q-learning, and theoretically prove their convergence to the optimal solution. We provide several widely used uncertainty sets as examples, including those defined by the contamination model, total variation, Chi-squared divergence, Kullback-Leibler (KL) divergence, and Wasserstein distance. Yue Wang 0068, Alvaro Velasquez, George Atia, Ashley Prater-Bennette, Shaofeng Zou |
ICML | 4 |
| 2022 | Scaling and Scalability: Provable Nonconvex Low-Rank Tensor CompletionabstractTensors, which provide a powerful and flexible model for representing multi-attribute data and multi-way interactions, play an indispensable role in modern data science across various fields in science and engineering. A fundamental task is tensor completion, which aims to faithfully recover the tensor from a small subset of its entries in a statistically and computationally efficient manner. Harnessing the low-rank structure of tensors in the Tucker decomposition, this paper develops a scaled gradient descent (ScaledGD) algorithm to directly recover the tensor factors with tailored spectral initializations, and shows that it provably converges at a linear rate independent of the condition number of the ground truth tensor for tensor completion as soon as the sample size is above the order of $n^{3/2}$ ignoring other parameter dependencies, where $n$ is the dimension of the tensor. To the best of our knowledge, ScaledGD is the first algorithm that achieves near-optimal statistical and computational complexities simultaneously for low-rank tensor completion with the Tucker decomposition. Our algorithm highlights the power of appropriate preconditioning in accelerating nonconvex statistical estimation, where the iteration-varying preconditioners promote desirable invariance properties of the trajectory with respect to the underlying symmetry in low-rank tensor factorization. Tian Tong, Cong Ma 0001, Ashley Prater-Bennette, Erin E. Tripp, Yuejie Chi |
AISTATS | 3 |
| 2022 | Incremental Task Learning with Incremental Rank Updates
Rakib Hyder, Ken Shao, Boyu Hou, Panos P. Markopoulos, Ashley Prater-Bennette, Muhammad Salman Asif |
ECCV (23) | 5 |
| 2022 | Scaling and Scalability: Provable Nonconvex Low-Rank Tensor Estimation from Incomplete MeasurementsabstractTensors, which provide a powerful and flexible model for representing multi-attribute data and multi-way interactions, play an indispensable role in modern data science across various fields in science and engineering. A fundamental task is to faithfully recover the tensor from highly incomplete measurements in a statistically and computationally efficient manner. Harnessing the low-rank structure of tensors in the Tucker decomposition, this paper develops a scaled gradient descent (ScaledGD) algorithm to directly recover the tensor factors with tailored spectral initializations, and shows that it provably converges at a linear rate independent of the condition number of the ground truth tensor for two canonical problems --- tensor completion and tensor regression --- as soon as the sample size is above the order of $n^{3/2}$ ignoring other parameter dependencies, where $n$ is the dimension of the tensor. This leads to an extremely scalable approach to low-rank tensor estimation compared with prior art, which suffers from at least one of the following drawbacks: extreme sensitivity to ill-conditioning, high per-iteration costs in terms of memory and computation, or poor sample complexity guarantees. To the best of our knowledge, ScaledGD is the first algorithm that achieves near-optimal statistical and computational complexities simultaneously for low-rank tensor completion with the Tucker decomposition. Our algorithm highlights the power of appropriate preconditioning in accelerating nonconvex statistical estimation, where the iteration-varying preconditioners promote desirable invariance properties of the trajectory with respect to the underlying symmetry in low-rank tensor factorization. Tian Tong, Cong Ma 0001, Ashley Prater-Bennette, Erin E. Tripp, Yuejie Chi |
J. Mach. Learn. Res. | 3 |
| 2020 | Low-Rank Tensor Ring Model for Completing Missing Visual DataabstractLow rank tensor factorization can be viewed as a higher order generalization of low-rank matrix factorization, both of which have been used for image and video representation and reconstruction from compressive measurements. In this paper, we present an algorithm for recovering low-rank tensors from massively under-sampled or missing data. We use low-rank tensor ring (TR) factorization to model images and videos. We observed that TR factorization models are robust to random missing entries but they fail in the cases when large blocks or slices of data are missing. We developed the following two types of algorithms to fill large missing blocks: An algorithm that incrementally updates the tensor rank and implicitly enforces correlations among different modes of the tensor. A framework to incorporate information about similarities between different modes of the tensor to enforce explicit similarity constraints between the missing and known parts of the tensors. We present simulation experiments on YaleB dataset to demonstrate the performance of our methods. Muhammad Salman Asif, Ashley Prater-Bennette |
ICASSP | 2 |
| 2020 | L1-Norm Higher-Order Orthogonal Iterations for Robust Tensor AnalysisabstractStandard Tucker tensor decomposition seeks to maximize the L2-norm of the compressed tensor; thus, it is very responsive to outlying/high-magnitude entries among the processed data. To counteract the impact of outliers in tensor data analysis, we propose L1-Tucker: a reformulation of standard Tucker decomposition, resulting by simple substitution of the outlier-responsive L2-norm by the sturdier L1-norm. Then, we propose the L1-norm Higher Order Orthogonal Iterations (L1-HOOI) algorithm for the approximate solution to L1-Tucker. Our numerical studies on data reconstruction and classification corroborate that L1-HOOI exhibits sturdy resistance against outliers compared to standard counterparts. Dimitris G. Chachlakis, Ashley Prater-Bennette, Panos P. Markopoulos |
ICASSP | 2 |
| 2019 | Multilinear Compressive Sensing With Tensor Ring FactorizationabstractTensor factorization has become a powerful tool for representation and analysis of multi-dimensional data. Low rank tensor factorization can be viewed as a higher order generalization of low-rank matrix factorization, both of which have been used for image and video representation and reconstruction from compressive measurements. In this paper, we present an algorithm for reconstructing images and videos from compressive measurements using tensor ring factorization model. We use a projected gradient descent approach that alternates between gradient descent over a loss function and projection onto a low-rank tensor ring structure. We also present a computationally efficient initialization step for the special case of multilinear compressive sensing. We present simulation results to demonstrate the performance of our algorithm on real images and videos. Muhammad Salman Asif, Ashley Prater-Bennette |
ICIP | 2 |
| 2017 | Spatiotemporal signal classification via principal components of reservoir states
Ashley Prater-Bennette |
Neural Networks | 1 |
| 2015 | Separation of undersampled composite signals using the Dantzig selector with overcomplete dictionariesabstractIn many applications, one may acquire a composition of several signals that may be corrupted by noise, and it is a challenging problem to reliably separate the components from one another without sacrificing significant details. Adding to the challenge, in a compressive sensing framework, one is given only an undersampled set of linear projections of the composite signal. In this study, the authors propose using the Dantzig selector model incorporating an overcomplete dictionary to separate a noisy undersampled collection of composite signals, and present an algorithm to efficiently solve the model. The Dantzig selector is a statistical approach to finding a solution to a noisy linear regression problem by minimising the ℓ 1 norm of candidate coefficient vectors while constraining the scope of the residuals. The Dantzig selector performs well in the recovery and separation of an unknown composite signal when the underlying coefficient vector is sparse. They propose a proximity operator‐based algorithm to recover and separate unknown noisy undersampled composite signals using the Dantzig selector. They present numerical simulations comparing the proposed algorithm with the competing alternating direction method, and the proposed algorithm is found to be faster, while producing similar quality results. In addition, they demonstrate the utility of the proposed algorithm by applying it in various applications including the recovery of complex‐valued coefficient vectors, the removal of impulse noise from smooth signals, and the separation and classification of a composition of handwritten digits. Ashley Prater-Bennette, Lixin Shen |
IET Signal Process. | 1 |