Mohammad Pedramfar

dblp:344/1771 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
4 papers
Mathematical optimization · 100%
Artificial intelligence
2 papers
Reinforcement learning · 44% Generative modeling · 25% Language models and text generation · 25%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
online optimization
3.042025
Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025
From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024
Unified Projection-Free Algorithms for Adversarial DR-Submodular Optimization · ICLR 2024
Mathematical optimization
submodular optimization
3.042025
Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025
From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024
Unified Projection-Free Algorithms for Adversarial DR-Submodular Optimization · ICLR 2024
Mathematical optimization › submodular optimization
DR-submodular maximization
2.332025
Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025
From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024
A Unified Approach for Maximizing Continuous DR-submodular Functions · NeurIPS 2023
Mathematical optimization › online optimization
regret bounds
1.622025
Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025
From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.912025
Diffusion Tree Sampling: Scalable inference‑time alignment of diffusion models · NeurIPS 2025
Natural language and speech › Language models and text generation › alignment
inference-time alignment
0.912025
Diffusion Tree Sampling: Scalable inference‑time alignment of diffusion models · NeurIPS 2025
Mathematical optimization › continuous optimization
convex optimization
0.912025
Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025
Mathematical optimization
nonconvex optimization
0.812024
From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024
Machine learning › Reinforcement learning
exploration
0.712023
Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning · NeurIPS 2023
Machine learning › Reinforcement learning
thompson sampling
0.712023
Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning · NeurIPS 2023
Mathematical optimization
frank-wolfe algorithm
0.712023
A Unified Approach for Maximizing Continuous DR-submodular Functions · NeurIPS 2023
Mathematical optimization › continuous optimization › convex optimization
oracle complexity
0.712023
A Unified Approach for Maximizing Continuous DR-submodular Functions · NeurIPS 2023
Machine learning › Reinforcement learning › regret minimization
bayesian regret
0.212023
Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning · NeurIPS 2023
Machine learning › Learning theory › online learning
regret bounds
0.212023
Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

bandit feedback · 3.0frank-wolfe · 2.2zeroth-order optimization · 1.6uniform wrapper · 0.9reward propagation · 0.9monte carlo tree search · 0.9projection-free algorithms · 0.8thompson sampling · 0.7stochastic optimization · 0.7information ratio · 0.7
YearPublicationVenuePosition
2025 Diffusion Tree Sampling: Scalable inference‑time alignment of diffusion models
abstract
Adapting a pretrained diffusion model to new objectives at inference time remains an open problem in generative modeling. Existing steering methods suffer from inaccurate value estimation, especially at high noise levels, which biases guidance. Moreover, information from past runs is not reused to improve sample quality, leading to inefficient use of compute. Inspired by the success of Monte Carlo Tree Search, we address these limitations by casting inference-time alignment as a search problem that reuses past computations. We introduce a tree-based approach that _samples_ from the reward-aligned target density by propagating terminal rewards back through the diffusion chain and iteratively refining value estimates with each additional generation. Our proposed method, Diffusion Tree Sampling (DTS), produces asymptotically exact samples from the target distribution in the limit of infinite rollouts, and its greedy variant Diffusion Tree Search (DTS*) performs a robust search for high reward samples. On MNIST and CIFAR-10 class-conditional generation, DTS matches the FID of the best-performing baseline with up to $5\times$ less compute. In text-to-image generation and language completion tasks, DTS* effectively searches for high reward samples that match best-of-N with $2\times$ less compute. By reusing information from previous generations, we get an _anytime algorithm_ that turns additional compute budget into steadily better samples, providing a scalable approach for inference-time alignment of diffusion models.
Vineet Jain, Kusha Sareen, Mohammad Pedramfar, Siamak Ravanbakhsh
NeurIPS3
2025 Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization
abstract
This paper presents novel contributions to the field of online optimization, particularly focusing on the adaptation of algorithms from concave optimization to more challenging classes of functions. Key contributions include the introduction of uniform wrappers, a class of meta-algorithms that could be used for algorithmic conversions such as converting algorithms for convex optimization into those for quadratizable optimization. Moreover, we propose a guideline that, given a base algorithm $\mathcal{A}$ for concave optimization and a uniform wrapper $\mathcal{W}$, describes how to convert a proof of the regret bound of $\mathcal{A}$ in the concave setting into a proof of the regret bound of $\mathcal{W}(\mathcal{A})$ for quadratizable setting. Through this framework, the paper demonstrates improved regret guarantees for various classes of DR-submodular functions under zeroth-order feedback. Furthermore, the paper extends zeroth-order online algorithms to bandit feedback and offline counterparts, achieving notable improvements in regret/sample complexity compared to existing approaches.
Mohammad Pedramfar, Christopher J. Quinn, Vaneet Aggarwal
NeurIPS1
2025 Learning-Based Two-Tiered Online Optimization of Region-Wide Datacenter Resource Allocation
abstract
Online optimization of resource management for large-scale data centers and infrastructures to meet dynamic capacity reservation demands and various practical constraints (e.g., feasibility and robustness) is a very challenging problem. Mixed Integer Programming (MIP) approaches suffer from recognized limitations in such a dynamic environment, while learning-based approaches may face with prohibitively large state/action spaces. To this end, this paper presents a novel two-tiered online optimization to enable a learning-based Resource Allowance System (RAS). To solve optimal server-to-reservation assignment in RAS in an online fashion, the proposed solution leverages a reinforcement learning (RL) agent to make high-level decisions, e.g., how much resource to select from the Main Switch Boards (MSBs), and then a low-level Mixed Integer Linear Programming (MILP) solver to generate the local server-to-reservation mapping, conditioned on the RL decisions. We take into account fault tolerance, server movement minimization, and network affinity requirements and apply the proposed solution to large-scale RAS problems. To provide interpretability, we further train a decision tree model to explain the learned policies and to prune unreasonable corner cases at the low-level MILP solver, resulting in further performance improvement. Extensive evaluations show that our two-tiered solution outperforms baselines such as pure MIP solver by over 15% while delivering$100\times $speedup in computation.
Chang-Lin Chen, Hanhan Zhou, Jiayu Chen 0006, Mohammad Pedramfar, Tian Lan 0001, Zheqing Zhu, Pol Mauri Ruiz, Neeraj Kumar 0004, Vaneet Aggarwal
IEEE Trans. Netw. Serv. Manag.4
2024 Unified Projection-Free Algorithms for Adversarial DR-Submodular Optimization
abstract
This paper introduces unified projection-free Frank-Wolfe type algorithms for adversarial continuous DR-submodular optimization, spanning scenarios such as full information and (semi-)bandit feedback, monotone and non-monotone functions, different constraints, and types of stochastic queries. For every problem considered in the non-monotone setting, the proposed algorithms are either the first with proven sub-linear $\alpha$-regret bounds or have better $\alpha$-regret bounds than the state of the art, where $\alpha$ is a corresponding approximation bound in the offline setting. In the monotone setting, the proposed approach gives state-of-the-art sub-linear $\alpha$-regret bounds among projection-free algorithms in 7 of the 8 considered cases while matching the result of the remaining case. Additionally, this paper addresses semi-bandit and bandit feedback for adversarial DR-submodular optimization, advancing the understanding of this optimization area.
Mohammad Pedramfar, Yididiya Y. Nadew, Christopher J. Quinn, Vaneet Aggarwal
ICLR1
2024 From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization
abstract
This paper introduces the notion of upper-linearizable/quadratizable functions, a class that extends concavity and DR-submodularity in various settings, including monotone and non-monotone cases over different types of convex sets. A general meta-algorithm is devised to convert algorithms for linear/quadratic maximization into ones that optimize upper-linearizable/quadratizable functions, offering a unified approach to tackling concave and DR-submodular optimization problems. The paper extends these results to multiple feedback settings, facilitating conversions between semi-bandit/first-order feedback and bandit/zeroth-order feedback, as well as between first/zeroth-order feedback and semi-bandit/bandit feedback. Leveraging this framework, new algorithms are derived using existing results as base algorithms for convex optimization, improving upon state-of-the-art results in various cases. Dynamic and adaptive regret guarantees are obtained for DR-submodular maximization, marking the first algorithms to achieve such guarantees in these settings. Notably, the paper achieves these advancements with fewer assumptions compared to existing state-of-the-art results, underscoring its broad applicability and theoretical contributions to non-convex optimization.
Mohammad Pedramfar, Vaneet Aggarwal
NeurIPS1
2023 Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning
abstract
In this paper, we prove state-of-the-art Bayesian regret bounds for Thompson Sampling in reinforcement learning in a multitude of settings. We present a refined analysis of the information ratio, and show an upper bound of order $\widetilde{O}(H\sqrt{d_{l_1}T})$ in the time inhomogeneous reinforcement learning problem where $H$ is the episode length and $d_{l_1}$ is the Kolmogorov $l_1-$dimension of the space of environments. We then find concrete bounds of $d_{l_1}$ in a variety of settings, such as tabular, linear and finite mixtures, and discuss how our results improve the state-of-the-art.
Ahmadreza Moradipari, Mohammad Pedramfar, Modjtaba Shokrian Zini, Vaneet Aggarwal
NeurIPS2
2023 A Unified Approach for Maximizing Continuous DR-submodular Functions
abstract
This paper presents a unified approach for maximizing continuous DR-submodular functions that encompasses a range of settings and oracle access types. Our approach includes a Frank-Wolfe type offline algorithm for both monotone and non-monotone functions, with different restrictions on the general convex set. We consider settings where the oracle provides access to either the gradient of the function or only the function value, and where the oracle access is either deterministic or stochastic. We determine the number of required oracle accesses in all cases. Our approach gives new/improved results for nine out of the sixteen considered cases, avoids computationally expensive projections in three cases, with the proposed framework matching performance of state-of-the-art approaches in the remaining four cases. Notably, our approach for the stochastic function value-based oracle enables the first regret bounds with bandit feedback for stochastic DR-submodular functions.
Mohammad Pedramfar, Christopher J. Quinn, Vaneet Aggarwal
NeurIPS1