VLDB 2026 Research / reviewers in the wild / expert
Mohammad Pedramfar
dblp:344/1771
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
4 papers |
Mathematical optimization · 100% | |
| Artificial intelligence
2 papers |
Reinforcement learning · 44% Generative modeling · 25% Language models and text generation · 25% |
Topics — the 14 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Mathematical optimization
online optimization |
3.0 | 4 | 2025 | Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025 From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024 Unified Projection-Free Algorithms for Adversarial DR-Submodular Optimization · ICLR 2024 |
Mathematical optimization
submodular optimization |
3.0 | 4 | 2025 | Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025 From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024 Unified Projection-Free Algorithms for Adversarial DR-Submodular Optimization · ICLR 2024 |
Mathematical optimization › submodular optimization
DR-submodular maximization |
2.3 | 3 | 2025 | Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025 From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024 A Unified Approach for Maximizing Continuous DR-submodular Functions · NeurIPS 2023 |
Mathematical optimization › online optimization
regret bounds |
1.6 | 2 | 2025 | Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025 From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Diffusion Tree Sampling: Scalable inference‑time alignment of diffusion models · NeurIPS 2025 |
Natural language and speech › Language models and text generation › alignment
inference-time alignment |
0.9 | 1 | 2025 | Diffusion Tree Sampling: Scalable inference‑time alignment of diffusion models · NeurIPS 2025 |
Mathematical optimization › continuous optimization
convex optimization |
0.9 | 1 | 2025 | Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online Optimization · NeurIPS 2025 |
Mathematical optimization
nonconvex optimization |
0.8 | 1 | 2024 | From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular Optimization · NeurIPS 2024 |
Machine learning › Reinforcement learning
exploration |
0.7 | 1 | 2023 | Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning · NeurIPS 2023 |
Machine learning › Reinforcement learning
thompson sampling |
0.7 | 1 | 2023 | Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning · NeurIPS 2023 |
Mathematical optimization
frank-wolfe algorithm |
0.7 | 1 | 2023 | A Unified Approach for Maximizing Continuous DR-submodular Functions · NeurIPS 2023 |
Mathematical optimization › continuous optimization › convex optimization
oracle complexity |
0.7 | 1 | 2023 | A Unified Approach for Maximizing Continuous DR-submodular Functions · NeurIPS 2023 |
Machine learning › Reinforcement learning › regret minimization
bayesian regret |
0.2 | 1 | 2023 | Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning · NeurIPS 2023 |
Machine learning › Learning theory › online learning
regret bounds |
0.2 | 1 | 2023 | Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
bandit feedback · 3.0frank-wolfe · 2.2zeroth-order optimization · 1.6uniform wrapper · 0.9reward propagation · 0.9monte carlo tree search · 0.9projection-free algorithms · 0.8thompson sampling · 0.7stochastic optimization · 0.7information ratio · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diffusion Tree Sampling: Scalable inference‑time alignment of diffusion modelsabstractAdapting a pretrained diffusion model to new objectives at inference time remains an open problem in generative modeling. Existing steering methods suffer from inaccurate value estimation, especially at high noise levels, which biases guidance. Moreover, information from past runs is not reused to improve sample quality, leading to inefficient use of compute. Inspired by the success of Monte Carlo Tree Search, we address these limitations by casting inference-time alignment as a search problem that reuses past computations. We introduce a tree-based approach that _samples_ from the reward-aligned target density by propagating terminal rewards back through the diffusion chain and iteratively refining value estimates with each additional generation. Our proposed method, Diffusion Tree Sampling (DTS), produces asymptotically exact samples from the target distribution in the limit of infinite rollouts, and its greedy variant Diffusion Tree Search (DTS*) performs a robust search for high reward samples. On MNIST and CIFAR-10 class-conditional generation, DTS matches the FID of the best-performing baseline with up to $5\times$ less compute. In text-to-image generation and language completion tasks, DTS* effectively searches for high reward samples that match best-of-N with $2\times$ less compute. By reusing information from previous generations, we get an _anytime algorithm_ that turns additional compute budget into steadily better samples, providing a scalable approach for inference-time alignment of diffusion models. Vineet Jain, Kusha Sareen, Mohammad Pedramfar, Siamak Ravanbakhsh |
NeurIPS | 3 |
| 2025 | Uniform Wrappers: Bridging Concave to Quadratizable Functions in Online OptimizationabstractThis paper presents novel contributions to the field of online optimization, particularly focusing on the adaptation of algorithms from concave optimization to more challenging classes of functions.
Key contributions include the introduction of uniform wrappers, a class of meta-algorithms that could be used for algorithmic conversions such as converting algorithms for convex optimization into those for quadratizable optimization.
Moreover, we propose a guideline that, given a base algorithm $\mathcal{A}$ for concave optimization and a uniform wrapper $\mathcal{W}$, describes how to convert a proof of the regret bound of $\mathcal{A}$ in the concave setting into a proof of the regret bound of $\mathcal{W}(\mathcal{A})$ for quadratizable setting.
Through this framework, the paper demonstrates improved regret guarantees for various classes of DR-submodular functions under zeroth-order feedback. Furthermore, the paper extends zeroth-order online algorithms to bandit feedback and offline counterparts, achieving notable improvements in regret/sample complexity compared to existing approaches. Mohammad Pedramfar, Christopher J. Quinn, Vaneet Aggarwal |
NeurIPS | 1 |
| 2025 | Learning-Based Two-Tiered Online Optimization of Region-Wide Datacenter Resource AllocationabstractOnline optimization of resource management for large-scale data centers and infrastructures to meet dynamic capacity reservation demands and various practical constraints (e.g., feasibility and robustness) is a very challenging problem. Mixed Integer Programming (MIP) approaches suffer from recognized limitations in such a dynamic environment, while learning-based approaches may face with prohibitively large state/action spaces. To this end, this paper presents a novel two-tiered online optimization to enable a learning-based Resource Allowance System (RAS). To solve optimal server-to-reservation assignment in RAS in an online fashion, the proposed solution leverages a reinforcement learning (RL) agent to make high-level decisions, e.g., how much resource to select from the Main Switch Boards (MSBs), and then a low-level Mixed Integer Linear Programming (MILP) solver to generate the local server-to-reservation mapping, conditioned on the RL decisions. We take into account fault tolerance, server movement minimization, and network affinity requirements and apply the proposed solution to large-scale RAS problems. To provide interpretability, we further train a decision tree model to explain the learned policies and to prune unreasonable corner cases at the low-level MILP solver, resulting in further performance improvement. Extensive evaluations show that our two-tiered solution outperforms baselines such as pure MIP solver by over 15% while delivering$100\times $speedup in computation. Chang-Lin Chen, Hanhan Zhou, Jiayu Chen 0006, Mohammad Pedramfar, Tian Lan 0001, Zheqing Zhu, Pol Mauri Ruiz, Neeraj Kumar 0004, Vaneet Aggarwal |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2024 | Unified Projection-Free Algorithms for Adversarial DR-Submodular OptimizationabstractThis paper introduces unified projection-free Frank-Wolfe type algorithms for adversarial continuous DR-submodular optimization, spanning scenarios such as full information and (semi-)bandit feedback, monotone and non-monotone functions, different constraints, and types of stochastic queries. For every problem considered in the non-monotone setting, the proposed algorithms are either the first with proven sub-linear $\alpha$-regret bounds or have better $\alpha$-regret bounds than the state of the art, where $\alpha$ is a corresponding approximation bound in the offline setting. In the monotone setting, the proposed approach gives state-of-the-art sub-linear $\alpha$-regret bounds among projection-free algorithms in 7 of the 8 considered cases while matching the result of the remaining case. Additionally, this paper addresses semi-bandit and bandit feedback for adversarial DR-submodular optimization, advancing the understanding of this optimization area. Mohammad Pedramfar, Yididiya Y. Nadew, Christopher J. Quinn, Vaneet Aggarwal |
ICLR | 1 |
| 2024 | From Linear to Linearizable Optimization: A Novel Framework with Applications to Stationary and Non-stationary DR-submodular OptimizationabstractThis paper introduces the notion of upper-linearizable/quadratizable functions, a class that extends concavity and DR-submodularity in various settings, including monotone and non-monotone cases over different types of convex sets. A general meta-algorithm is devised to convert algorithms for linear/quadratic maximization into ones that optimize upper-linearizable/quadratizable functions, offering a unified approach to tackling concave and DR-submodular optimization problems. The paper extends these results to multiple feedback settings, facilitating conversions between semi-bandit/first-order feedback and bandit/zeroth-order feedback, as well as between first/zeroth-order feedback and semi-bandit/bandit feedback. Leveraging this framework, new algorithms are derived using existing results as base algorithms for convex optimization, improving upon state-of-the-art results in various cases. Dynamic and adaptive regret guarantees are obtained for DR-submodular maximization, marking the first algorithms to achieve such guarantees in these settings. Notably, the paper achieves these advancements with fewer assumptions compared to existing state-of-the-art results, underscoring its broad applicability and theoretical contributions to non-convex optimization. Mohammad Pedramfar, Vaneet Aggarwal |
NeurIPS | 1 |
| 2023 | Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement LearningabstractIn this paper, we prove state-of-the-art Bayesian regret bounds for Thompson Sampling in reinforcement learning in a multitude of settings. We present a refined analysis of the information ratio, and show an upper bound of order $\widetilde{O}(H\sqrt{d_{l_1}T})$ in the time inhomogeneous reinforcement learning problem where $H$ is the episode length and $d_{l_1}$ is the Kolmogorov $l_1-$dimension of the space of environments. We then find concrete bounds of $d_{l_1}$ in a variety of settings, such as tabular, linear and finite mixtures, and discuss how our results improve the state-of-the-art. Ahmadreza Moradipari, Mohammad Pedramfar, Modjtaba Shokrian Zini, Vaneet Aggarwal |
NeurIPS | 2 |
| 2023 | A Unified Approach for Maximizing Continuous DR-submodular FunctionsabstractThis paper presents a unified approach for maximizing continuous DR-submodular functions that encompasses a range of settings and oracle access types. Our approach includes a Frank-Wolfe type offline algorithm for both monotone and non-monotone functions, with different restrictions on the general convex set. We consider settings where the oracle provides access to either the gradient of the function or only the function value, and where the oracle access is either deterministic or stochastic. We determine the number of required oracle accesses in all cases. Our approach gives new/improved results for nine out of the sixteen considered cases, avoids computationally expensive projections in three cases, with the proposed framework matching performance of state-of-the-art approaches in the remaining four cases. Notably, our approach for the stochastic function value-based oracle enables the first regret bounds with bandit feedback for stochastic DR-submodular functions. Mohammad Pedramfar, Christopher J. Quinn, Vaneet Aggarwal |
NeurIPS | 1 |