Saaduddin Mahmud

dblp:248/9157 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 36% Language models and text generation · 25% Trustworthy machine learning · 20%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.012026
Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models · AAAI 2026
Machine learning › Trustworthy machine learning › interpretability
explainable reinforcement learning
1.012026
Causal Explanations for Sequential Decision Making (Abstract Reprint) · AAAI 2026
Machine learning › Reinforcement learning
markov decision process
1.012026
Causal Explanations for Sequential Decision Making (Abstract Reprint) · AAAI 2026
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization
1.012026
Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models · AAAI 2026
Machine learning › Reinforcement learning › preference learning
active preference learning
0.912025
MAPLE: A Framework for Active Preference Learning Guided by Large Language Models · AAAI 2025
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
bayesian active learning
0.912025
MAPLE: A Framework for Active Preference Learning Guided by Large Language Models · AAAI 2025
Knowledge, reasoning and agents › Multi-agent systems
distributed constraint optimization
0.922020
Learning Optimal Temperature Region for Solving Mixed Integer Functional DCOPs · IJCAI 2020
A Particle Swarm Based Algorithm for Functional Distributed Constraint Optimization Problems · AAAI 2020
Machine learning › Reinforcement learning
preference learning
0.912025
MAPLE: A Framework for Active Preference Learning Guided by Large Language Models · AAAI 2025
Machine learning › Trustworthy machine learning › interpretability
explanation-based learning
0.712023
Explanation-Guided Reward Alignment · IJCAI 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Explanation-Guided Reward Alignment · IJCAI 2023
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.712023
Explanation-Guided Reward Alignment · IJCAI 2023
Machine learning › Reinforcement learning › reinforcement learning from human feedback
reward alignment
0.712023
Explanation-Guided Reward Alignment · IJCAI 2023
Machine learning › Optimization for machine learning › optimization
continuous optimization
0.412020
A Particle Swarm Based Algorithm for Functional Distributed Constraint Optimization Problems · AAAI 2020
Mathematical optimization › metaheuristic optimization
simulated annealing
0.412020
Learning Optimal Temperature Region for Solving Mixed Integer Functional DCOPs · IJCAI 2020
Natural language and speech › Language models and text generation › decoding › decoding strategy
best-of-n sampling
0.312026
Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models · AAAI 2026
Natural language and speech › Language models and text generation
test-time scaling
0.312026
Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models · AAAI 2026

Methods — techniques the papers use, named apart from their topics

structural causal model · 1.0shapley value · 1.0sample complexity analysis · 1.0prompt optimization · 1.0majority voting · 1.0causal inference · 1.0best-of-n sampling · 1.0PSST · 1.0large language model · 0.9bayesian active learning · 0.9simulated annealing · 0.4parameter learning · 0.4
YearPublicationVenuePosition
2026 Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models
abstract
Prompt optimization methods have demonstrated significant effectiveness in aligning black-box large language models (LLMs). In parallel, inference scaling strategies such as Best-of-N Sampling and Majority Voting have likewise been shown to improve alignment and performance by trading additional computation for better output. However, existing prompt optimization approaches are inference strategy agnostic; that is, they optimize prompts without accounting for the inference strategy. This constitutes a significant methodological gap, as our empirical and theoretical analysis reveals a strong interdependence between these two paradigms. Moreover, we find that user preferences regarding trade-offs among multiple objectives and inference budgets substantially influence the choice of prompt and inference configuration. To address this gap, we introduce a novel unified framework named IAPO (Inference-Aware Prompt Optimization) that jointly optimizes the prompt and inference scale, while being aware of the inference budget and different task objectives. We then develop a fixed-budget training algorithm for IAPO, called PSST (Prompt Scaling via Sequential Trimming), and establish finite-budget guarantees on the error probability. Finally, we evaluate the effectiveness of PSST on six tasks, including multi-objective text generation and reasoning, and demonstrate the critical role of incorporating inference-awareness in aligning black-box LLMs using prompt optimization.
Saaduddin Mahmud, Mason Nakamura, Kyle Hollins Wray, Shlomo Zilberstein
AAAI1
2026 Causal Explanations for Sequential Decision Making (Abstract Reprint)
abstract
Stochastic sequential decision-making systems — such as Markov decision processes and their variants — are increasingly used in areas such as transportation, healthcare, and communication. However, the ability to explain these systems’ outputs to non-technical end users has not kept pace with their widespread adoption. This paper addresses that gap by extending prior work and presenting a unified framework for generating causal explanations of agent behavior in sequential decision-making settings, grounded in the structural causal model (SCM) paradigm. Our framework supports the generation of multiple, semantically distinct explanations for agent actions — capabilities that were previously unattainable. In addition to introducing a novel taxonomy of explanations for MDPs to guide empirical investigation, we develop both exact and approximate causal inference methods within the SCM framework. We analyze their applicability and derive run-time bounds for each. This leads to the proposed algorithm, MeanRESP, which operates flexibly across a spectrum of approximations tailored to external constraints. We further analyze the sample complexity and error rates of approximate MeanRESP, and provide a detailed comparison of its outputs — under varying definitions of responsibility — with popular Shapley-value-based methods. Empirically, we performed a series of experiments to evaluate the practicality and effectiveness of the proposed system, focusing on real-world computational demands and the validity and reliability of metrics for comparing approximate and exact causal methods. Finally, we present two user studies that reveal user preferences for certain types of explanations and demonstrate a strong preference for explanations generated by our framework compared to those from other state-of-the-art systems.
Samer B. Nashed, Saaduddin Mahmud, Claudia V. Goldman, Shlomo Zilberstein
AAAI2
2025 MAPLE: A Framework for Active Preference Learning Guided by Large Language Models
abstract
The advent of large language models (LLMs) has sparked significant interest in using natural language for preference learning. However, existing methods often suffer from high computational burdens, taxing human supervision, and lack of interpretability. To address these issues, we introduce MAPLE, a framework for large language model-guided Bayesian active preference learning. MAPLE leverages LLMs to model the distribution over preference functions, conditioning it on both natural language feedback and conventional preference learning feedback, such as pairwise trajectory rankings. MAPLE employs active learning to systematically reduce uncertainty in this distribution and incorporates a language-conditioned active query selection mechanism to identify informative and easy-to-answer queries, thus reducing the burden on humans. We evaluate MAPLE's sample efficiency and preference inference quality across two benchmarks, including a real-world vehicle route planning benchmark using OpenStreetMap data. Our results demonstrate that MAPLE accelerates the learning process and effectively improves humans' ability to answer queries.
Saaduddin Mahmud, Mason Nakamura, Shlomo Zilberstein
AAAI1
2025 Causal Explanations for Sequential Decision Making
abstract
Stochastic sequential decision-making systems — such as Markov decision processes and their variants — are increasingly used in areas such as transportation, healthcare, and communication. However, the ability to explain these systems’ outputs to non-technical end users has not kept pace with their widespread adoption. This paper addresses that gap by extending prior work and presenting a unified framework for generating causal explanations of agent behavior in sequential decision-making settings, grounded in the structural causal model (SCM) paradigm. Our framework supports the generation of multiple, semantically distinct explanations for agent actions — capabilities that were previously unattainable. In addition to introducing a novel taxonomy of explanations for MDPs to guide empirical investigation, we develop both exact and approximate causal inference methods within the SCM framework. We analyze their applicability and derive run-time bounds for each. This leads to the proposed algorithm, MeanRESP, which operates flexibly across a spectrum of approximations tailored to external constraints. We further analyze the sample complexity and error rates of approximate MeanRESP, and provide a detailed comparison of its outputs—under varying definitions of responsibility—with popular Shapley-value-based methods. Empirically, we performed a series of experiments to evaluate the practicality and effectiveness of the proposed system, focusing on real-world computational demands and the validity and reliability of metrics for comparing approximate and exact causal methods. Finally, we present two user studies that reveal user preferences for certain types of explanations and demonstrate a strong preference for explanations generated by our framework compared to those from other state-of-the-art systems.
Samer B. Nashed, Saaduddin Mahmud, Claudia V. Goldman, Shlomo Zilberstein
J. Artif. Intell. Res.2
2023 Explanation-Guided Reward Alignment
abstract
Agents often need to infer a reward function from observations to learn desired behaviors. However, agents may infer a reward function that does not align with the original intent because there can be multiple reward functions consistent with its observations. Operating based on such misaligned rewards can be risky. Furthermore, black-box representations make it difficult to verify the learned rewards and prevent harmful behavior. We present a framework for verifying and improving reward alignment using explanations and show how explanations can help detect misalignment and reveal failure cases in novel scenarios. The problem is formulated as inverse reinforcement learning from ranked trajectories. Verification tests created from the trajectory dataset are used to iteratively validate and improve reward alignment. The agent explains its learned reward and a tester signals whether the explanation passes the test. In cases where the explanation fails, the agent offers alternative explanations to gather feedback, which is then used to improve the learned reward. We analyze the efficiency of our approach in improving reward alignment using different types of explanations and demonstrate its effectiveness in five domains.
Saaduddin Mahmud, Sandhya Saisubramanian, Shlomo Zilberstein
IJCAI1
2023 Learning Constraints on Autonomous Behavior from Proactive Feedback
abstract
Learning from feedback is a common paradigm to acquire information that is hard to specify a priori. In this work, we consider an agent with a known nominal reward model that captures its high-level task objective. Furthermore, the agent operates subject to constraints that are unknown a priori and must be inferred from human interventions. Unlike existing methods, our approach does not rely on full or partial demonstration trajectories or assume a fully reactive human. Instead, we assume access only to sparse interventions, which may in fact be generated proactively by the human, and we only make minimal assumptions about the human. We provide both theoretical bounds on performance and empirical validations of our method. We show that our method enables an agent to learn a constraint set with high accuracy that generalizes well to new environments within a domain, whereas methods that only consider reactive feedback learn an incorrect constraint set that does not generalize well, making constraint violations more likely in new environments.
Connor Basich, Saaduddin Mahmud, Shlomo Zilberstein
IROS2
2020 A Particle Swarm Based Algorithm for Functional Distributed Constraint Optimization Problems
abstract
Distributed Constraint Optimization Problems (DCOPs) are a widely studied constraint handling framework. The objective of a DCOP algorithm is to optimize a global objective function that can be described as the aggregation of several distributed constraint cost functions. In a DCOP, each of these functions is defined by a set of discrete variables. However, in many applications, such as target tracking or sleep scheduling in sensor networks, continuous valued variables are more suited than the discrete ones. Considering this, Functional DCOPs (F-DCOPs) have been proposed that can explicitly model a problem containing continuous variables. Nevertheless, state-of-the-art F-DCOPs approaches experience onerous memory or computation overhead. To address this issue, we propose a new F-DCOP algorithm, namely Particle Swarm based F-DCOP (PFD), which is inspired by a meta-heuristic, Particle Swarm Optimization (PSO). Although it has been successfully applied to many continuous optimization problems, the potential of PSO has not been utilized in F-DCOPs. To be exact, PFD devises a distributed method of solution construction while significantly reducing the computation and memory requirements. Moreover, we theoretically prove that PFD is an anytime algorithm. Finally, our empirical results indicate that PFD outperforms the state-of-the-art approaches in terms of solution quality and computation overhead.
Moumita Choudhury, Saaduddin Mahmud, Md. Mosaddek Khan
AAAI2
2020 Learning Optimal Temperature Region for Solving Mixed Integer Functional DCOPs
abstract
Distributed Constraint Optimization Problems (DCOPs) are an important framework for modeling coordinated decision-making problems in multi-agent systems with a set of discrete variables. Later works have extended DCOPs to model problems with a set of continuous variables, named Functional DCOPs (F-DCOPs). In this paper, we combine both of these frameworks into the Mixed Integer Functional DCOP (MIF-DCOP) framework that can deal with problems regardless of their variables' type. We then propose a novel algorithm - Distributed Parallel Simulated Annealing (DPSA), where agents cooperatively learn the optimal parameter configuration for the algorithm while also solving the given problem using the learned knowledge. Finally, we empirically evaluate our approach in DCOP, F-DCOP, and MIF-DCOP settings and show that DPSA produces solutions of significantly better quality than the state-of-the-art non-exact algorithms in their corresponding settings.
Saaduddin Mahmud, Md. Mosaddek Khan, Moumita Choudhury, Long Tran-Thanh, Nicholas R. Jennings
IJCAI1