EDBT 2026 Demo / reviewers in the wild / expert
Setareh Maghsudi
dblp:30/10806
· DBLP profile ↗
47ranked-venue papers
11as first author
34since 2021 · last 2026
0000-0002-0647-611XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 28 · 10 first-author · 16 since 2021Artificial intelligence and machine learning · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Optimization and Learning-Based Hide-and-Seek for Resilient Network Design
Setareh Maghsudi |
ICC | 2 |
| 2026 | Optimal Radio Resource Management for ISAC Under Imperfect Information: A Resource Economy-Driven PerspectiveabstractThis work investigates the radio resource management (RRM) design for downlink integrated sensing and communications (ISAC) systems, jointly optimizing timeslot allocation, beam adaptation, functionality selection, and user-target pairing, with the goal of economizing resource consumption under imperfect information. Timeslot allocation assigns a number of discrete channel uses to targets and users, while beam adaptation selects transmit and receive beams with suitable directions, power levels, and beamwidths. Functionality selection determines whether each timeslot is used for sensing, communication, or their simultaneous operation, while user-target pairing specifies which users and targets are jointly served within the same timeslot. To ensure reliable operation, information imperfections arising from motion, quantization, feedback delays, and hardware limitations are considered. Resource economization is achieved by minimizing energy and time consumption through a multi-objective function, with strict prioritization of time savings. The resulting RRM problem is formulated as a semi-infinite, nonconvex mixed-integer nonlinear program (MINLP). Given the lack of generic methods for solving such problems, we propose a tailor-made approach that exploits the underlying structure of the problem to uncover hidden convexities. This enables an exact reformulation as a mixed-integer semidefinite program (MISDP), which can be solved to global optimality. Simulations reveal important interdependencies among the considered RRM components and show that the proposed approach achieves substantial performance improvements over baseline schemes, with gains up to 88%. Luis F. Abanto-Leon, Setareh Maghsudi |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | Service Placement in Small Cell Networks Using Distributed Best arm Identification in Linear BanditsabstractAs users in small cell networks increasingly rely on computation-intensive services, cloud-based access often results in high latency. Multi-access edge computing (MEC) mitigates this by bringing computational resources closer to end users, with small base stations (SBSs) serving as edge servers to enable low-latency service delivery. However, limited edge capacity makes it challenging to decide which services to deploy locally versus in the cloud, especially under unknown service demand and dynamic network conditions. To tackle this problem, we model service demand as a linear function of service attributes and formulate the service placement task as a linear bandit problem, where SBSs act as agents and services as arms. The goal is to identify the service that, when placed at the edge, offers the greatest reduction in total user delay compared to cloud deployment. We propose a distributed and adaptive multi-agent best-arm identification (BAI) algorithm under a fixed-confidence setting, where SBSs collaborate to accelerate learning. Simulations show that our algorithm identifies the optimal service with the desired confidence and achieves near-optimal speedup, as the number of learning rounds decreases proportionally with the number of SBSs. We also provide theoretical analysis of the algorithm's sample complexity and communication overhead. Mariam Yahya, Aydin Sezgin, Setareh Maghsudi |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | Efficient Resource Allocation Under Adversary Attacks: A Decomposition-Based ApproachabstractWe address the problem of allocating limited resources in a network under persistent yet statistically unknown adversarial attacks. Each node in the network may be degraded, but not fully disabled, depending on its available defensive resources. The objective is twofold: to minimize total system damage and to reduce cumulative resource allocation and transfer costs over time. We model this challenge as a bi-objective optimization problem and propose a decomposition-based solution that integrates chance-constrained programming with network flow optimization. The framework separates the problem into two interrelated subproblems: determining optimal node-level allocations across time slots, and computing efficient inter-node resource transfers. We theoretically prove the convergence of our method to the optimal solution that would be obtained with full statistical knowledge of the adversary. We further establish anO(√Tlog(nT)) regret bound, showing that the average per-round performance gap shrinks asO(1/√T). Extensive simulations demonstrate that our method efficiently learns the adversarial patterns and achieves substantial gains in minimizing both damage and operational costs, comparing three benchmark strategies under various parameter settings. Mansoor Davoodi Monfared, Setareh Maghsudi |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2025 | Robust Inverse Reinforcement Learning Under State Adversarial PerturbationsabstractState-adversarial perturbations—arising from sensor spoofing, environmental interference, or targeted attacks—corrupt observations and invalidate state-wise optimality assumptions commonly made in IRL. We study inverse reinforcement learning (IRL) in state-adversarial MDPs (SA-MDPs) where only perturbed states are observable and propose SAMM-IRL, a max-margin IRL framework that operates purely in the belief (perturbed) space without access to clean states. In contrast to point-wise, state-wise optimality, we adopt a robust optimality notion based on the expected return over the initial-state distribution, which is well-posed under adversarial observation mappings. We prove (i) the existence of robust optimal policies in SA-MDPs, (ii) the contraction properties of intermediate RL operators under fixed and adaptive adversaries, and the iteration bounds for SAMM-IRL max-margin updates in belief space. Empirically, in discrete GridWorld and continuous control, SAMM-IRL achieves stronger reward recovery and imitation performance under adversarial observations than baselines, while maintaining stable policy updates. We further report perturbation parameters and ablation results in the main text to support reproducibility and practical deployment. Mine Melodi Caliskan, Saeed Ghoorchian, Setareh Maghsudi |
ECAI | 3 |
| 2025 | Emergence of Fair Leaders via Mediators in Multi-Agent Reinforcement LearningabstractStackelberg games and their resulting equilibria have received increasing attention in the multi-agent reinforcement learning literature. Each stage of a traditional Stackelberg game involves a leader(s) acting first, followed by the followers. In situations where the roles of leader(s) and followers can be interchanged, the designated role can have considerable advantages, for example, in first-mover advantage settings. Then the question arises: Who should be the leader and when? A bias in the leader selection process can lead to unfair outcomes. This problem is aggravated if the agents are self-interested and care only about their goals and rewards. We formally define this leader selection problem and show its relation to fairness in agents’ returns. Furthermore, we propose a multi-agent reinforcement learning framework that maximizes fairness by integrating mediators. Mediators have previously been used in the simultaneous action setting with varying levels of control, such as directly performing agents’ actions or just recommending them. Our framework integrates mediators in the Stackelberg setting with minimal control (leader selection). We show that the presence of mediators leads to self-interested agents taking fair actions, resulting in higher overall fairness in agents’ returns. Akshay Dodwadmath, Setareh Maghsudi |
ECAI | 2 |
| 2025 | Federated Learning with Heterogeneous Feature Adaptation for Human Activity RecognitionabstractFederated learning promotes knowledge sharing in data-sensitive domains, such as Human Activity Recognition (HAR). However, data heterogeneity, namely, non-iid feature, can degrade the performance by causing client drift. We propose an effective knowledge distillation method incorporating a novel batch normalization setup within the federated learning aggregation process. This approach enables the global model to align closely with client models in parameter and feature spaces. Our method demonstrates superior performance and warm start capability compared to other approaches across various non-iid HAR datasets. Tobias Schlagenhauf, Setareh Maghsudi |
ICASSP | 3 |
| 2025 | Memristor-Based Meta-Learning for Fast mmWave Beam Prediction in Non-Stationary EnvironmentsabstractTraditional machine learning techniques have achieved great success in improving data-rate performance and reducing latency in millimeter wave (mmWave) communications. However, these methods still face two key challenges: (i) their reliance on large-scale paired data for model training and tuning, which limits performance gains and makes beam predictions outdated, especially in multi-user mmWave systems with large antenna arrays, and (ii) meta-learning (ML)-based beamforming solutions are prone to overfitting when trained on a limited number of tasks. To address these issues, we propose a memristorbased meta-learning (M-ML) framework for predicting mmWave beam in real time. The M-ML framework generates optimal initialization parameters during the training phase, providing a strong starting point for adapting to unknown environments during the testing phase. By leveraging memory to store key data, M-ML ensures the predicted beamforming vectors are wellsuited to episodically dynamic channel distributions, even when testing and training environments do not align. Simulation results show that our approach delivers high prediction accuracy in new environments, without relying on large datasets. Moreover, MML enhances the model's generalization ability and adaptability. Wenqin Lu, Tomoaki Ohtsuki, Setareh Maghsudi, Xueqin Jiang 0001, Charalampos Tsimenidis |
ICC | 4 |
| 2025 | Generation of Programmatic Rules for Document Forgery Detection Using Large Language ModelsabstractDocument forgery poses a growing threat to legal, economic, and governmental processes, requiring increasingly sophisticated verification mechanisms. One approach involves the use of plausibility checks, rule-based procedures that assess the correctness and internal consistency of data, to detect anomalies or signs of manipulation. Although these verification procedures are essential for ensuring data integrity, existing plausibility checks are manually implemented by software engineers, which is time-consuming. Recent advances in code generation with large language models (LLMs) offer new potential for automating and scaling the generation of these checks. However, adapting LLMs to the specific requirements of an unknown domain remains a significant challenge. This work investigates the extent to which LLMs, adapted on domain-specific code and data through different fine-tuning strategies, can generate rule-based plausibility checks for forgery detection on constrained hardware resources. We fine-tune open-source LLMs, Llama 3.1 8B and OpenCoder 8B, on structured datasets derived from real-world application scenarios and evaluate the generated plausibility checks on previously unseen forgery patterns. The results demonstrate that the models are capable of generating executable and effective verification procedures. This also highlights the potential of LLMs as scalable tools to support human decision-making in security-sensitive contexts where comprehensibility is required. Valentin Schmidberger, Manuel Eberhardinger, Setareh Maghsudi, Johannes Maucher |
ICMLA | 3 |
| 2025 | Pareto Multi-objective Alignment for Language Models
Setareh Maghsudi |
ECML/PKDD (4) | 2 |
| 2024 | Meta Learning in Bandits within shared affine SubspacesabstractWe study the problem of meta-learning several contextual stochastic bandits tasks by leveraging their concentration around a low dimensional affine subspace, which we learn via online principal component analysis to reduce the expected regret over the encountered bandits. We propose and theoretically analyze two strategies that solve the problem: One based on the principle of optimism in the face of uncertainty and the other via Thompson sampling. Our framework is generic and includes previously proposed approaches as special cases. Besides, the empirical results show that our methods significantly reduce the regret on several bandit tasks. Steven Bilaj, Sofien Dhouib, Setareh Maghsudi |
AISTATS | 3 |
| 2024 | Age-Based Federated Learning Approach to In-Network Caching: An Online Scheduling PolicyabstractWe develop an accurate real-time scheduling framework for federated learning (FL) in wireless caching networks to guarantee the successful delivery of files at a low cost and with a short delay. The following persisting challenges motivated our work: i) Enforcing an excessive number of FL model update per communication round is infeasible due to the limited backhaul spectrum; ii) Naive scheduling policy accounting for FL model update renders service backlogs, thus leading to network parameter staleness and in-network caching utility (ICU) deterioration. Optimal scheduling in FL is challenging, as the mobile users' preferences for content, request patterns, and network traffic are dynamic and unknown. To tackle that challenge, we first formulate an instantaneous ICU optimization problem against the stale FL models. Afterward, based on the concept of age-of-update (AoU), we propose a federated learning with an unsatisfactory set selection (FedUSS) approach capable of executing the multiple-tasks of short-term predictions and making cache replacement decisions at low cost. Theoretical and numerical analyses manifest the effectiveness of our approach. Setareh Maghsudi, Tomoaki Ohtsuki |
ICC | 2 |
| 2024 | Adaptive Regularization of Representation Rank as an Implicit Constraint of Bellman EquationabstractRepresentation rank is an important concept for understanding the role of Neural Networks (NNs) in Deep Reinforcement learning (DRL), which measures the expressive capacity of value networks. Existing studies focus on unboundedly maximizing this rank; nevertheless, that approach would introduce overly complex models in the learning, thus undermining performance. Hence, fine-tuning representation rank presents a challenging and crucial optimization problem. To address this issue, we find a guiding principle for adaptive control of the representation rank. We employ the Bellman equation as a theoretical foundation and derive an upper bound on the cosine similarity of consecutive state-action pairs representations of value networks. We then leverage this upper bound to propose a novel regularizer, namely BEllman Equation-based automatic rank Regularizer (BEER). This regularizer adaptively regularizes the representation rank, thus improving the DRL agent's performance. We first validate the effectiveness of automatic control of rank on illustrative experiments. Then, we scale up BEER to complex continuous control tasks by combining it with the deterministic policy gradient method. Among 12 challenging DeepMind control tasks, BEER outperforms the baselines by a large margin. Besides, BEER demonstrates significant advantages in Q-value approximation. Our code is available at https://github.com/sweetice/BEER-ICLR2024. Tianyi Zhou 0001, Setareh Maghsudi |
ICLR | 4 |
| 2024 | Gradual Change Detection in Covariance Matrix: A Lazy ApproachabstractThanks to its slow-varying characteristic and relatively low requirement for estimation overhead, the covariance matrix has been extensively researched in sixth-generation (6G) wireless systems. Nevertheless, user mobility in practice will cause a gradual change in the covariance matrix, thereby deteriorating the system's performance if no update of the covariance matrix is applied. In this paper, we study the problem of efficient detection of gradual changes in the covariance matrix. We first introduce four change-point detectors that directly map the observations to change in our target KPI. Then, we propose a low-overhead detection algorithm that omits unnecessary channel estimations by adapting an AoA-based estimation trigger. Simulation results show that our proposed scheme can provide near-optimal performance while drastically reducing the estimation and computation overhead. Sida Dai, Ehsan Tohidi, Setareh Maghsudi, Lars Thiele, Slawomir Stanczak |
WCNC | 3 |
| 2024 | Budgeted Recommendation with Delayed Feedback
Kwei-guu Liu, Setareh Maghsudi, Makoto Yokoo |
WorldCIST (3) | 2 |
| 2024 | Mobility-Aware Routing and Caching in Small Cell Networks Using Federated LearningabstractWe consider a service cost minimization problem for resource-constrained small-cell networks with caching, where the challenge mainly stems from (i) the insufficient backhaul capacity and limited network bandwidth and (ii) the limited storing capacity of small-cell base stations (SBSs). Besides, the optimization problem is NP-hard since both the users’ mobility patterns and content preferences are unknown. In this paper, we develop a novel mobility-aware joint routing and caching strategy to address the challenges. The designed framework divides the entire geographical area into small sections containing one SBS and several mobile users (MUs). Based on the concept of one-stop-shop (OSS), we propose a federated routing and popularity learning (FRPL) approach in which the SBSs cooperatively learn the routing and preference of their respective MUs and make a caching decision. The FRPL method completes multiple tasks in one shot, thus reducing the average processing time per global aggregation of learning. By exploiting the outcomes of FRPL together with the estimated service edge of SBSs, the proposed cache placement solution greedily approximates the minimizer of the challenging service cost optimization problem. Theoretical and numerical analyses show the effectiveness of our proposed approaches. Setareh Maghsudi, Tomoaki Ohtsuki, Tony Q. S. Quek |
IEEE Trans. Commun. | 2 |
| 2024 | A Repeated Auction Model for Load-Aware Dynamic Resource Allocation in Multi-Access Edge ComputingabstractMulti-access edge computing (MEC) is one of the enabling technologies for high-performance computing at the edge of the 6 G networks, supporting high data rates and ultra-low service latency. Although MEC is a remedy to meet the growing demand for computation-intensive applications, the scarcity of resources at the MEC servers degrades its performance. Hence, effective resource management is essential; nevertheless, state-of-the-art research lacks efficient economic models to support the exponential growth of the MEC-enabled applications market. We focus on designing a MEC offloading service market based on a repeated auction model with multiple resource sellers (e.g., network operators and service providers) that compete to sell their computing resources to the offloading users. We design a computationally-efficient modified Generalized Second Price (GSP)-based algorithm that decides on pricing and resource allocation by considering the dynamic offloading requests arrival and the servers' computational workloads. Besides, we propose adaptive best-response bidding strategies for the resource sellers, satisfying the symmetric Nash equilibrium (SNE) and individual rationality properties. Finally, via intensive numerical results, we show the effectiveness of our proposed resource allocation mechanism. Ummy Habiba, Setareh Maghsudi, Ekram Hossain 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Federated Learning in UAV-Enhanced Networks: Joint Coverage and Convergence Time OptimizationabstractFederated learning (FL) involves several devices that collaboratively train a shared model without transferring their local data. FL reduces the communication overhead, making it a promising learning method in UAV-enhanced wireless networks with scarce energy resources. Despite the potential, implementing FL in UAV-enhanced networks is challenging, as conventional UAV placement methods that maximize coverage increase the FL delay significantly. Moreover, the uncertainty and lack of a priori information about crucial variables, such as channel quality, exacerbate the problem. In this paper, we first analyze the statistical characteristics of a UAV-enhanced wireless sensor network (WSN) with energy harvesting. We then develop a model and solution based on the multi-objective multi-armed bandit theory to maximize the network coverage while minimizing the FL delay. Besides, we propose another solution that is particularly useful with large action sets and strict energy constraints at the UAVs. Our proposal uses a scalarized best-arm identification algorithm to find the optimal arms that maximize the ratio of the expected reward to the expected energy cost by sequentially eliminating one or more arms in each round. Then, we derive the upper bound on the error probability of our multi-objective and cost-aware algorithm. Numerical results show the effectiveness of our approach. Mariam Yahya, Setareh Maghsudi, Slawomir Stanczak |
IEEE Trans. Wirel. Commun. | 2 |
| 2023 | Piecewise-Stationary Combinatorial Semi-Bandit with Causally Related RewardsabstractWe study the piecewise stationary combinatorial semi-bandit problem with causally related rewards. In our nonstationary environment, variations in the base arms’ distributions, causal relationships between rewards, or both, change the reward generation process. In such an environment, an optimal decision-maker must follow both sources of change and adapt accordingly. The problem becomes aggravated in the combinatorial semi-bandit setting, where the decision-maker only observes the outcome of the selected bundle of arms. The core of our proposed policy is the Upper Confidence Bound (UCB) algorithm. We assume the agent relies on an adaptive approach to overcome the challenge. More specifically, it employs a change-point detector based on the Generalized Likelihood Ratio test. Besides, we introduce the notion of group restart as a new alternative restarting strategy in the decision making process in structured environments. Finally, our algorithm integrates a mechanism to trace the variations of the underlying graph structure, which captures the causal relationships between the rewards in the bandit setting. Theoretically, we establish a regret upper bound that reflects the effects of the number of structural- and distribution changes on the performance. The outcome of our numerical experiments in real-world scenarios exhibits applicability and superior performance of our proposal compared to the state-of-the-art benchmarks. Steven Bilaj, Amir Rezaei Balef, Setareh Maghsudi |
ECAI | 3 |
| 2023 | Cooperative Thresholded Lasso for Sparse Linear BanditabstractWe present a novel approach to address the multi-agent sparse contextual linear bandit problem, in which the feature vectors have a high dimension d whereas the reward function depends on only a limited set of features - precisely s0 ≪ d. Furthermore, the learning follows under information-sharing constraints. The proposed method employs Lasso regression for dimension reduction, allowing each agent to independently estimate an approximate set of main dimensions and share that information with others depending on the network’s structure. The information is then aggregated through a specific process and shared with all agents. Each agent then resolves the problem with ridge regression focusing solely on the extracted dimensions. We represent algorithms for both a star-shaped network and a peer-to-peer network. The approaches effectively reduce communication costs while ensuring minimal cumulative regret per agent. Theoretically, we show that our proposed methods have a regret bound of order O(s0 log d + s0 √T) with high probability, where T is the time horizon. To our best knowledge, it is the first algorithm that tackles row-wise distributed data in sparse linear bandits, achieving comparable performance compared to the state-of-the-art single and multi-agent methods. Besides, it is widely applicable to high-dimensional multi-agent problems where efficient feature extraction is critical for minimizing regret. To validate the effectiveness of our approach, we present experimental results on both synthetic and real-world datasets. Xiaotong Cheng, Setareh Maghsudi |
ECAI | 2 |
| 2023 | A Deep Reinforcement Learning Approach for Load Balancing in Open Radio Access NetworksabstractThe Open RAN paradigm offers data-driven, intelligent optimization of the radio access network (RAN). The disaggregated nature of the Open RAN combined with virtualization on general-purpose CPUs with limited computation capacity creates different load types at multiple levels, making it more challenging to balance the load within the network. This paper proposes a learning framework that learns the assignment of users (UEs) to network nodes to balance the communication and computation load in the network. The framework incorporates communication resources consumed by the users in the radio unit (RU), and computation resources needed for baseband processing in the virtualized distributed unit (DU). The goal is thus to balance the communication load between RUs and the computation load between DUs to avoid overloading network elements or to handle higher peak data rate demands when new users arrive in the network. We apply a novel utility-based approach to jointly optimize the UE-RU and RU-DU assignments taking into account the users' QoS (quality of service) requirements. Simulations demonstrate that the proposed method generates the assignments that significantly improve the network load conditions compared to baseline schemes, thereby enabling more available communication and computation resources for incoming peak data rate users in the network. Hammad Zafar, Martin Kasparick 0001, Setareh Maghsudi, Slawomir Stanczak |
GLOBECOM | 3 |
| 2023 | A Bandit Online Convex Optimization Approach To Distributed Energy Management In Networked SystemsabstractModern power systems integrate renewable distributed energy resources (DERs) as an environment-friendly enhancement to meet the ever-increasing demands. However, due to the inherent unreliability of renewable energy, it is imperative to develop effective algorithms for DER management. In this work, we study the energy-sharing problem in a system consisting of several DERs. Each agent harvests and distributes renewable energy in its neighborhood to optimize the network's performance while minimizing energy waste. We model this problem as a bandit convex optimization problem with constraints, where the constraints correspond to each node's limitations for energy production. We propose a distributed decision-making policy to solve the formulated problem, that achieves ${\mathcal{O}}\left( {{T^{\frac{3}{4}}}} \right)$ regret bound and ${\mathcal{O}}\left( {{T^{\frac{3}{4}}}} \right)$ constraint violations. To reduce the constraint violations, we suggest two decision-making variations. Numerical experiments using a real-world dataset show superior performance of our proposal compared to state-of-the-art methods. Ioannis Tsetis, Xiaotong Cheng, Setareh Maghsudi |
ICASSP | 3 |
| 2023 | Multi-Agent Learning from LearnersabstractA large body of the "Inverse Reinforcement Learning" (IRL) literature focuses on recovering the reward function from a set of demonstrations of an expert agent who acts optimally or noisily optimally. Nevertheless, some recent works move away from the optimality assumption to study the "Learning from a Learner (LfL)" problem, where the challenge is inferring the reward function of a learning agent from a sequence of demonstrations produced by progressively improving policies. In this work, we take one of the initial steps in addressing the multi-agent version of this problem and propose a new algorithm, MA-LfL (Multiagent Learning from a Learner). Unlike the state-of-the-art literature, which recovers the reward functions from trajectories produced by agents in some equilibrium, we study the problem of inferring the reward functions of interacting agents in a general sum stochastic game without assuming any equilibrium state. The MA-LfL algorithm is rigorously built on a theoretical result that ensures its validity in the case of agents learning according to a multi-agent soft policy iteration scheme. We empirically test MA-LfL and we observe high positive correlation between the recovered reward functions and the ground truth. Mine Melodi Caliskan, Francesco Chini, Setareh Maghsudi |
ICML | 3 |
| 2023 | Parallel Online Clustering of Bandits via Hedonic GameabstractContextual bandit algorithms appear in several applications, such as online advertisement and recommendation systems like personalized education or personalized medicine. Individually-tailored recommendations boost the performance of the underlying application; nevertheless, providing individual suggestions becomes costly and even implausible as the number of users grows. As such, to efficiently serve the demands of several users in modern applications, it is imperative to identify the underlying users’ clusters, i.e., the groups of users for which a single recommendation might be (near-)optimal. We propose CLUB-HG, a novel algorithm that integrates a game-theoretic approach into clustering inference. Our algorithm achieves Nash equilibrium at each inference step and discovers the underlying clusters. We also provide regret analysis within a standard linear stochastic noise setting. Finally, experiments on synthetic and real-world datasets show the superior performance of our proposed algorithm compared to the state-of-the-art algorithms. Xiaotong Cheng, Setareh Maghsudi |
ICML | 3 |
| 2023 | Eigensubspace of Temporal-Difference Dynamics and How It Improves Value Approximation in Reinforcement Learning
Tianyi Zhou 0001, Setareh Maghsudi |
ECML/PKDD (4) | 4 |
| 2023 | Distributed Cooperation Under Uncertainty in Drone-Based Wireless Networks: A Bayesian Coalitional GameabstractWe study the resource sharing problem in a drone-based wireless network by considering a distributed control setting under uncertainty (e.g., due to lack of full information). The drones cooperate in serving the users while pooling their spectrum and energy resources in the absence of prior knowledge about different system characteristics such as the amount of available power at the other drones. Compared to the state-of-the-art research in drone-based wireless networks, which is mainly based on the assumption of accurate global information availability at every drone, our setting is realistic and practical. We cast the efficient resource pooling problem as a Bayesian cooperative game in which the agents (drones) engage in a coalition formation process, where the goal is to maximize the overall transmission rate of the network. The drones update their beliefs by using a novel technique that combines the maximum likelihood estimation with Kullback-Leibler divergence. We propose a decision-making strategy for repeated coalition formation that converges to a stable coalition structure. We analyze the performance of the proposed approach by both theoretical analysis and simulations. We provide the comparison of our scheme with the baseline and the socially optimal solution obtained from the exhaustive search. Simulation results demonstrate the superior performance of the proposed method in terms of the sum-rate of the network, the individual rate of the drones, and convergence properties. Vandana Mittal, Setareh Maghsudi, Ekram Hossain 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Data-Driven Online Recommender Systems With Costly Information AcquisitionabstractIn numerous recommender systems, collecting useful information from users is costly, implying that the recommender system has to make active choices by simultaneously learning the observations of the features' states to make useful recommendations to users among available products and services. This paper integrates information acquisition decisions into recommender system. To solve the aforementioned dual learning problem, we propose two different algorithms, namely Sim-OOS and Seq-OOS, where observations are made simultaneously and sequentially, respectively. We prove that both algorithms guarantee a sub linear regret. The developed recommender system can be applied to a variety of real-world applications, including medical informatics, smart transportation, finance, and cyber-security where collecting information before making decisions results in an excessive cost. We validate and evaluate our proposed policies in a medical decision-support system that recommends tests and treatments for breast cancer patients. Onur Atan, Saeed Ghoorchian, Setareh Maghsudi, Mihaela van der Schaar |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Linear Combinatorial Semi-Bandit with Causally Related RewardsabstractIn a sequential decision-making problem, having a structural dependency amongst the reward distributions associated with the arms makes it challenging to identify a subset of alternatives that guarantees the optimal collective outcome. Thus, besides individual actions' reward, learning the causal relations is essential to improve the decision-making strategy. To solve the two-fold learning problem described above, we develop the 'combinatorial semi-bandit framework with causally related rewards', where we model the causal relations by a directed graph in a stationary structural equation model. The nodal observation in the graph signal comprises the corresponding base arm's instantaneous reward and an additional term resulting from the causal influences of other base arms' rewards. The objective is to maximize the long-term average payoff, which is a linear function of the base arms' rewards and depends strongly on the network topology. To achieve this objective, we propose a policy that determines the causal relations by learning the network's topology and simultaneously exploits this knowledge to optimize the decision-making process. We establish a sublinear regret bound for the proposed algorithm. Numerical experiments using synthetic and real-world datasets demonstrate the superior performance of our proposed method compared to several benchmarks. Behzad Nourani-Koliji, Saeed Ghoorchian, Setareh Maghsudi |
IJCAI | 3 |
| 2022 | A Learning-Based Approach to Approximate Coded ComputationabstractLagrange coded computation (LCC) is essential to solving problems about matrix polynomials in a coded distributed fashion; nevertheless, it can only solve the problems that are representable as matrix polynomials. In this paper, we propose AICC, an AI-aided learning approach that is inspired by LCC but also uses deep neural networks (DNNs). It is appropriate for coded computation of more general functions. Numerical simulations demonstrate the suitability of the proposed approach for the coded computation of different matrix functions that are often utilized in digital signal processing. Navneet Agrawal, Yuqin Qiu, Matthias Frey, Igor Bjelakovic, Setareh Maghsudi, Slawomir Stanczak, Jingge Zhu |
ITW | 5 |
| 2022 | Hypothesis Transfer in Bandits by Weighted Models
Steven Bilaj, Sofien Dhouib, Setareh Maghsudi |
ECML/PKDD (4) | 3 |
| 2022 | Joint Coverage and Resource Allocation for Federated Learning in UAV-Enabled NetworksabstractThanks to its communication efficiency and low latency, federated learning (FL) has emerged as a promising learning paradigm in the unmanned aerial vehicle (UAV)-enabled networks; nevertheless, the great potential of FL in UAV networks is realizable only upon optimizing crucial factors such as coverage and transmission delay. In this paper, we study the problem of joint coverage optimization and efficient radio resource allocation. The objective is to minimize the convergence time of FL in a UAV-enabled network, where UAVs perform learning over an inhomogeneous sensor network. To this end, we develop a method that minimizes the FL computation and communication time in each global iteration: First, the algorithm adjusts the UAVs’ locations to control the average number of sensors associated with each UAV to maximize the coverage and to reduce the overall computation time. The UAVs’ locations also affect their transmission delay. Thus, in the second step, the method uses a fair resource allocation scheme for channel allocation and power control to minimize the FL communication time while retaining the efficiency of resource expenditure. Mariam Yahya, Setareh Maghsudi |
WCNC | 2 |
| 2022 | Computation Offloading in Heterogeneous Vehicular Edge Networks: On-Line and Off-Policy Bandit SolutionsabstractWith the rapid advancement of intelligent transportation systems (ITS) and vehicular communications, vehicular edge computing (VEC) is emerging as a promising technology to support low-latency ITS applications and services. In this paper, we consider the computation offloading problem from mobile vehicles/users in a heterogeneous VEC scenario, and focus on the network- and base station selection problems, where different networks have different traffic loads. In a fast-varying vehicular environment, computation offloading experience of users is strongly affected by the latency due to the congestion at the edge computing servers co-located with the base stations. However, as a result of the non-stationary property of such an environment and also information shortage, predicting this congestion is an involved task. To address this challenge, we propose an on-line learning algorithm and an off-policy learning algorithm based on multi-armed bandit theory. To dynamically select the least congested network in a piece-wise stationary environment, these algorithms predict the latency that the offloaded tasks experience using the offloading history. In addition, to minimize the task loss due to the mobility of the vehicles, we develop a method for base station selection. Moreover, we propose a relaying mechanism for the selected network, which operates based on the sojourn time of the vehicles. Through intensive numerical analysis, we demonstrate that the proposed learning-based solutions adapt to the traffic changes of the network by selecting the least congested network, thereby reducing the latency of offloaded tasks. Moreover, we demonstrate that the proposed joint base station selection and the relaying mechanism minimize the task loss in a vehicular environment. Arash Bozorgchenani, Setareh Maghsudi, Daniele Tarchi, Ekram Hossain 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2021 | Mobility-Aware Routing and Caching: A Federated Learning Assisted ApproachabstractWe develop mobility-aware routing and caching strategies to solve the network cost minimization problem for dense small-cell networks. The challenge mainly stems from the insufficient backhaul capacity of small-cell networks and the limited storing capacity of small-cell base stations (SBSs). The optimization problem is NP-hard since both the mobility patterns of the mobilized users (MUs), as well as the MUs’ preference for contents, are unknown. To tackle this problem, we start by dividing the entire geographical area into small sections, each of which containing one SBS and several MUs. Based on the concept of one-stop-shop (OSS), we propose a federated routing and popularity learning (FRPL) approach in which the SBSs cooperatively learn the routing and preference of their respective MUs, and make caching decision. Notably, FRPL enables the completion of the multi-tasks in one shot, thereby reducing the average processing time per global aggregation.1Theoretical and numerical analyses show the effectiveness of our proposed approach. Setareh Maghsudi, Tomoaki Ohtsuki |
ICC | 2 |
| 2021 | EdgeDASH: Exploiting Network-Assisted Adaptive Video Streaming for Edge Caching
Suzan Bayhan, Setareh Maghsudi, Anatolij Zubow |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2020 | A Non-Stationary Bandit-Learning Approach to Energy-Efficient Femto-Caching With Rateless-Coded TransmissionabstractThe ever-increasing demand for media streaming together with limited backhaul capacity renders developing efficient file-delivery methods imperative. One such method is femto-caching, which, despite its great potential, imposes several challenges such as efficient resource management. We study a resource allocation problem for joint caching and transmission in small cell networks, where the system operates in two consecutive phases: (i) cache placement, and (ii) joint file- and transmit power selection followed by broadcasting. We define the utility of every small base station in terms of the number of successful reconstructions per unit of transmission power. We then formulate the problem as to select a file from the cache together with a transmission power level for every broadcast round so that the accumulated utility over the horizon is maximized. The former problem boils down to a stochastic knapsack problem, and we cast the latter as a multi-armed bandit problem. We develop a solution to each problem and provide theoretical and numerical evaluations. In contrast to the state-of-the-art research, the proposed approach is especially suitable for networks with time-variant statistical properties. Moreover, it is applicable and operates well even when no initial information about the statistical characteristics of the random parameters such as file popularity and channel quality is available. Setareh Maghsudi, Mihaela van der Schaar |
IEEE Trans. Wirel. Commun. | 1 |
| 2019 | A Reverse Auction Model for Efficient Resource Allocation in Mobile Edge Computation OffloadingabstractMobile edge computing (MEC) enables mobile users to offload their computationally-intensive tasks to the servers located at the network's edge. One of the fundamental challenges of MEC is to develop methods to efficiently allocate the limited computational resources of the edge servers to the offloading users. To address this challenge, in this paper, we propose a reverse auction framework based on position auction consisting of pricing, bidding strategy optimization, and winner determination. The proposed solution allows the edge servers to maximize their utility through strategic participation. Moreover, it ensures users' satisfaction by taking the users' preferences into account. The solution has polynomial-time complexity and enjoys desirable economical characteristics including envy-free and individual rationality. In addition to the theoretical analysis, numerical results establish the sound performance of the proposed framework in terms of the system's resource utilization as well as the users' satisfaction level. Ummy Habiba, Setareh Maghsudi, Ekram Hossain 0001 |
GLOBECOM | 2 |
| 2019 | A Bandit Learning Approach to Energy-Efficient Femto-Caching under UncertaintyabstractWe address a resource allocation problem for joint caching and broadcast transmission in small cell networks with time-varying statistical properties. Each small base station (SBS) selects some files to store in its capacity-limited cache, given no prior information about the random and dynamic parameters such as file popularity, channel quality, and network traffic. Moreover, at consecutive rounds, a file is selected from the cache to broadcast. We define the utility of the SBS in terms of the number of successful file receptions per power consumption. The problem is formulated as to place the cache, and afterward select a file from the cache together with a transmission power for every broadcast round. The goal is to maximize the accumulated utility over the horizon. Therefore, we decompose the initial problem into two sub- problems: (i) cache placement, and (ii) joint file- and transmit power selection. The former problem boils down to a stochastic knapsack problem with stationary items' value, whereas the latter is cast as a multi-armed bandit problem with mortal arms. We develop a solution to each problem and evaluate the proposed solutions by theoretical and numerical analysis. Setareh Maghsudi, Mihaela van der Schaar |
GLOBECOM | 1 |
| 2018 | Distributed Task Management in Cyber-Physical Systems: How to Cooperate Under Uncertainty?abstractWe consider the problem of task allocation in a network of cyber-physical systems (CPSs). The network can have different states, and the tasks are of different types. The task arrival is stochastic and state-dependent. Every CPS is capable of performing each type of task with some specific state-dependent efficiency. The CPSs have to agree on task allocation prior to knowing about the realized network's state and/or the arrived tasks. We model the problem as a multistate stochastic cooperative game with state uncertainty. We then use the concept of deterministic equivalence and sequential core to solve the problem. We establish the non-emptiness of the strong sequential core in our designed task allocation game and investigate its characteristics including uniqueness and optimality. Moreover, we prove that in the task allocation game, the strong sequential core is equivalent to Walrasian equilibrium under state uncertainty; consequently, it can be implemented by using the Walras' tatonnement process. Setareh Maghsudi, Mihaela van der Schaar |
GLOBECOM | 1 |
| 2018 | Cheat-Proof Distributed Power Control in Full-Duplex Small Cell Networks: A Repeated Game With Imperfect Public MonitoringabstractWe address the problem of distributed power control in a two-tier cellular network, where full-duplex small cells underlay a macro cell in a co-channel deployment scenario. We first formulate the distributed power control problem as a non-cooperative game and then extend it to a repeated game with imperfect public monitoring. The repeated game formulation prevents deceitful small cells from deviating from the social optimal solution for their own benefit. We establish the existence and uniqueness of the Nash equilibrium in the formulated non-cooperative game. We also characterize the set of public perfect equilibrium for the repeated game. A two-phase distributed algorithm is proposed to achieve and enforce a Pareto optimal transmit power profile. The solution obtained by this algorithm is also social optimal. Phase 1 of the algorithm is a fully distributed learning phase based on perturbed Markov chains, where each base station individually learns a Pareto optimal operating point. Phase 2 is composed of two rules: 1) a detection rule based on Page-Hinckley test to detect cheating and 2) a punishment rule to motivate cheating base stations to cooperate. Through theoretical analysis, we prove that the proposed distributed power control mechanism achieves a public perfect equilibrium point of the formulated repeated game. The power control algorithm is also cheat-proof and needs only a small amount of information exchange among network nodes. The effectiveness of the algorithm is shown through numerical analysis. Our proposed model, algorithm, and analysis are also valid for a half-duplex system as a special case. Prabodini Semasinghe, Ekram Hossain 0001, Setareh Maghsudi |
IEEE Trans. Commun. | 3 |
| 2017 | Distributed User Association in Energy Harvesting Small Cell Networks: A Probabilistic Bandit ModelabstractWe investigate a distributed downlink user association problem in a dynamic small cell network, where every small base station (SBS) obtains its required energy through ambient energy harvesting. On the one hand, energy harvesting is inherently opportunistic, so that the amount of available energy is a random variable. On the other hand, users arrive at random and require different wireless services, rendering the energy consumption a random variable. In this paper, we develop a probabilistic framework to mathematically model and analyze the random behavior of energy harvesting and energy consumption. We further analyze the probability of QoS satisfaction (success probability), for each user with respect to every SBS. The proposed user association scheme is distributed in the sense that every user independently selects its corresponding SBS with the success probability serving as the performance metric. The success probability however depends on a variety of random factors such as energy harvesting, channel quality, and network traffic, whose distribution or statistical characteristics might not be known at users. Since acquiring the knowledge of these random variables (even statistical) is very costly in a dense network, we develop a bandit-theoretical formulation for distributed SBS selection when no prior information is available at users. The performance is analyzed both theoretically and numerically. Setareh Maghsudi, Ekram Hossain 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2016 | Distributed downlink user association in small cell networks with energy harvestingabstractWe consider a user assignment problem in small cell networks, where small cells obtain the required energy through ambient energy harvesting. We model the network as a competitive market with uncertainty, where small cells, represented as consumers, are willing to maximize their utility scores by selecting users, represented as commodities. Small cells are uncertain about the amount of harvested energy, formulated as natures' state. The solution is the general equilibrium under uncertainty, also called Arrow-Debreu equilibrium. We show that in our setting such equilibrium exists, and is Pareto optimal in terms of expected aggregate network utility. Besides, we use the Walras' tatonnement process to implement equilibrium efficiently. Setareh Maghsudi, Ekram Hossain 0001 |
ICC | 1 |
| 2015 | Joint channel allocation and power control for underlay D2D transmissionabstractWe study a joint channel allocation and power control problem for device-to-device (D2D) transmission underlaying a conventional single-cell cellular network. In such networks, direct transmissions are allowed among device pairs with local needs, provided that the adverse effects of D2D communications on cellular users is negligible and cellular users are given the priority in using limited wireless resources. Moreover, as D2D users are not in contact with the base station (BS), providing them with channel and/or network knowledge imposes excessive overhead. As a result, it becomes imperative to seek for new resource management mechanisms that fit the limitations of this concept. In this paper we consider a realistic model with respect to the information availability, and propose a joint channel allocation and power control scheme by using game- and graph theory. In particular, we first decompose the resource management problem into two cascaded channel allocation and power control problems, by proving a lower bound on the aggregate utility of cellular users. Afterwards we propose a centralized graph-theoretical channel allocation approach jointly for D2D and cellular users. Given the channel allocation, the subsequent power control problem is modeled as a game with incomplete information. We analyze the characteristics of this game and solve it in a distributed manner, by using a multi-agent Q-learning strategy. We evaluate the proposed resource allocation scheme both analytically and numerically. Setareh Maghsudi, Slawomir Stanczak |
ICC | 1 |
| 2015 | Channel Selection for Network-Assisted D2D Communication via No-Regret Bandit Learning With Calibrated ForecastingabstractWe consider the distributed channel selection problem in the context of device-to-device (D2D) communication as an underlay to a cellular network. Underlaid D2D users communicate directly by utilizing the cellular spectrum, but their decisions are not governed by any centralized controller. Selfish D2D users that compete for access to the resources form a distributed system where the transmission performance depends on channel availability and quality. This information, however, is difficult to acquire. Moreover, the adverse effects of D2D users on cellular transmissions should be minimized. In order to overcome these limitations, we propose a network-assisted distributed channel selection approach in which D2D users are only allowed to use vacant cellular channels. This scenario is modeled as a multi-player multi-armed bandit game with side information, for which a distributed algorithmic solution is proposed. The solution is a combination of no-regret learning and calibrated forecasting, and can be applied to a broad class of multi-player stochastic learning problems, in addition to the formulated channel selection problem. Theoretical analysis shows that the proposed approach not only yields vanishing regret in comparison to the global optimal solution but also guarantees that the empirical joint frequencies of the game converge to the set of correlated equilibria. Setareh Maghsudi, Slawomir Stanczak |
IEEE Trans. Wirel. Commun. | 1 |
| 2014 | Transmission mode selection for network-assisted device to device communication: A Levy-bandit approachabstractThis paper studies device-to-device (D2D) communication underlaying cellular infrastructure, where each device pair is provided with two transmission modes: indirect and direct. Indirect transmission is a two-hop interference-free transmission via a base station. Despite being interference-free, this transmission type might be inefficient in communications scenarios where short-distance connections can be established. Moreover, the need for centralized resource allocation and utilizing extra hardware may lead to excessive complexity and unacceptable costs. In such scenarios, direct transmissions can utilize the proximity- and hop gains to achieve higher rates and lower end-to-end latencies. While having a potential for huge performance gains, direct D2D communications poses some fundamental challenges resulting from the absence of a devoted controller such as uncoordinated interference and unavailability of permanent direct channels. Roughly speaking, in an average sense, while indirect transmission pays safe and steady reward, direct transmission is risky, yielding a stochastic reward which might be lower than the guaranteed reward of indirect transmission, despite the proximity-and hop gains. Transmitters should therefore choose the most efficient transmission mode in the presence of limited information. This paper characterizes the reward process for each transmission mode to model the mode selection problem as a two-armed Levy-bandit game. Accordingly, the reward of the risky arm (direct mode) is considered to be a pure-jump Levy process, following compound Poisson distribution. Mathematical results from bandit and learning theories are used to solve the selection problem. Numerical results complete the paper. Setareh Maghsudi, Slawomir Stanczak |
ICASSP | 1 |
| 2013 | Dynamic bandit with covariates: Strategic solutions with application to wireless resource allocationabstractMulti-armed bandit (MAB) problems form a class of sequential optimization problems, in which a player sequentially pulls an arm, selected from a known and finite set of arms, in order to achieve an initially unknown reward. The player aims at maximizing the accumulated reward over a predefined game horizon. Clearly, in bandit setting, a dilemma appears between pushing the currently most promising arm, i.e. the arm with the highest empirical mean reward, on the one hand (exploitation) and on the other sampling arms in order to improve the estimation of the reward generating processes of arms (exploration). In this paper we study a specific subset of MAB problems, namely stochastic covariate bandits, where it is assumed that the series of instantaneous rewards generated by each arm can be attributed to a specific distribution, and that some side information (covariate) is revealed to the player at the beginning of each game trial. In this setting, we address the exploitation-exploration dilemma by proposing two strategies for arm selection (allocation rule). Provided that the underlying regression process is trust-worthy, the proposed strategies are strongly consistent, in the sense that the accumulated reward is equivalent to that based on the best arm, asymptotically almost surely. Further, it is illustrated that the covariate bandit model and our allocation strategies are applicable to wireless networking scenarios by considering the relay selection problem as case study. Setareh Maghsudi, Slawomir Stanczak |
ICC | 1 |
| 2013 | Relay selection with no side information: An adversarial bandit approachabstractMulti-armed bandit games form a class of sequential optimization problems, in which a player sequentially pulls an arm, selected from a known and finite set of arms, in order to receive an a priori unknown reward. Since the player does not know the arm with the highest reward in advance, it utilizes a well-designed selection strategy to minimize the so-called regret, which, roughly speaking, results from the lack of this information. This paper studies cooperative transmission in a dense mobile network, where users compete for utilizing a number of relays to improve the quality of transmissions. Under the assumption of no side information available to the users, the relay selection and assignment problem is formulated as an adversarial multi-player multi-armed bandit game. Based on this formulation, a selection strategy is proposed that is shown to guarantee the convergence of the empirical frequencies of the game to a correlated equilibrium. Moreover, applying the experimental regret testing protocol shows that the empirical frequencies of the relay selection game converges to Nash equilibrium. Finally, experimental evaluations are carried out to compare the performance of various selection strategies and with it to demonstrate the effectiveness of the proposed approach. The proposed game model and selection strategies can be used in a wide range of wireless networking scenarios, such as spectrum pulling in cognitive radio networks and base station assignment in cellular networks. Setareh Maghsudi, Slawomir Stanczak |
WCNC | 1 |
| 2012 | A delay-constrained rateless coded incremental relaying protocol for two-hop transmissionabstractWe develop an efficient relay selection protocol for two-hop transmission in cognitive radio networks where secondary users interfere with a primary user. Our protocol provides joint benefits of two well-known relaying strategies, namely relay sub-set selection and incremental relaying. It aims at minimizing the outage probability for delay-constrained applications by utilizing rateless codes. The advantages of the proposed scheme compared to conventional relay selection protocols can be summarized as follows: (1) feedback overhead is significantly reduced, (2) resources are exploited efficiently, (3) outage probability is minimized. In order to evaluate the proposed strategy, analytical and numerical results are presented. Setareh Maghsudi, Slawomir Stanczak |
WCNC | 1 |