VLDB 2026 Research / reviewers in the wild / expert
Andrea Ortiz
dblp:163/8755
· DBLP profile ↗
30ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-6544-1152ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 26 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Joint Age of Information and Energy Minimization in Two-hop Status Update Systems
Yuhui Sun, Andrea Ortiz, Chao Wang 0015 |
ICC | 2 |
| 2026 | Risk-Aware Learning for Digital Twin Placement in Non-Stationary Hierarchical Edge ComputingabstractA Digital Twin (DT) is a virtual replica of a physical system (PS), such as an Internet-of-Things (IoT) device, enabling the analysis and prediction of the PS's dynamics. To maintain accuracy, a DT frequently synchronizes with its PS, by processing status updates sent from the PS. Therefore, the incurred synchronization latency is impacted by both, communication and computation resources. Due to being computationally demanding, DTs are commonly hosted on edge servers (ESs) in Mobile Edge Computing (MEC). To increase scalability and flexibility of MEC systems, hierarchical architectures distribute computation resources across multiple tiers of ESs between base stations and cloud. Hierarchical MEC gives rise to the DT placement problem, i.e., the selection of an ES for hosting the DT. The selection must consider the synchronization latency, which should remain below critical thresholds to ensure DT accuracy and thus, successful synchronization. DT placement is challenging, because the synchronization latency is impacted by the dynamic and statistically non-stationary MEC system characteristics, such as communication channels and server loads. Furthermore, additional overhead is caused by migrating DTs between ESs. To address these challenges, we propose a novel risk-aware Multi-Armed Bandit algorithm that adapts to piece-wise stationary environments. Our approach minimizes synchronization cost, defined as a weighted sum of synchronization latency and energy consumption, while reducing the risk of synchronization failures and limiting DT migration overhead. Simulation results demonstrate that our proposed algorithm achieves a similar synchronization cost, compared to state-of-the-art baselines, while achieving a significantly smaller synchronization failure rate and migration overhead. Maximilian Wirth, Andrea Ortiz, Anja Klein 0002 |
WCNC | 2 |
| 2026 | Federated Reinforcement Learning for Efficient Mobile Crowdsensing Under Incomplete InformationabstractMobile crowdsensing (MCS) is a distributed sensing architecture that utilizes existing sensors on mobile units (MUs) to perform sensing tasks. A mobile crowdsensing platform (MCSP) publishes the sensing tasks and the MUs decide if they want to participate in their execution in exchange for money. The MCS system is characterized by its dynamic nature in which the task requirements, the MUs’ availability, and their available resources change over time. The MUs aim to find an efficient task participation strategy to maximize their income while the MCSP focuses on maximizing the number of completed tasks. As optimal task participation strategies require perfect non-causal information about the MCS system, which is unavailable in realistic scenarios, the main challenge in MCS is to find an efficient task participation strategy for the MUs under incomplete information. To this aim, a novel fully decentralized federated deep reinforcement learning algorithm, termed FDRL-PPO is proposed. FDRL-PPO enables every MU to learn its own task participation strategy based on its experiences, available resources, and preferences, without relying on perfect non-causal information about the MCS system. To replenish their batteries, the MUs rely on energy harvesting. As a result, their available energy varies over time, leading to varying availability and fragmented learning experiences. To mitigate these challenges, the proposed approach leverages federated learning, enabling MUs to collaboratively improve their models without having to share private raw data like their own experiences. By exchanging only learned models, MUs collectively compensate for individual limitations, and find more scalable, robust, and efficient task participation strategies. Comprehensive evaluations on both synthetic and real-world datasets show that FDRL-PPO consistently outperforms benchmark algorithms in terms of task completion ratio, fairness in task completion, energy consumption, and number of conflicting proposals. Sumedh J. Dongare, Patrick Weber 0001, Andrea Ortiz, Walid Saad 0001, Oliver Hinz, Anja Klein 0002 |
IEEE Internet Things J. | 3 |
| 2026 | Deep Sleep Scheduling for Satellite IoT via Simulation-Based OptimizationabstractThe Satellite Internet of Things (S-IoT) enables global connectivity for remote sensing devices that must operate energy-efficiently over long time spans. We consider an S-IoT system consisting of a sender-receiver pair connected by a data channel and a feedback channel and capture its dynamics using a Markov Decision Process (MDP). To extend battery life, the sender has to decide on deep-sleep durations. Deep-sleep scheduling is the primary lever to reduce energy consumption, since sleeping devices consume only a fraction of their idle power. By choosing its deep-sleep duration online, the sender has to find a trade-off between energy consumption and data quality degradation at the receiver, captured by a weighted sum of costs. We quantify data quality degradation via the recently introduced Goal-Oriented Tensor (GoT) metric, which can take both age and content of delivered data into account. We assume a Markovian observed process and Markov channels with time-varying delay and erasure rates. The challenge is that content awareness of the GoT metric makes periodic transmissions inherently inefficient. Additionally, optimal sleep durations depends on the (unknown) future states of the observed process and the channels, both of which must be inferred online. We propose a novel algorithm using probabilistic simulation-based optimization (PSBO). With PSBO, the sensor forecasts future states based on estimated transition probabilities, and uses these forecasts to select the optimal deep-sleep duration. Extensive simulations demonstrate the strong performance of PSBO across diverse conditions. In S-IoT hardware experiments, PSBO reduces costs by 59% versus a threshold-based solution and by 89% versus Q-learning. Wanja de Sombre, Monika Tomová, Marek Galinski, Anja Klein 0002, Andrea Ortiz |
IEEE Internet Things J. | 5 |
| 2026 | SkyLink: Scalable and Resilient Link Management in LEO Satellite NetworksabstractThe rapid growth of space-based services has established Low Earth Orbit (LEO) satellite networks as a promising option for global broadband connectivity. Next-generation LEO networks leverage inter-satellite links (ISLs) to provide faster and more reliable communications compared to traditional bent-pipe architectures, even in remote regions. However, the high mobility of satellites, dynamic traffic patterns, and potential link failures pose significant challenges for efficient and resilient routing. To address these challenges, we model the LEO satellite network as a time-varying graph comprising a constellation of satellites and ground stations. Our objective is to minimize a weighted sum of average delay and packet drop rate. Each satellite independently decides how to distribute its incoming traffic to neighboring nodes in real time. Given the infeasibility of finding optimal solutions at scale, due to the exponential growth of routing options and uncertainties in link capacities, we propose SKYLINK, a novel fully distributed learning strategy for link management in LEO satellite networks. SKYLINK enables each satellite to adapt to the time-varying network conditions, ensuring real-time responsiveness, scalability to millions of users, and resilience to network failures, while maintaining low communication overhead and computational complexity. To support the evaluation of SKYLINK at global scale, we develop a new simulator for large-scale LEO satellite networks. For 25.4 million users, SKYLINK reduces the weighted sum of average delay and drop rate by 29% compared to the bent-pipe approach, and by 92% compared to Dijkstra. It lowers drop rates by 95% relative to k-shortest paths, 99% relative to Dijkstra, and 74% compared to the bent-pipe baseline, while achieving up to 46% higher throughput. At the same time, SKYLINK maintains constant computational complexity with respect to constellation size. Wanja de Sombre, Arash Asadi, Debopam Bhattacherjee, Deepak Vasisht, Andrea Ortiz |
IEEE Trans. Commun. | 5 |
| 2025 | Momentum Survey Propagation: A Statistical Physics Approach to Resource Allocation in mMTCabstractThe importance of massive machine-type communications (mMTCs) in Beyond 5G and 6G networks is supported by the ever-increasing number of connected devices in what are known as massive Internet of Things (IoT) networks. These networks bring unprecedented challenges for the distribution of the available communication resources because the allocation problems often lead to combinatorial optimization formulations which are known to be NP-hard. A fact that limits the performance of state-of-the-art techniques when the network size increases. To address this challenge, we take a new direction and propose a method based on statistical physics to address resource allocation problems in large networks. To this aim, we first show that resource allocation problems have the same structure as the problem of finding specific configurations in spin glasses, a type of disordered physical systems. Based on this parallel, we propose Momentum Survey Propagation, a resource allocation method to minimize the interference in mMTC networks. Our proposed approach extends the Survey Propagation method of statistical physics. Specifically, it exploits the so-called momentum technique, widely used in the context of neural networks, to improve the convergence properties of Survey Propagation. Our implementation is the first application of Survey Propagation to a wireless communication network. Through numerical simulations we show that Momentum Survey Propagation is a promising tool for the efficient allocation of communication resources in mMTC. Andrea Ortiz, Rostyslav Olshevskyi, Daniel Barragan-Yani |
IEEE Internet Things J. | 1 |
| 2024 | Minimizing the Age of Incorrect Information for Status Update Systems with Energy HarvestingabstractStatus Update Systems (SUSs) are central components in applications like environmental sensing or smart cities. They consist of a sender monitoring a remote process and sending the sensed information to a receiver. The sender aims to deliver fresh information about the monitored process's state to allow the receiver to timely respond to the process's changes. In SUSs, the sender is usually battery operated. Therefore, to increase the available energy we consider Energy Harvesting (EH). Moreover, as at the receiver the information transmitted by the sender is only relevant when the process's state changes, we measure the information's freshness using Age of Incorrect Information (AoII). Finding the optimal transmission strategy at the sender that minimizes the AoII requires perfect system knowledge, i.e., the behavior of the monitored process, the channel quality, and the available energy. However, in real applications this knowledge is usually not available. To overcome this challenge, we first establish the optimality of threshold-based policies for AoII minimization in SUSs with EH capabilities by proving that there exists an AoII value depending on the observed state of the monitored process, the battery level and the receiver's estimation of the monitored process's state beyond which transmitting is preferable over idling. Next, we exploit the threshold-based policies' structure and deploy a learning algorithm based on Finite-Difference Policy Gradient (FDPG). Our proposed approach finds the AoII thresholds without requiring perfect system knowledge. Simulations show that our approach outperforms reference algorithms by at least 20% and efficiently learns near-optimal policies for AoII minimization. Sumedh J. Dongare, Aleksandar Jovovic, Wanja de Sombre, Andrea Ortiz, Anja Klein 0002 |
ICC | 4 |
| 2024 | Two-Sided Learning: A Techno-Economic View of Mobile Crowdsensing Under Incomplete InformationabstractIn Mobile Crowdsensing (MCS) a mobile crowd-sensing platform (MCSP) collects sensing data from mobile units (MUs) in exchange for payment. The MCSP broadcasts a list of available sensing tasks. Based on this list, each MU solves a task proposal problem to decide which task it is willing to perform and sends a proposal to the MCSP. Based on the MUs' proposals, the MCSP solves a task assignment problem. There are two challenges when finding efficient task proposal strategies for the MUs and an efficient task assignment strategy for the MCSP (i) The techno-economic perspective of MCS: From the technical perspective, MCS should maximize the data quality while minimizing time and energy consumption. From the economic perspective, there are two sides, the MUs and the MCSP which act as selfish decision-makers, who aim at maximizing their own income. (ii) Incomplete information at two sides: Initially, the MCSP does not know the expected data quality and the MUs do not know the expected effort required for task completion. To overcome these challenges, we propose a novel Two-Sided Learning (TSL) approach. At the MU side, TSL is based on an innovative gradient-based multi-armed bandit solution to maximize the MUs' utility under incomplete information about the strategies of other MUs. At the MCSP side, a learning strategy is used to find the task assignment strategy that maximizes its utility. Simulation results show that TSL achieves near-optimal social welfare, which is the sum of MUs' and MCSP's utilities, and a near-optimal energy efficiency. Sumedh J. Dongare, Bernd Simon, Andrea Ortiz, Anja Klein 0002 |
ICC | 3 |
| 2024 | Age of Information Minimization in Status Update Systems with Imperfect Feedback ChannelabstractStatus Update System (SUS) are monitoring applications of Internet of Things (IoT). They are formed by a sender that monitors a remote process and sends status updates to a receiver over a wireless channel. For successful monitoring, the sender must keep the status updates at the receiver fresh. This freshness is generally measured using the Age of Information (AoI) metric. The aim of the sender is to find a monitoring and transmission strategy that minimizes the AoI. To find the optimal strategy, the sender needs to accurately track the AoI at the receiver, i.e., it needs to perfectly know whether a transmitted status update is correctly received or not. This knowledge can be achieved by using a feedback channel between receiver and sender to send acknowledge (ACK) or negative acknowledge (NACK) messages. However, in real applications, the feedback channel is not perfect, and the transmission of ACK/NACK messages might fail. This means, the monitoring and transmission decisions have to be made under uncertainty about the receiver's AoI. To overcome this challenge, we introduce the concept of a socalled belief distribution and propose a joint monitoring and transmission strategy at the sender based on reinforcement learning. Our approach, termed Belief Learning, exploits the belief distribution to minimize the AoI at the receiver. Through numerical simulations we show that Belief Learning enables the sender to achieve near-optimal performance with respect to the perfect feedback channel case. Friedrich Pyttel, Wanja de Sombre, Andrea Ortiz, Anja Klein 0002 |
ICC | 3 |
| 2024 | Decentralized Online Learning in Task Assignment Games for Mobile CrowdsensingabstractThe problem of coordinated data collection is studied for a mobile crowdsensing (MCS) system. A mobile crowdsensing platform (MCSP) sequentially publishes sensing tasks to the available mobile units (MUs) that signal their willingness to participate in a task by sending sensing offers back to the MCSP. From the received offers, the MCSP decides the task assignment. A stable task assignment must address two challenges: the MCSP’s and MUs’ conflicting goals, and the uncertainty about the MUs’ required efforts and preferences. To overcome these challenges a novel decentralized approach combining matching theory and online learning, called collision-avoidance multi-armed bandit with strategic free sensing (CA-MAB-SFS), is proposed. The task assignment problem is modeled as a matching game considering the MCSP’s and MUs’ individual goals while the MUs learn their efforts online. Our innovative “free-sensing” mechanism significantly improves the MU’s learning process while reducing collisions during task allocation. The stable regret of CA-MAB-SFS, i.e., the loss of learning, is analytically shown to be bounded by a sublinear function, ensuring the convergence to a stable optimal solution. Simulation results show that CA-MAB-SFS increases the MUs’ and the MCSP’s satisfaction compared to state-of-the-art methods while reducing the average task completion time by at least 16%. Bernd Simon, Andrea Ortiz, Walid Saad 0001, Anja Klein 0002 |
IEEE Trans. Commun. | 2 |
| 2024 | Risk-Averse Learning for Reliable mmWave Self-BackhaulingabstractWireless backhauling at millimeter-wave frequencies (mmWave) in static scenarios is a well-established practice in cellular networks. However, highly directional and adaptive beamforming in today’s mmWave systems have opened new possibilities for self-backhauling. Tapping into this potential, 3GPP has standardized Integrated Access and Backhaul (IAB) allowing the same base station to serve both access and backhaul traffic. Although much more cost-effective and flexible, resource allocation and path selection in IAB mmWave networks is a formidable task. To date, prior works have addressed this challenge through a plethora of classic optimization and learning methods, generally optimizing Key Performance Indicators (KPIs) such as throughput, latency, and fairness, and little attention has been paid to the reliability of the KPI. We propose Safehaul, a risk-averse learning-based solution for IAB mmWave networks. In addition to optimizing the average performance, Safehaul ensures reliability by minimizing the losses in the tail of the performance distribution. We develop a novel simulator and show via extensive simulations that Safehaul not only reduces the latency by up to 43.2% compared to the benchmarks, but also exhibits significantly more reliable performance, e.g., 71.4% less variance in latency. Amir Ashtari Gargari, Andrea Ortiz, Matteo Pagin, Wanja de Sombre, Michele Zorzi, Arash Asadi |
IEEE/ACM Trans. Netw. | 2 |
| 2023 | Federated Deep Reinforcement Learning for Task Participation in Mobile CrowdsensingabstractMobile Crowdsensing (MCS) is a promising distributed sensing architecture that harnesses the power of sensors on mobile units (MUs) to perform sensing tasks. The MCS is a dynamic system in which the requirements of the sensing tasks, the MUs' conditions and the available resources change over time. The performance of an MCS system depends on the selection of the MUs participating in each sensing task. However, this is not a trivial problem. An optimal task participation strategy requires non-causal knowledge about the dynamic MCS system, a requirement that cannot be fulfilled in real implementations. Moreover, centralized optimization-based approaches do not scale with increasing number of participating MUs and often ignore the MUs' preferences. To overcome these challenges, in this paper we propose a novel multi-agent federated deep reinforcement learning algorithm (FDRL-PPO) which does not need this perfect non-causal knowledge, but instead, enables the MUs to learn their own task participation strategies based on their own conditions, available resources, and preferences. Through federated learning, the MUs share their learned strategies without disclosing sensitive information, enabling a robust and scalable task participation scheme. Numerical evaluations validate the effectiveness and efficiency of FDRL-PPO in comparison with reference schemes. Sumedh J. Dongare, Andrea Ortiz, Anja Klein 0002 |
GLOBECOM | 2 |
| 2023 | A Unified Approach to Learn Transmission Strategies Using Age-Based Metrics in Point-to-Point Wireless CommunicationabstractBased on the Age of Information as an optimization criterion, proposals for further age-based metrics have been made in recent years in the Internet of Things (IoT) domain. The research community's great interest in age-based metrics for point-to-point wireless communication has led to a multitude of different scenarios being investigated, including energy optimization, sensing, and risk-sensitivity. All these scenarios involve a sender-receiver pair and revolve around finding appropriate times for the sender to communicate status updates to the receiver. We propose a unified and modular framework that represents the aforementioned options in various combinations and enables transferring solutions developed for specific cases to a variety of scenarios. We generalize an existing optimization approach, which decides to transmit based on a threshold for the age-based metric, using this framework. We develop a unified and extended Q-learning-based algorithm with mechanisms to learn suitable solutions for all scenarios derived from our framework. These mechanisms accelerate the learning process and result in improved algorithmic performance compared to traditional Q-learning. Furthermore, we demonstrate the effectiveness of our solution in numerical simulations. Our unified solution outperforms several reference schemes in terms of age-based metrics, energy consumption, and risk. We present our findings as a starting point to investigate transmission strategies for more general settings with a more efficient approach. Wanja de Sombre, Felipe Marques, Friedrich Pyttel, Andrea Ortiz, Anja Klein 0002 |
GLOBECOM | 4 |
| 2023 | Contextual Multi-Armed Bandits for Non-Stationary Heterogeneous Mobile Edge ComputingabstractBase station (BS) selection for task offloading in Mobile Edge Computing (MEC) is a challenging problem due to the dynamic nature of MEC systems. The wireless channel as well as the load of BSs are stochastic quantities that can change in a statistically non-stationary fashion. Moreover, the computation capabilities of the BSs are heterogeneous. As the dynamic behaviour of a MEC system is, in practical scenarios, not known in advance, deciding where to offload has to be done under uncertainty about the MEC system and considering its non-stationary and heterogeneous characteristics. This paper investigates latency minimization in MEC with heterogeneous BSs. In order to meet low latency demands, a mobile unit (MU) has to quickly identify the best BS for offloading different computation tasks while facing uncertainty about the non-stationary system dynamics. To solve this problem, we propose a novel piece-wise stationary contextual Multi-Armed Bandit (MAB) algorithm that treats different task types as context and detects non-stationary changes in the BSs' performance. With the use of extensive simulations, we show that our proposed approach outperforms state-of-the-art algorithms, as it quickly adapts to changes in the MEC system and exhibits no penalty during stationary phases. Maximilian Wirth, Andrea Ortiz, Anja Klein 0002 |
GLOBECOM | 2 |
| 2023 | Safehaul: Risk-Averse Learning for Reliable mmWave Self-Backhauling in 6G NetworksabstractWireless backhauling at millimeter-wave frequencies (mmWave) in static scenarios is a well-established practice in cellular networks. However, highly directional and adaptive beamforming in today’s mmWave systems have opened new possibilities for self-backhauling. Tapping into this potential, 3GPP has standardized Integrated Access and Backhaul (IAB) allowing the same base station to serve both access and backhaul traffic. Although much more cost-effective and flexible, resource allocation and path selection in IAB mmWave networks is a formidable task. To date, prior works have addressed this challenge through a plethora of classic optimization and learning methods, generally optimizing a Key Performance Indicator (KPI) such as throughput, latency, and fairness, and little attention has been paid to the reliability of the KPI. We propose Safehaul, a risk-averse learning-based solution for IAB mmWave networks. In addition to optimizing average performance, Safehaul ensures reliability by minimizing the losses in the tail of the performance distribution. We develop a novel simulator and show via extensive simulations that Safehaul not only reduces the latency by up to 43.2% compared to the benchmarks, but also exhibits significantly more reliable performance, e.g., 71.4% less variance in achieved latency. Amir Ashtari Gargari, Andrea Ortiz, Matteo Pagin, Anja Klein 0002, Matthias Hollick, Michele Zorzi, Arash Asadi |
INFOCOM | 2 |
| 2023 | Demo:[SeBaSi] system-level Integrated Access and Backhaul simulator for self-backhaulingabstractmillimeter wave (mmWave) and sub-terahertz (THz) communications have the potential of increasing mobile network throughput drastically. However, the challenging propagation conditions experienced at mmWave and beyond frequencies can potentially limit the range of the wireless link down to a few meters, compared to up to kilometers for sub-6GHz links. Thus, increasing the density of base station deployments is required to achieve sufficient coverage in the Radio Access Network (RAN). To such end, 3rd Generation Partnership Project (3GPP) introduced wireless backhauled base stations with Integrated Access and Backhaul (IAB), a key technology to achieve dense networks while preventing the need for costly fiber deployments. In this paper, we introduce SeBaSi, a system-level simulator for IAB networks, and demonstrate its functionality by simulating IAB deployments in Manhattan, New York City and Padova. Finally, we show how SeBaSi can represent a useful tool for the performance evaluation of self-backhauled cellular networks, thanks to its high level of network abstraction, coupled with its open and customizable design, which allows users to extend it to support novel technologies such as Reconfigurable Intelligent Surfaces (RISs). Amir Ashtari Gargari, Matteo Pagin, Andrea Ortiz, Nairy Moghadas-Gholian, Michele Polese, Michele Zorzi |
WoWMoM | 3 |
| 2022 | Deep Reinforcement Learning for Task Allocation in Energy Harvesting Mobile CrowdsensingabstractMobile crowd-sensing (MCS) is an upcoming sensing architecture which provides better coverage, accuracy, and requires lower costs than traditional wireless sensor networks. It utilizes a collection of sensors, or crowd, to perform various sensing tasks. As the sensors are battery operated and require a mechanism to recharge them, we consider energy harvesting (EH) sensors to form a sustainable sensing architecture. The execution of the sensing tasks is controlled by the mobile crowd-sensing platform (MCSP) which makes task allocation decisions, i.e., it decides whether or not to perform a task depending on the available resources, and if the task is to be performed, assigns it to suitable sensors. To make optimal allocation decisions, the MCSP requires perfect non-causal knowledge regarding the channel coefficients of the wireless links to the sensors, the amounts of energy the sensors harvest and the sensing tasks to be performed. However, in practical scenarios this non-causal knowledge is not available at the MCSP. To overcome this problem, we propose a novel Deep-Q-Network solution to find the task allocation strategy that maximizes the number of completed tasks using only realistic causal knowledge of the battery statuses of the available sensors. Through numerical evaluations we show that our proposed approach performs only 7.8% lower than the optimal solution. Moreover, it outperforms the myopically optimal and the random task allocation schemes. Sumedh J. Dongare, Andrea Ortiz, Anja Klein 0002 |
GLOBECOM | 2 |
| 2022 | Delay- and Incentive-Aware Crowdsensing: A Stable Matching Approach for Coverage MaximizationabstractMobile crowdsensing (MCS) is a novel approach to increase the coverage, lower the costs, and increase the accuracy of sensing data. Its main idea is to collect sensor data using mobile units (MUs). The sensing is controlled by a mobile crowdsensing platform (MCSP) through the assignment of delay-sensitive sensing tasks to the MUs. Although promising, research effort in MCS is still needed to find task assignment solutions that maximize the coverage while considering the cost incurred by the MCSPs, the preferences of the MUs and the limited communication resources available. Specifically, we identify two main challenges: (i) A task assignment problem which incorporates the MCSP’s utility and the preferences of the MUs. (ii) An underlying communication resource allocation problem formulating the requirement of the timely transmission of sensing results given the limited communication resources. To address these challenges, we propose a novel two-stage matching algorithm. In the first stage, potential MU-task pairs are constructed considering the preferences of the MUs and the utility of the MCSP. In the second stage, the communication resource allocation is done based on potential MU-task pairs from the first stage. Through numerical simulations, we show that our proposed approach outperforms state-of-the-art methods in terms of the MCSP’s utility, coverage and MU’s satisfaction. Bernd Simon, Sumedh J. Dongare, Tobias Mahn, Andrea Ortiz, Anja Klein 0002 |
ICC | 4 |
| 2022 | Risk-Aware Multi-Armed Bandits for Vehicular CommunicationsabstractThe importance of vehicular communications has grown significantly in recent years. Potential use cases of vehicular communications are manifold and range from sharing information for driver assistance to entertainment purposes. This means that each connected vehicle has an individual data requirement for the communication infrastructure. However, due to the dynamic wireless environment, the simultaneous fulfillment of such requirements cannot be guaranteed. Therefore, novel solutions should not only consider the requirements of each user but also the risk of not being able to fulfill them. In this paper, we consider a vehicular communication scenario consisting of a base station that serves the vehicles in its coverage area using 5G millimeter wave (mmWave) narrow beams. The problem boils down to finding an optimal policy for the selection of the narrow beams. This should be done carefully, as the choice of the used beams greatly impacts the performance. For this purpose, we propose a risk-aware contextual Multi-Armed Bandit (MAB) online learning algorithm. Using this algorithm, the base station autonomously learns its environment and selects the best set of beams based on the vehicles located in its coverage area. In order to achieve a large risk awareness, this work focuses on two pillars. Firstly, the notion of risk is integrated in the proposed contextual MAB algorithm by exploiting the concepts of Mean-Variance and Conditional Value at Risk for the evaluation of the decisions made by the algorithm. Secondly, we introduce mechanisms that can detect non-stationarities and swiftly adapt to them in order to make the proposed approach robust against volatile environments that violate stationarity assumptions. By using extensive simulations, the effectiveness of the aforementioned approaches are proven numerically. Maximilian Wirth, Anja Klein 0002, Andrea Ortiz |
VTC Spring | 3 |
| 2021 | Energy-Optimal Short Packet Transmission for Time-Critical ControlabstractIn this paper, the transmission energy for reliable communications with short packets and low latency requirements, e.g. for control applications, is minimized. Since the dynamics of the agents determine the allowed latencies for receiving control inputs, the requirements on latency and allowable packet error rate are individual, depending on the machine type. We consider a centralized environment with a single controller transmitting control commands wireless to multiple agents with given latency requirements. Also, the channel conditions are individual for each agent. Therefore, the optimal time-frequency resource allocation is derived for continuous time-frequency resource allocation. Since the resource allocation in OFDM systems like 5G is discrete, an algorithm to select the allocation from a resource grid with different resolutions is proposed and shown to achieve solutions with less than 0.5 dB increase in energy consumption compared to the continuous results. With numerical evaluation, the benefit of a channel-state- and deadline-aware solution is shown for a resource grid based on the 5G frame structure. On average, the gain of the proposed algorithm to an allocation only balancing the number of resources for each agent, as far as the deadlines allow, is about 50% energy saving. Kilian Kiekenap, Andrea Ortiz, Anja Klein 0002 |
VTC Fall | 2 |
| 2020 | Delay Minimization for Edge Computing with Dynamic Server Computing Capacity: A Learning ApproachabstractThe offloading decisions of K mobile users (MUs) aiming at minimizing the execution delay in a Mobile Edge Computing (MEC) scenario with non-orthogonal multiple access is considered. In this work, we assume a time-varying MEC server computing capacity which exploits additional computing resources that are freed over time, but are not known beforehand. In this setting, the optimal offloading decision depends on the different tasks of MUs, their channel fading processes and the MEC server computing capacity. We first formulate the optimization problem and identify two main challenges, namely, how to exploit the given incomplete knowledge to minimize the delay and how to handle the high dimensionality of the problem. To address these challenges, we propose a novel reinforcement learning (RL) algorithm, termed combinatorial offloading learning (COL). The name stands for its ability to handle the combinatorial nature of the solutions. Exploiting the available knowledge, we learn the offloading decision policy aiming at minimizing the delay. Furthermore, we handle the curse of dimensionality, typical of combinatorial problems, by splitting the learning task, solving K+1 smaller RL problems and using linear function approximation. Through numerical simulations, we show that COL could perform similar to a short term optimal solution with complete information and exhaustive search, and outperforms known strategies like the greedy approach. Burak Yilmaz, Andrea Ortiz, Anja Klein 0002 |
GLOBECOM | 2 |
| 2019 | Optimal Resource Allocation Policy for Multi-Rate Opportunistic ForwardingabstractMany opportunistic routing protocols for wireless multi-hop networks rely on a fixed channel rate and a fixed priority order to manage the access to the channel by the involved nodes. Thereby, the actual channel capacities in the network are not considered and the diversity of links is not fully exploited. Furthermore, the data buffer of the nodes is not taken into account. In this work, we consider a wireless multihop scenario consisting of multiple cooperative nodes within each hop that share channel resources and adapt their channel rates based on local channel knowledge. A Markov Decision Process (MDP) model is used to derive an optimal resource allocation policy that minimizes the number of required time slots to forward all data packets to the next hop. Furthermore, we propose a state approximation technique that limits the required number of states, but captures the most important features of the problem. Simulation results demonstrate that the proposed policy achieves throughput gains of up to 25% compared to a fixed order transmission policy and up to 49% compared to a unipath approach. Fabian Hohmann, Andrea Ortiz, Anja Klein 0002 |
WCNC | 2 |
| 2019 | CBMoS: Combinatorial Bandit Learning for Mode Selection and Resource Allocation in D2D SystemsabstractThe complexity of the mode selection and resource allocation (MS&RA) problem has hampered the commercialization progress of Device-to-Device (D2D) communication in 5G networks. Furthermore, the combinatorial nature of MS&RA has forced the majority of existing proposals to focus on constrained scenarios or offline solutions to contain the size of the problem. Given the real-time constraints in actual deployments, a reduction in computational complexity is necessary. Adaptability is another key requirement for mobile networks that are exposed to constant changes such as channel quality fluctuations and mobility. In this article, we propose an online learning technique (i.e., CBMoS) which leverages combinatorial multi-armed bandits (CMAB) to tackle the combinatorial nature of MS&RA. Furthermore, our two-stage CMAB design results in a tight model, which eliminates the theoretically feasible but practicality invalid options from the solution space. We prototype the first SDR-based D2D testbed to verify the performance of CBMoS under real-world conditions. The simulations confirm that the fast learning speed of CBMoS leads to outperforming the benchmark schemes by up to 132%. In experiments, CBMoS exhibits even higher performance (up to 142%) than in the simulations. This stems from the adaptability/fast learning speed of CBMoS in presence of high channel dynamics which cannot be captured via statistical channel models used in the simulators. Andrea Ortiz, Arash Asadi, Max Engelhardt, Anja Klein 0002, Matthias Hollick |
IEEE J. Sel. Areas Commun. | 1 |
| 2019 | SCAROS: A Scalable and Robust Self-Backhauling Solution for Highly Dynamic Millimeter-Wave NetworksabstractMillimeter-wave (mmWave) backhauling is key to ultra-dense deployments in beyond-5G networks because providing every base station with a dedicated fiber-optic backhaul link to the core network is technically too complicated and economically too costly. Self-backhauling allows the operators to provide fiber connectivity only to a small subset of base stations (Fiber-BSs), whereas the rest of the base stations reach the core network via a (multi-hop) wireless link towards the Fiber-BS. Although a very attractive architecture, self-backhauling is proven to be an NP-hard route selection and resource allocation problem. The existing self-backhauling solutions lack practicality because:$(i)$they require solving a fairly complex combinatorial problem every time there is a change in the network (e.g., channel fluctuations), or$(ii)$they ignore the impact of network dynamics which are inherent to mobile networks. In this article, we propose SCAROS which is a semi-distributed learning algorithm that aims at minimizing the end-to-end latency as well as enhancing the robustness against network dynamics including load imbalance, channel variations, and link failures. We benchmark SCAROS against state-of-the-art approaches under a real-world deployment scenario in Manhattan and using realistic beam patterns obtained from off-the-shelf mmWave devices. The evaluation demonstrates that SCAROS achieves the lowest latency, at least$1.8\times $higher throughput, and the highest flexibility against variability or link failures in the system. Andrea Ortiz, Arash Asadi, Allyson Sim, Daniel Steinmetzer, Matthias Hollick |
IEEE J. Sel. Areas Commun. | 1 |
| 2018 | A Two-Layer Reinforcement Learning Solution for Energy Harvesting Data Dissemination ScenariosabstractA data dissemination scenario is considered. The transmitter harvests energy from the environment and uses it to transmit individual data to multiple receivers. We consider a realistic scenario in which only causal knowledge regarding the energy harvesting, the channel fading and the data arrival processes is available. Our goal is to find a power allocation policy aiming at maximizing the throughput. We propose a two-layer reinforcement learning algorithm which divides the learning task into two sub-tasks, namely, how much power to use in each time interval and how to split the power among the data to be transmitted. By dividing the task, we increase the learning speed as compared to the standard reinforcement learning algorithms Q-Iearning and SARSA. Moreover, the proposed algorithm outperforms reference policies that deplete the battery in every time interval. Andrea Ortiz, Anja Klein 0002 |
ICASSP | 1 |
| 2018 | Design and Implementation of a Novel Semi-Active Hybrid Unilateral Stance Control Knee Ankle Foot OrthosisabstractThis work presents the design and development of a semi-active hybrid orthotic system for support and facilitation of unilateral pathological human walking. The system is based on a novel lower limb orthosis with mechanical control and its combination with non-invasive muscle electrostimulation. The paper presents the concept design and realization of a novel Stance Control Knee Ankle Foot Orthoses (SCKAFO) and the design of Functional Electrical Stimulation (FES) strategies for artificial activation of ankle joint muscles in this hybrid scheme for gait support. In particular, we present the investigation of the effects of electrical stimulation patterns of ankle muscles synchronized in real-time with gait events., on the possible biomechanical alterations to gait in able-bodied individuals (n=8). Bilateral 3D ground reaction forces (GRF) were analyzed between barefoot overground walking with and without FES of ankle muscles. The observed effects of the tested FES strategy are coherent with a physiological gait strategy and compatible with the proposed SCKAFO. Juan J. Gil, Maria Carmen Sanchez-Villamañan, Jesús Gómez, Andrea Ortiz, José Luis Pons Rovira, Juan C. Moreno 0001, Antonio J. del Ama |
IROS | 4 |
| 2017 | DisVis 2.0: Decision Support for Rescue Missions Using Predictive Disaster Simulations with Human-Centric ModelsabstractIn disaster situations, the coordination of rescue missions is a difficult task. The person in charge makes decisions under pressure, which could lead to inappropriate instructions to rescuers and could cost many lives. The aim of DisVis 2.0 approach is to release pressure from those responsible by providing decision support using predictive human-centric disaster simulations for infrastructure-less emergency scenarios. More precisely, we enhance existing simulation frameworks only relying on ad-hoc network and mobility models by (1) human-centric models considering cognitive processes and behavior reactions, as well as (2) a novel automated decision support considering real-world sensor data. Christian Meurisch, The An Binh Nguyen, Martin Kromm, Andrea Ortiz, Ragnar Mogk, Max Mühlhäuser |
ICCCN | 4 |
| 2016 | A Learning Based Solution for Energy Harvesting Decode-and-Forward Two-Hop CommunicationsabstractEnergy harvesting (EH) two-hop communications are considered. The transmitter and the relay harvest energy from the environment and use it exclusively for transmitting data. A data arrival process is assumed at the transmitter. At the relay, a finite data buffer is used to store the received data. We consider a realistic scenario in which the EH nodes have only local causal information, i.e., at any time instant, each EH node only knows the current value of its EH process, channel state and data arrival process. Our goal is to find a power allocation policy to maximize the throughput at the receiver. We show that because the EH nodes have local causal information, the two-hop communication problem can be separated into two point-to-point problems. Consequently, independent power allocation problems are solved at each EH node. To find the power allocation policy, reinforcement learning with linear function approximation is applied. Moreover, to perform function approximation two feature functions which consider the data arrival process are introduced. Numerical results show that the proposed approach has only a small degradation as compared to the offline optimum case. Furthermore, we show that with the use of the proposed feature functions a better performance is achieved compared to standard approximation techniques. Andrea Ortiz, Hussein Al-Shatri, Xiang Li 0004, Anja Klein 0002 |
GLOBECOM | 1 |
| 2016 | Reinforcement learning for energy harvesting point-to-point communicationsabstractEnergy harvesting point-to-point communications are considered. The transmitter harvests energy from the environment and stores it in a finite battery. It is assumed that the transmitter has always data to transmit and the harvested energy is used exclusively for data transmission. As in practical scenarios prior knowledge about the energy harvesting process might not be available, we assume that at each time instant only information about the current state of the transmitter is available, i.e., harvested energy, battery level and channel coefficient. We model the scenario as a Markov decision process and we implement reinforcement learning at the transmitter to find a power allocation policy that aims at maximizing the throughput. To overcome the limitations of traditional reinforcement learning algorithms, we apply the concept of function approximation and we propose a set of binary functions to approximate the expected throughput given the state of the transmitter. Numerical results show that the performance of the proposed approach, which requires only causal knowledge of the energy harvesting process and channel coefficients, has only a small degradation compared to the optimum case which requires perfect non-causal knowledge. Additionally, the proposed approach outperforms naïve policies that assume only causal knowledge at the transmitter. Andrea Ortiz, Hussein Al-Shatri, Xiang Li 0004, Anja Klein 0002 |
ICC | 1 |
| 2015 | A resource requirement aware transmit strategy for non-regenerative multi-way relayingabstractNon-regenerative multi-antenna multi-way relaying is considered. The scenario consists of a group of single-antenna nodes that want to communicate using multiple subcarriers. Each node has a message which every other node in the group should receive. The communications between the nodes are performed via a half-duplex multi-antenna relay station. It is assumed that the nodes have individual resource requirements. A new transmit strategy, in which the individual resources requirements of the nodes are considered, is proposed. To ensure that the relay station has enough spatial dimensions to separate the incoming signals, it is proposed that the number of transmitting nodes per subcarrier is equal to the number of antennas at the relay station. Moreover, it is proposed that the required numbers of subcarriers are calculated according to the buffer level of the nodes, reflecting the amount of data a node has to transmit. To allocate the required number of subcarriers to each node, the proposed transmit strategy performs an efficient subcarrier allocation. Numerical results show that the proposed strategy outperforms existing transmit strategies especially when the number of antennas at the relay station is smaller than the number of nodes. Andrea Ortiz, Holger Degenhardt, Anja Klein 0002 |
WCNC | 1 |