EDBT 2026 Demo / reviewers in the wild / expert
Mirco Theile
dblp:233/1962
· DBLP profile ↗
11ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-1574-8858ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Objective Memory Bandwidth Regulation and Cache Partitioning for Multicore Real-Time Systems
Binqi Sun, Zhihang Wei, Andrea Bastoni, Debayan Roy, Mirco Theile, Tomasz Kloda, Rodolfo Pellizzoni, Marco Caccamo |
ECRTS | 5 |
| 2025 | Position paper: deep reinforcement learning for real-time resource managementabstractAbstract Many real-time problems can be characterized as combinatorial optimization problems where exact solutions are infeasible at scale. As problem complexity grows, handcrafted heuristics become increasingly difficult to design. Reinforcement learning (RL) has emerged as a promising alternative, enabling the discovery of decision-making policies without requiring explicit supervision. While RL does not guarantee optimality, it provides adaptive heuristics to solve complex problems. This paper explores the potential of RL for real-time resource management, outlining key principles, demonstrating an application to directed acyclic graph (DAG) scheduling, and identifying open challenges for future research. Mirco Theile, Binqi Sun, Marco Caccamo |
Real Time Syst. | 1 |
| 2025 | Quasi-Static Scheduling for Deterministic Timed Concurrent Models on Multi-Core HardwareabstractTo design performant, expressive, and reliable cyber-physical systems (CPSs), researchers extensively perform quasi-static scheduling for concurrent models of computation (MoCs) on multi-core hardware. However, these quasi-static scheduling approaches are developed independently for their corresponding MoCs, despite commonality in the approaches. To help generalize the use of quasi-static scheduling to new and emerging MoCs, this article proposes a unified approach for a class of deterministic timed concurrent models (DTCMs), including prominent models such as synchronous dataflow (SDF), Boolean-controlled dataflow (BDF), scenario-aware dataflow (SADF), and Logical Execution Time (LET). In contrast to scheduling techniques tailored exclusively to specific MoCs, our unified approach leverages a common intermediate formalism called state space finite automata (SSFA), bridging the gap between high-level MoCs and executable schedules. Once identified as DTCMs, new MoCs can directly adopt SSFA-based scheduling, significantly easing adoption. We show that quasi-static schedules facilitated by SSFA are provably free from timing anomalies and enable straightforward worst-case makespan analysis. We demonstrate the approach using the reactor model—an emerging discrete-event MoC—programmed using the Lingua Franca ( LF ) language. Experiments show that quasi-statically scheduled LF programs exhibit lower runtime overhead compared to the dynamically scheduled LF programs, and that the analyzable worst-case makespans enable compile-time deadline checking. Shaokai Lin, Erling Rennemo Jellum, Mirco Theile, Tassilo Tanneberger, Binqi Sun, Chadlia Jerad, Yimo Xu, Guangyu Feng, Magnus Mæhlum, Jian-Jia Chen, Martin Schoeberl, Linh T. X. Phan, Jerónimo Castrillón, Sanjit A. Seshia, Edward A. Lee |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Equivariant Ensembles and Regularization for Reinforcement Learning in Map-based Path PlanningabstractIn reinforcement learning (RL), exploiting environmental symmetries can significantly enhance efficiency, robustness, and performance. However, ensuring that the deep RL policy and value networks are respectively equivariant and invariant to exploit these symmetries is a substantial challenge. Related works try to design networks that are equivariant and invariant by construction, limiting them to a very restricted library of components, which in turn hampers the expressiveness of the networks. This paper proposes a method to construct equivariant policies and invariant value functions without specialized neural network components, which we term equivariant ensembles. We further add a regularization term for adding inductive bias during training. In a map-based path planning case study, we show how equivariant ensembles and regularization benefit sample efficiency and performance. Mirco Theile, Hongpeng Cao, Marco Caccamo, Alberto L. Sangiovanni-Vincentelli |
IROS | 1 |
| 2024 | Edge Generation Scheduling for DAG Tasks Using Deep Reinforcement LearningabstractDirected acyclic graph (DAG) tasks are currently adopted in the real-time domain to model complex applications from the automotive, avionics, and industrial domains that implement their functionalities through chains of intercommunicating tasks. This paper studies the problem of scheduling real-time DAG tasks by presenting a novel schedulability test based on the concept oftrivial schedulability. Using this schedulability test, we propose a new DAG scheduling framework (edge generation scheduling—EGS) that attempts to minimize the DAG width by iteratively generating edges while guaranteeing the deadline constraint. We study how to efficiently solve the problem of generating edges by developing a deep reinforcement learning algorithm combined with a graph representation neural network to learn an efficient edge generation policy for EGS. We evaluate the effectiveness of the proposed algorithm by comparing it with state-of-the-art DAG scheduling heuristics and an optimal mixed-integer linear programming baseline. Experimental results show that the proposed algorithm outperforms the state-of-the-art by requiring fewer processors to schedule the same DAG tasks.https://github.com/binqi-sun/egs Binqi Sun, Mirco Theile, Ziyuan Qin 0002, Daniele Bernardini 0002, Debayan Roy, Andrea Bastoni, Marco Caccamo |
IEEE Trans. Computers | 2 |
| 2022 | Cloud-Edge Training Architecture for Sim-to-Real Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) is a promising approach to solve complex control tasks by learning policies through interactions with the environment. However, the training of DRL policies requires large amounts of training experiences, making it impractical to learn the policy directly on physical systems. Sim-to-real approaches leverage simulations to pretrain DRL policies and then deploy them in the real world. Unfortunately, the direct real-world deployment of pretrained policies usually suffers from performance deterioration due to the different dynamics, known as the reality gap. Recent sim-to-real methods, such as domain randomization and domain adaptation, focus on improving the robustness of the pretrained agents. Nevertheless, the simulation-trained policies often need to be tuned with real-world data to reach optimal performance, which is challenging due to the high cost of real-world samples. This work proposes a distributed cloud-edge architecture to train DRL agents in the real world in real-time. In the architecture, the inference and training are assigned to the edge and cloud, separating the real-time control loop from the computationally expensive training loop. To overcome the reality gap, our architecture exploits sim-to-real transfer strategies to continue the training of simulation-pretrained agents on a physical system. We demonstrate its applicability on a physical inverted-pendulum control system, analyzing critical parameters. The real-world experiments show that our architecture can adapt the pretrained DRL agents to unseen dynamics consistently and efficiently.11A video showing a real-world training process under the proposed method can be found from https://youtu.be/hMY9-c0SST0. Hongpeng Cao, Mirco Theile, Federico G. Wyrwal, Marco Caccamo |
IROS | 2 |
| 2020 | UAV Path Planning for Wireless Data Harvesting: A Deep Reinforcement Learning ApproachabstractAutonomous deployment of unmanned aerial vehicles (UAVs) supporting next-generation communication networks requires efficient trajectory planning methods. We propose a new end-to-end reinforcement learning (RL) approach to UAV-enabled data collection from Internet of Things (IoT) devices in an urban environment. An autonomous drone is tasked with gathering data from distributed sensor nodes subject to limited flying time and obstacle avoidance. While previous approaches, learning and non-learning based, must perform expensive recomputations or relearn a behavior when important scenario parameters such as the number of sensors, sensor positions, or maximum flying time, change, we train a double deep Q-network (DDQN) with combined experience replay to learn a UAV control policy that generalizes over changing scenario parameters. By exploiting a multi-layer map of the environment fed through convolutional network layers to the agent, we show that our proposed network architecture enables the agent to make movement decisions for a variety of scenario parameters that balance the data collection goal with flight time efficiency and safety constraints. Considerable advantages in learning efficiency from using a map centered on the UAV's position over a non-centered map are also illustrated. Harald Bayerlein, Mirco Theile, Marco Caccamo, David Gesbert |
GLOBECOM | 2 |
| 2020 | UAV Coverage Path Planning under Varying Power Constraints using Deep Reinforcement LearningabstractCoverage path planning (CPP) is the task of designing a trajectory that enables a mobile agent to travel over every point of an area of interest. We propose a new method to control an unmanned aerial vehicle (UAV) carrying a camera on a CPP mission with random start positions and multiple options for landing positions in an environment containing no-fly zones. While numerous approaches have been proposed to solve similar CPP problems, we leverage end-to-end reinforcement learning (RL) to learn a control policy that generalizes over varying power constraints for the UAV. Despite recent improvements in battery technology, the maximum flying range of small UAVs is still a severe constraint, which is exacerbated by variations in the UAV's power consumption that are hard to predict. By using map-like input channels to feed spatial information through convolutional network layers to the agent, we are able to train a double deep Q-network (DDQN) to make control decisions for the UAV, balancing limited power budget and coverage goal. The proposed method can be applied to a wide variety of environments and harmonizes complex goal structures with system constraints. Mirco Theile, Harald Bayerlein, Richard Nai, David Gesbert, Marco Caccamo |
IROS | 1 |
| 2020 | Latency-Aware Generation of Single-Rate DAGs from Multi-Rate Task SetsabstractModern automotive and avionics embedded systems integrate several functionalities that are subject to complex timing requirements. A typical application in these fields is composed of sensing, computation, and actuation. The ever increasing complexity of heterogeneous sensors implies the adoption of multi-rate task models scheduled onto parallel platforms. Aspects like freshness of data or first reaction to an event are crucial for the performance of the system. The Directed Acyclic Graph (DAG) is a suitable model to express the complexity and the parallelism of these tasks. However, deriving age and reaction timing bounds is not trivial when DAG tasks have multiple rates. In this paper, a method is proposed to convert a multi-rate DAG task-set with timing constraints into a single-rate DAG that optimizes schedulability, age and reaction latency, by inserting suitable synchronization constructs. An experimental evaluation is presented for an autonomous driving benchmark, validating the proposed approach against state-of-the-art solutions. Micaela Verucchi, Mirco Theile, Marco Caccamo, Marko Bertogna |
RTAS | 2 |
| 2019 | Trajectory Estimation for Geo-Fencing Applications on Small-Size Fixed-Wing UAVsabstractThe steadily increasing popularity of Unmanned Aerial Vehicles (UAVs) is creating new opportunities in diverse fields of technology and business. However, this increase of popularity also raises safety concerns. To tackle the primary concern of keeping the UAV inside a designated region, a novel trajectory estimation algorithm for geo-fencing applications is proposed. We derive the Beta-Trajectory that takes into account constraints in curvature as well as constraints in the change of curvature which is bounded by the maximum roll-rate of the aircraft. We incorporate the Beta-Trajectory into a geo-fencing algorithm. By using our open-source uavAP autopilot, the applicability and necessity of accurate trajectory estimation algorithms for geo-fencing applications are shown on small fixed-wing aircraft. The model and algorithm are validated in high-fidelity simulations as well as in real flight testing. Mirco Theile, Simon Yu, Or D. Dantsker, Marco Caccamo |
IROS | 1 |
| 2018 | uavEE: A Modular, Power-Aware Emulation Environment for Rapid Prototyping and Testing of UAVsabstractState of the art design and testing of avionics for unmanned aircraft is an iterative process that involves many test flights, interleaved with multiple revisions of the flight management software and hardware. To significantly reduce flight test time and software development costs, we have developed a real-time UAV Emulation Environment (uavEE) using ROS that interfaces with high fidelity simulators to simulate the flight behavior of the aircraft. Our uavEE emulates the avionics hardware by interfacing directly with the embedded hardware used in real flight. The modularity of uavEE allows the integration of countless test scenarios and applications. Furthermore, we present an accurate data driven approach for modeling of propulsion power of fixed-wing UAVs, which is integrated into uavEE. Finally, uavEE and the proposed UAV Power Model have been experimentally validated using a fixed-wing UAV testbed. Mirco Theile, Or D. Dantsker, Richard Nai, Marco Caccamo |
RTCSA | 1 |