VLDB 2026 Research / reviewers in the wild / expert
Mohammed Bakr Sikal
dblp:305/9495
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-9788-9026ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 10 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: A Holistic and Open-Source Approach to Efficient, Secure and Reliable AI Hardware Deployment in DI-EDAIabstractArtificial Intelligence (AI) has demonstrated strong capabilities across various domains over the past decade. Edge and specifically mission-critical applications, such as automotive and aerospace, require both high performance and efficiency without compromises in security and reliability. This stems from tightly constrained power consumption, failures that can have catastrophic consequences and devices that may be physically accessible to malicious actors. AI algorithm deployment to hardware also presents significant barriers, requiring specialized knowledge and expensive development tools. The DI-EDAI project aims to offer a holistic approach for connecting high-level AI algorithms with hardware implementations while tackling the aforementioned issues. Unlike other approaches that address individual aspects of the AI deployment flow, we investigate solutions across multiple layers of the design stack. Through our work we develop efficient hardware, map AI algorithms to hardware while simultaneously ensuring security and reliability. Furthermore, we leverage AI-techniques to assist with Electronic Design Automation (EDA) workflows for design optimization, verification and implementation. Our open source approach aims to reduce entry barriers, promote transparency and education, and spark innovation. This paper presents the current state of the DI-EDAI project at midterm, highlighting our latest contributions, identifying limitations in existing state-of-the-art approaches, and outlining ongoing work to address these gaps. Georgios Sotiropoulos, Felix Frombach, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Henrik Iver Thorøe, Vincent Meyers, Mehdi Baradaran Tahoori, Zeynep Demirdag, Mohammed Bakr Sikal, Hassan Nassar, Heba Khdr, Jörg Henkel, Christopher Wolters, Philipp van Kempen, Johannes Geier, Ulf Schlichtmann, Batuhan Sesli, Muhammad Sabih, Jakob Wittmann, Frank Hannig, Jürgen Teich, Lukas Steiner, Norbert Wehn, Mohamed Shelkamy Ali, Philipp Schmitz, Wolfgang Kunz, Stefan Koegler, Georg Sigl |
DATE | 10 |
| 2026 | ARDiS: A Portable and Unified Resource Management Framework in Real Hardware SystemsabstractDesigning efficient RM strategies is a cornerstone of modern computing, driving innovations in performance optimization, energy efficiency, and security. While simulators have long been the go-to tools for RM research, they fail to balance accuracy and practicality: high-fidelity simulators are excruciatingly slow, and low-fidelity ones compromise on reliability. Real hardware offers unparalleled precision and accuracy but remains underutilized due to significant barriers, including fragmented implementations, lack of portability, and prohibitive development overhead. We present ARDiS, the first open-source 1 and portable framework to provide a unified, architecture-agnostic platform for running system-level resource management (RM) techniques directly on real hardware. ARDiS eliminates the need to “reinvent the wheel,” enabling researchers to design, implement, and evaluate sophisticated RM strategies—including machine learning-based approaches—with minimal effort and maximum reproducibility. To demonstrate its versatility, we evaluate ARDiS on two real-world hardware platforms: a server-grade heterogeneous processor (Intel i9-12900) and a resource-constrained embedded system (NVIDIA Jetson TX2). Through extensive experimentation, we validate the ability of ARDiS to deliver accurate, scalable, and reproducible results across diverse platforms and application domains. By lowering the barriers to hardware-based RM research, ARDiS empowers the design automation community to explore new frontiers in system-level optimization and innovation. Mohammed Bakr Sikal, Jeferson González-Gómez, Andreas Noebel, Heba Khdr, Jörg Henkel |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2025 | Contention-Aware Forecasting of Energy Efficiency through Sequence-Based Models in Modern Heterogeneous ProcessorsabstractWe present EffiCast, the first methodology for contentionaware energy efficiency forecasting in clustered heterogeneous processors using sequence-based models. Through extensive experimental analysis of energy efficiency sensitivities across core types, voltage/frequency (V/f) levels, application phases, and resource contention scenarios, EffiCast uncovers key factors driving energy efficiency variability in modern heterogeneous processors. Leveraging structured data generation and advanced LSTM- and Transformer-based models, EffiCast achieves unprecedented accuracy while outperforming state-of-the-art predictive techniques. Deployed on a real heterogenous processor with Intel’s oneDNN acceleration, EffiCast delivers inference latencies as low as 1.82 ms per sequence, enabling seamless integration into proactive resource management frameworks. With the ability to forecast future system states under dynamic workloads, EffiCast sets a new standard for energy efficiency optimization in energy-constrained application domains. Mohammed Bakr Sikal, Jeferson González-Gómez, Heba Khdr, Jörg Henkel |
DAC | 1 |
| 2024 | Multi-Agent Reinforcement Learning for Thermally-Restricted Performance Optimization on ManycoresabstractThe problem of performance maximization under a thermal constraint has been tackled by means of dynamic voltage and frequency scaling (DVFS) in many system-level optimization techniques. State-of-the-art ones have exploited Su-pervised Learning (SL) to develop models that predict power and performance characteristics of applications and temperature of the cores. Such predictions enable proactive and efficient optimization decisions that exploit performance potentials under a temperature constraint. SL- based models are built at design time based on training data generated considering specific environment settings, i.e., processor architecture, cooling system, ambient temperature, etc. Hence, these models cannot adapt at runtime to different environment settings. In contrast, Reinforcement Learning (RL) employs an agent that explores and learns the environment at runtime, and hence can adapt to its potential changes. Nonetheless, using an RL agent to perform optimization on manycores is challenging because of the inherent large state/action spaces that might hinder the agent's ability to converge. To get the advantages of RL while tackling this challenge, we employ for the first time multi -agent RL to perform thermally-restricted performance optimization for manycores through DVFS. We investigated two RL algorithms-Table-based Q-Learning (TQL) and Deep Q-Learning (DQL)-and demonstrated that the latter outperforms the former. Compared to the state of the art, our DQL delivers a significant performance improvement of 34.96% on average, while also guaranteeing thermally -safe operation on the manycore. Our evaluation reveals the runtime adaptability of our DQL to varying workloads and ambient temperatures. Heba Khdr, Mustafa Enes Batur, Kanran Zhou, Mohammed Bakr Sikal, Jörg Henkel |
DATE | 4 |
| 2024 | Balancing Security and Efficiency: System-Informed Mitigation of Power-Based Covert ChannelsabstractAs the digital landscape continues to evolve, the security of computing systems has become a critical concern. Power-based covert channels (e.g., thermal covert channel s (TCCs)), a form of communication that exploits the system resources to transmit information in a hidden or unintended manner, have been recently studied as an effective mechanism to leak information between malicious entities via the modulation of CPU power. To this end, dynamic voltage and frequency scaling (DVFS) has been widely used as a countermeasure to mitigate TCCs by directly affecting the communication between the actors. Although this technique has proven effective in neutralizing such attacks, it introduces significant performance and energy penalties, that are particularly detrimental to energy-constrained embedded systems. In this article, we propose different system-informed countermeasures to power-based covert channels from the heuristic and machine learning (ML) domains. Our proposed techniques leverage task migration and DVFS to jointly mitigate the channels and maximize energy efficiency. Our extensive experimental evaluation on two commercial platforms: 1) the NVIDIA Jetson TX2 and 2) Jetson Orin shows that our approach significantly improves the overall energy efficiency of the system compared to the state-of-the-art solution while nullifying the attack at all times. Jeferson González-Gómez, Mohammed Bakr Sikal, Heba Khdr, Lars Bauer, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | ML-Based Thermal and Cache Contention Alleviation on Clustered Manycores With 3-D HBMabstractEnabled by the recent advancements in 2.5D/3-D integration and packaging, the integration of clustered manycore processors with high-bandwidth memory (HBM) is gaining prominence to satisfy the increasing memory bandwidth demands. Although this integration can offer significant performance gains, it is still limited by cache contention in the final-level cache on the clusters and by the thermal issues in the 3-D HBM. While the existing state-of-the-art resource management techniques have tackled these issues in isolation, we argue that the cache contention and the temperature of both the manycore and the HBM must be considered jointly to harness the full performance potential of such modern architectures. To cover this gap in the literature, we present MTCM, the first resource management technique that considers the cache contention in maximizing the system performance, while maintaining the thermal safety across both the manycore and the HBM stack. Enabled by our accurate, yet lightweight, neural network models, our proposed task migration and dynamic voltage and frequency scaling policies can accurately predict the impact of runtime decisions on the performance and temperature of both the subsystems. Our extensive evaluation experiments reveal a significant performance improvement over existing state of the art by up to$1\times $, while maintaining thermal safety of both the manycore and the HBM. Mohammed Bakr Sikal, Heba Khdr, Lokesh Siddhu, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Smart Detection of Obfuscated Thermal Covert Channel Attacks in Many-core ProcessorsabstractIn thermal covert channel (TCC) attacks, malicious applications seek to leak private information in a stealthy and hard-to-detect manner. State-of-the-art approaches for TCC detection employ the Discrete Fourier Transform (DFT) combined with heuristics to identify possible channels. However, as we demonstrate in this paper, these approaches are limited when detecting short-duration attacks, where an attacker intentionally halts the transmission for a time interval to avoid the detection. In order to overcome this limitation of the state-of-the-art solutions, we propose the first detection method for short-duration TCC attacks. Our solution, Dotecca, is a machine learning-based technique that employs short windows of time-domain measurements instead of the DFT to detect TCCs. To evaluate our solution, we introduce a new obfuscated short-duration attack that disguises as a regular application from the perspective of a DFT spectrum. Our experiments show that the new obfuscated attack is able to remain undetected even under advanced DFT-based state-of-the-art detection approaches, reducing their detection accuracy to about 18 %. In contrast, our smart detection approach is able to detect state-of-the-art and new obfuscated attacks with an accuracy of 99 %. Moreover, our solution reduces the overhead of the DFT-based state-of-the-art solution by more than 14 ×. Jeferson González-Gómez, Mohammed Bakr Sikal, Heba Khdr, Lars Bauer, Jörg Henkel |
DAC | 2 |
| 2023 | Machine Learning-based Thermally-Safe Cache Contention Mitigation in Clustered ManycoresabstractWe present the first technique that mitigates cache contention under thermal constraints in clustered manycores. We show by means of extensive experiments that significant performance gains in this scenario can be achieved. The background is that concurrently-running applications on manycore clusters compete for the shared cache, slowing down their execution. In addition, heavy parallel computations on physically-close cores increase temperatures to non-sustainable levels, which in turn triggers a throttle down of voltage/frequency levels and hence performance is compromised. These problems are not unknown, but as our analysis shows, tackling them independently is sub-optimal. We introduce the first task migration technique that jointly mitigates cache contention while enforcing the thermal constraint at the same time. It works in conjunction with cluster-level dynamic voltage and frequency scaling. Our technique needs to predict the impact of task migration on performance considering cache contention. Since it is impossible to derive an analytical model for cache contention that is both sufficiently accurate and practically feasible to implement, we employ an accurate, yet lightweight neural network (NN) model. As a result, we can operate the manycore system at higher performance while safely staying within thermal constraints. We report a significant step forward in this paper and unveil new potentials for performance optimization. Mohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg Henkel |
DAC | 1 |
| 2022 | Thermal- and Cache-Aware Resource Management based on ML- Driven Cache Contention PredictionabstractWhile on-chip many-core systems enable a large number of applications to run in parallel, the increased overall performance may come at the cost of complicating the performance constraints of individual applications due to contention on shared resources. For instance, the competition for last-level cache by concurrently-running applications may lead to slowing down the execution and to potentially violating individual performance constraints. Clustered many-cores reduce cache contention at chip level by sharing caches only at cluster level. To reduce cache con-tention within a cluster, state-of-the art techniques aim to co-map a memory-intensive application with a compute-intensive application onto one cluster. However, compute-intensive applications typ-ically consume high power, and therefore, executing another application in their nearby cores may lead to high temperatures. Hence, there is a trade-off between cache contention and temperature. This paper is the first to consider this trade-off through a novel thermal- and cache-aware resource management technique. We build a neural network (NN)-based model to predict the slowdown of the application execution induced by cache contention feeding our resource management technique that then optimizes the application mapping and selects the voltage/frequency levels of the clus-ters to compensate for the potential contention-induced slowdown. Thereby, it meets the performance constraints, while minimizing temperature. Compared to the state of the art, our technique significantly reduces the temperature by 30% on average, while satisfying performance constraints of all individual applications. Mohammed Bakr Sikal, Heba Khdr, Martin Rapp, Jörg Henkel |
DATE | 1 |
| 2021 | SmartBoost: Lightweight ML-Driven Boosting for Thermally-Constrained Many-Core ProcessorsabstractDynamic voltage and frequency scaling (DVFS)-based boosting is indispensable for optimizing the performance of thermally-constrained many-core processors. State-of-the-art techniques employ the voltage/frequency (V/D sensitivity of the performance of an application as a boosting metric. This paper demonstrates that this leads to suboptimal boosting decisions because the sensitivities of power and temperature also play a profound impact and need to be included within the optimization. Therefore, we introduce a novel boosting metric that integrates all relevant metrics: the application-dependent V/f sensitivities of performance and power, and the core-dependent sensitivity of the temperature. This new boosting metric is derived at run-time using machine learning via a neural network (NN) model, which accurately estimates the V/f sensitivities of performance and power of a priori unknown applications with diverse and time-varying characteristics. This new metric enables to build a smart, yet lightweight, boosting technique to maximize the performance under a temperature constraint. The experimental results demonstrate a 21 % average improvement of the system performance over the state-of-the-art at a negligible run-time overhead of 0.8 %. Martin Rapp, Mohammed Bakr Sikal, Heba Khdr, Jörg Henkel |
DAC | 2 |