VLDB 2026 Research / reviewers in the wild / expert
Saleh Mulhem
dblp:208/0372
· DBLP profile ↗
15ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0001-7380-5270ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 1 first-author · 13 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IMS: Intelligent Hardware Monitoring System for Secure SoCsabstractIn the modern Systems-on-Chip (SoC), the Advanced eXtensible Interface (AXI) protocol exhibits security vulnerabilities, enabling partial or complete denial-of-service (DoS) through protocol-violation attacks. The recent counter-measures lack a dedicated real-time protocol semantic analysis and evade protocol compliance checks. This paper tackles this AXI vulnerability issue and presents an intelligent hardware monitoring system (IMS) for real-time detection of AXI protocol violations. IMS is a hardware module leveraging neural networks to achieve high detection accuracy. For model training, we perform DoS attacks through header-field manipulation and systematic malicious operations, while recording AXI transactions to build a training dataset. We then deploy a quantization-optimized neural network, achieving 98.7% detection accuracy with2.5 million inferences/s. We subsequently integrate this IMS into a RISC-V SoC as a memory-mapped IP core to monitor its AXI bus. For demonstration and initial assessment for later ASIC integration, we implemented this IMS on an AMD Zynq UltraScale+ MPSoC ZCU104 board, showing an overall small hardware footprint (9.04% look-up-tables (LUTs), 0.23% DSP slices, and 0.70% flip-flops) and negligible impact on the overall design’s achievable frequency. This demonstrates the feasibility of lightweight, security monitoring for resource-constrained edge environments. Wadid Foudhaili, Aykut Rencber, Anouar Nechi, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
DATE | 7 |
| 2026 | Multi-Partner Project: CeCaS Accelerator Design for Efficient Supercomputing in Automotive SystemsabstractModern vehicles integrate an increasing amount of computational functionality, driven by the growing complexity of in-vehicle applications. At the same time, automotive system architectures are becoming more centralized, requiring powerful HPC platforms at the core. These platforms must deliver the performance needed for ADAS, AI, and autonomous driving, while also meeting stringent energy efficiency and safety requirements.The CeCaS project addresses these challenges across a wide range of topics and domains of expertise, including processor design in advanced FinFET technology, the transformation of the E/E architecture, and advanced packaging for automotive supercomputing platforms. Within CeCaS, our work focuses on application-specific accelerator design to enable efficient processing of compute-intensive workloads. In this paper, we present our contributions in this area, including the design of hardware accelerators for both conventional and neuromorphic AI workloads, the development and evaluation of representative AI benchmarks, and the use of virtual platforms for early design-space exploration and hardware/software co-design. Annina Gutermann, Alexey Serdyuk, Fabian Lesniak, Julian Höfer, Hella Toto-Kiesa, Tanja Harbaum, Jürgen Becker 0001, Brian Pachideh, Sven Nitzsche, Moritz Neher, Carmen Weigelt, Jann Krausse, Victor Pazmino Betancourt, Klaus Knobloch, Lukas Groth, Andrija Neskovic, Saleh Mulhem, Mladen Berekovic |
DATE | 17 |
| 2026 | Evaluating Generative AI for Functional Safety Analysis of Integrated Circuits
Mouadh Ayache, Alessandra Nardi, Aditya Raj Singh, Brian Davenport, Ganapathy Parthasarathy, Teo Cupaiuolo, Mladen Berekovic, Saleh Mulhem |
IOLTS | 8 |
| 2026 | Fast Reinforcement Learning for Robust Beam Codebooks in Future Communication SystemsabstractMillimeter wave (mmWave) and terahertz (THz) MIMO systems typically rely on predefined beamforming codebooks for initial access and data transmission. However, these codebooks are often unoptimized for specific conditions, leading to large sizes and significant beam training overhead, thereby complicating support for highly mobile applications. This paper introduces a reinforcement learning framework that optimizes beam patterns using only receive power measurements, adapting to the environment, user distribution, and hardware constraints without prior channel knowledge. The framework explores three reinforcement learning algorithms: Deep Deterministic Policy Gradient (DDPG), Twin Delayed Deep Deterministic Policy Gradient (TD3), and Soft Actor-Critic (SAC). While reinforcement learning has shown promise for beamforming, a comprehensive comparative analysis of advanced RL algorithms under a combination of realistic challenges, such as Non-Line-of-Sight (NLoS) conditions and hardware impairments for adaptive beam codebook design in mmWave/THz systems has been largely unexplored. This paper presents the first such in-depth comparative study. Simulation results demonstrate the superiority of the SAC algorithm, achieving higher beamforming gain and faster convergence compared to DDPG and TD3 in various scenarios, including LoS and NLoS conditions, even with hardware impairments. Anouar Nechi, Zakaria Narjis, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
IEEE Trans. Commun. | 5 |
| 2025 | Nail: Not Another Fault-Injection Framework for Chisel-generated RTLabstractFault simulation and emulation are essential techniques for evaluating the dependability of integrated circuits, enabling early-stage vulnerability analysis and supporting the implementation of effective mitigation strategies. High-level hardware description languages such as Chisel facilitate the rapid development of complex fault scenarios with minimal modification to the design. However, existing Chisel-based fault injection (FI) frameworks are limited by their coarse-grained, instruction-level controllability, which restricts the precision of fault modeling. This work introduces Nail, a Chisel-based open-source FI framework that overcomes these limitations by introducing statebased faults. This approach allows fault scenarios based on specific system states instead of just instruction-level triggers, removing the need for precise timing of fault activation. For greater controllability, Nail allows users to arbitrarily modify internal trigger states via software at runtime. To support this, Nail automatically generates a software interface, offering straightforward access to the instrumented design. This enables fine-tuning of fault parameters during active fault-injection campaigns, a feature particularly beneficial for FPGA emulation, where synthesis is time-consuming. Utilizing these features, Nail narrows the gap between the high speed of emulation-based FI frameworks, the usability of software-based approaches, and the controllability achieved in simulation. We demonstrate Nail’s state-based fault injection and software framework by modeling a faulty general-purpose register in a RISC-V processor. Although this might appear straightforward, it requires statedependent fault injection and was previously impossible without fundamental changes to the design. The approach was validated in both simulation and FPGA emulation, where the addition of Nail introduced less than $1 \%$ resource overhead. Robin Sehm, Christian Ewert, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
DSD | 5 |
| 2025 | A Lightweight Peripheral Design for RRAM-based LUTsabstractResistive random-access memory (RRAM) has garnered increasing interest due to its compact structure size, low energy requirements, and non-volatility. Recently, RRAM crossbars have shown potential for implementing lookup tables (LUTs) for reconfigurable logic. However, due to its analog properties, RRAM suffers from extensive peripherals, such as analog-to-digital converters (ADCs). In this paper, we tackle this problem by combining signal amplification, logic disjunction, and conversion to the digital domain in a single compact gate. As a result, the proposed RRAM-based LUT (R-LUT) design reduces the area per LUT by a factor of four and the peripheral overhead by more than six times. It is also almost three times smaller than an SRAM-based LUT implemented in the same technology. At the same time, it achieves a competitive access frequency of more than 3 GHz while consuming less than 240 fJ per inference, improving all major aspects of state-of-the-art R-LUT designs. Philipp Grothe, Christoph Hübner, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
ISCAS | 5 |
| 2025 | Lightweight Authenticated Integration and In-Field Secure Operation of System-in-PackageabstractSystem in Package (SiP) relies on integrating different chiplets potentially involving many third-party devices and chiplet foundries. This type of advanced packaging technology opens up numerous threat scenarios, especially: (a) the inauthentic and untraceable integration of chiplets into a SiP, (b) the insecure integration of malicious chiplets, which leads to a severe impact on the SiP security in the field. The current solutions require many hardware cryptographic primitives, making them costly and power-hungry. Therefore, a new lightweight solution is needed to ensure secure chiplet integration and secure SiP operation. In this article, we deal with these problems and introduce iTrustlet , as a combination of a physical unclonable function and an authenticated encryption scheme to ensure an authenticated and traceable chiplet integration. We propose a chiplet integration protocol based on iTrustlet and a classical root-of-trust (RoT) to ensure the integrated chiplets are unaltered and unreplaced. To guarantee SiP in-field security, iTrustlet with a hardware firewall (HWF) is proposed. Their interaction leads to two security features: (i) HWF provides a SiP protection mechanism, and (ii) iTrustlet secures the update of HWF rules. In particular, we provide a multilevel solution centralized around iTrustlet , focusing on lightweightness. The implementation results show that area and power overheads are 1.24% and 1.84% in the case of FPGA and 0.49% and 1.2% for ASIC implementation. Christian Ewert, Andrija Neskovic, Carsten Heinz, Felix Muuss, Alexander Treff, Marc Gourjon, Rainer Buchty, Thomas Eisenbarth 0001, Andreas Koch 0001, Mladen Berekovic, Saleh Mulhem |
ACM Trans. Design Autom. Electr. Syst. | 11 |
| 2025 | A Systematic Mapping Study on SystemC/TLM Modeling Capabilities in New Research DomainsabstractWith increasingly complex circuits and systems, the need for advanced design methodologies is growing. These methodologies shift the designers’ focus from technology-specific implementations to more abstract electronic system design (ESL). SystemC was developed to address this need. Being an open standard based on C++, SystemC facilitates hardware and software modeling across multiple levels of abstraction, with a particular emphasis on ESL. It is further enhanced by including the transaction-level modeling (TLM) layer, strengthening its capability to model communication between components, and even full-system simulators. Traditionally, SystemC/TLM has been deployed to provide hardware prototypes for software development early in the design process. However, surveys and literature reviews showing other capabilities of SystemC/TLM are scarce. Hence, it is essential to explore SystemC/TLM’s new capabilities in different domains such as in-circuit fault propagation, security assessment, and verification. In this article, we conduct a systematic mapping study (SMS) of SystemC/TLM modeling capabilities in certain research domains. We elaborate on the state-of-the-art ESL with an emphasis on SystemC/TLM-based system modeling. Subsequently, we present how such technologies can be applied to the new research domains within the field of circuit and system modeling, namely: (D.1) architecture exploration, (D.2) power estimation, (D.3) fault-injection analysis, (D.4) functional and security verification, and (D.5) side-channel analysis. This SMS highlights the advantages and disadvantages of the investigated SystemC/TLM capabilities and addresses the open challenges in these domains, concluding that SystemC/TLM offers significant potential in performance evaluation, verification, and security assessment of circuits and systems at ESL. Ahmed Mahmoudi, Andrija Neskovic, Celine Thermann, Robin Sehm, Christoph Hübner, Tavia Plattenteich, Rolf Meyer, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2024 | EMDRIVE Architecture: Embedded Distributed Computing and Diagnostics from Sensor to EdgeabstractFuture automotive architectures are expected to transition from a network-centric to a domain-centered architecture featuring central compute units. Powerful domain controllers or smart sensors alleviate the load on these central units and communication systems. These controllers execute tasks with varying criticalities on heterogeneous multicore processors, and are ideally capable of dynamically balancing the computing load between the central unit and sensors. Here, Artificial Intelligence (AI) capabilities playa crucial role, as it is in high demand for such an automotive architecture. However, AI still requires specialized accelerators to improve their computation performance. Task-oriented distributed computing with criticalities up to ASIL-D necessitates the development and utilization of specialized methodologies, such as safety, through the isolation and abstraction of low-level hardware concepts. Meanwhile, online monitoring and diagnostics become vital features to detect errors during operation. The EMDRIVE architecture includes methods, components, and strategies to enhance the performance, safety, and security of such distributed computing platforms. The nationally funded EMDRIVE project connects its twelve partners from academia and industry and is currently in its intermediate stage. Patrick Schmidt 0003, Iuliia Topko, Matthias Stammler, Tanja Harbaum, Jürgen Becker 0001, Rico Berner, Omar Ahmed, Jakub Jagielski, Thomas Seidler, Markus Abel, Marius Kreutzer, Maximilian Kirschner, Victor Pazmino Betancourt, Robin Sehm, Lukas Groth, Andrija Neskovic, Rolf Meyer, Saleh Mulhem, Mladen Berekovic, Matthias Probst, Manuel Brosch, Georg Sigl, Thomas Wild, Matthias Ernst, Andreas Herkersdorf, Florian Aigner, Stefan Hommes, Sebastian Lauer, Maximilian Seidler, Thomas Raste, Gasper Skvarc Bozic, Ibai Irigoyen Ceberio, Albrecht Mayer |
DATE | 18 |
| 2024 | Machine Learning for SRAM Stability AnalysisabstractSRAM stability is a critical challenge in technology scaling due to process variations. In this paper, we introduce a cutting-edge approach leveraging machine learning based on device and bitcell simulation to predict SRAM behavior in high sigma local and global variations. Our focus includes both high-density (HDC) and Low Voltage Cell (LVC) analysis, revealing the Extreme Gradient Boosting Regressor (XGBR) as the top performer for both. This research demonstrates the superior accuracy of the XGBR regressor in predicting key SRAM metrics, such as Access Disturb Margin (ADM), Write Margin (WRM), and Ireadmin, offering a compelling alternative to traditional statistical simulations. The purpose of such prediction is to revolutionize the design process and speed up designers’ decisionmaking. Jihene Bouhlila, Felix Last, Rainer Buchty, Mladen Berekovic, Saleh Mulhem |
ISCAS | 5 |
| 2024 | Holistic Framework for Evaluating the Trustworthiness of Integrated CircuitsabstractNew applications such as autonomous driving, cyber-physical systems, or remote surgeries demand integrated circuits (ICs) with an ever-lower tolerance for failure. Typical IC design focuses on the targets of functionality and power, performance, and area. An emerging topic in IC design is trustworthiness. It attempts to unify the various interdependent functional and non-functional aspects, such as correct functionality, reliability, security, and functional safety. Existing methodologies and standards focus on evaluating trustworthiness issues (TIs), i.e., causes of faults, and their effects on only one particular attribute. Instead, TIs should be evaluated on their effect on trustworthiness as a whole, which demands a holistic approach. In this paper, we make two main contributions. The first contribution is a framework with a set of unified evaluation criteria that can be applied across all trustworthiness attributes, and a metric, called the Residual Risk Value (RRV). The latter can be used to assess the residual risk of a TI, where a low RRV indicates low risk remaining, and vice versa. RRV considers the impact and likelihood, and the potential for risk reduction enabled by implementing countermeasures. The second contribution is a questionnaire-based measure that ranks TIs according to the priority of addressing them. The results highlight that TIs that emerge during the early stages of IC development should be treated with greater priority. Further, there is a tendency to prioritize security-related TIs as a greater risk to trustworthy ICs. Meanwhile, TIs affecting well-established aspects of IC design and verification are given a lower priority. Mouadh Ayache, Enkele Rama, Saleh Mulhem, Mladen Berekovic, Matthias Korb |
VLSI-SoC | 3 |
| 2024 | Secure Software/Hardware Hybrid In-Field Testing for System-on-ChipabstractModern Systems-on-Chip (SoCs) incorporate built-in self-test (BIST) modules deeply integrated into the device's intellectual property (IP) blocks. Such modules handle hardware faults and defects during device operation. As such, BIST results potentially reveal the internal structure and state of the device under test (DUT) and hence open attack vectors. So-called result compaction can overcome this vulnerability by hiding the BIST chain structure but introduces the issues of aliasing and invalid signatures. Software-BIST provides a flexible solution, that can tackle these issues, but suffers from limited observability and fault coverage. In this paper, we hence introduce a low-overhead software/hardware hybrid approach that overcomes the mentioned limitations. It relies on ($a$) keyed-hash message authentication code (KMAC) available on the$S$oC providing device-specific secure and valid signatures with zero aliasing and (b) the$S$oC processor for test scheduling hence increasing DUT availability. The proposed approach offers both on-chip- and remote-testing capabilities. We showcase a RISC-V-based$S$oC to demonstrate our approach, discussing system overhead and resulting compaction rates. Saleh Mulhem, Christian Ewert, Andrija Neskovic, Amrit Sharma Poudel |
VLSI-SoC | 1 |
| 2023 | SystemC Model of Power Side-Channel Attacks Against AI Accelerators: Superstition or not?abstractAs training artificial intelligence (AI) models is a lengthy and hence costly process, leakage of such a model's internal parameters is highly undesirable. In the case of AI accelerators, side-channel information leakage opens up the threat scenario of extracting the internal secrets of pre-trained models. Therefore, sufficiently elaborate methods for design verification as well as fault and security evaluation at the electronic system level are in demand. In this paper, we propose estimating information leakage from the early design steps of AI accelerators to aid in a more robust architectural design. We first introduce the threat scenario before diving into SystemC as a standard method for early design evaluation and how this can be applied to threat modeling. We present two successful side-channel attack methods executed via SystemC-based power modeling: correlation power analysis and template attack, both leading to total information leakage. The presented models are verified against an industry-standard netlist-level power estimation to prove general feasibility and determine accuracy. Consequently, we explore the impact of additive noise in our simulation to establish indicators for early threat evaluation. The presented approach is again validated via a model-vs-netlist comparison, showing high accuracy of the achieved results. This work hence is a solid step towards fast attack deployment and, subsequently, the design of attack-resilient AI accelerators. Andrija Neskovic, Saleh Mulhem, Alexander Treff, Rainer Buchty, Thomas Eisenbarth 0001, Mladen Berekovic |
ICCAD | 2 |
| 2023 | Practical Trustworthiness Model for DNN in Dedicated 6G ApplicationabstractArtificial intelligence (AI) is considered an efficient response to several challenges facing 6G technology. However, AI still suffers from a huge trust issue due to its ambiguous way of making predictions. Therefore, there is a need for a method to evaluate the AI’s trustworthiness in practice for future 6G applications. This paper presents a practical model to analyze the trustworthiness of AI in a dedicated 6G application. In particular, we present two customized deep neural networks (DNNs) to solve the automatic modulation recognition (AMR) problem in Terahertz communications-based 6G technology. Then, a specific trustworthiness model and its attributes, namely data robustness, parameter sensitivity, and security covering adversarial examples, are introduced. The evaluation results indicate that the proposed trustworthiness attributes are crucial to evaluate the trustworthiness of DNN for this 6G application. Anouar Nechi, Ahmed Mahmoudi, Christoph Herold, Daniel Widmer, Thomas Kürner, Mladen Berekovic, Saleh Mulhem |
WiMob | 7 |
| 2023 | FPGA-based Deep Learning Inference Accelerators: Where Are We Standing?abstractRecently, artificial intelligence applications have become part of almost all emerging technologies around us. Neural networks, in particular, have shown significant advantages and have been widely adopted over other approaches in machine learning. In this context, high processing power is deemed a fundamental challenge and a persistent requirement. Recent solutions facing such a challenge deploy hardware platforms to provide high computing performance for neural networks and deep learning algorithms. This direction is also rapidly taking over the market. Here, FPGAs occupy the middle ground regarding flexibility, reconfigurability, and efficiency compared to general-purpose CPUs, GPUs, on one side, and manufactured ASICs on the other. FPGA-based accelerators exploit the features of FPGAs to increase the computing performance for specific algorithms and algorithm features. Filling a gap, we provide holistic benchmarking criteria and optimization techniques that work across several classes of deep learning implementations. This article summarizes the current state of deep learning hardware acceleration: More than 120 FPGA-based neural network accelerator designs are presented and evaluated based on a matrix of performance and acceleration criteria, and corresponding optimization techniques are presented and discussed. In addition, the evaluation criteria and optimization techniques are demonstrated by benchmarking ResNet-2 and LSTM-based accelerators. Anouar Nechi, Lukas Groth, Saleh Mulhem, Farhad Merchant, Rainer Buchty, Mladen Berekovic |
ACM Trans. Reconfigurable Technol. Syst. | 3 |