Muhammad Sabih

dblp:150/0573 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0003-2066-646XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Multi-Partner Project: A Holistic and Open-Source Approach to Efficient, Secure and Reliable AI Hardware Deployment in DI-EDAI
abstract
Artificial Intelligence (AI) has demonstrated strong capabilities across various domains over the past decade. Edge and specifically mission-critical applications, such as automotive and aerospace, require both high performance and efficiency without compromises in security and reliability. This stems from tightly constrained power consumption, failures that can have catastrophic consequences and devices that may be physically accessible to malicious actors. AI algorithm deployment to hardware also presents significant barriers, requiring specialized knowledge and expensive development tools. The DI-EDAI project aims to offer a holistic approach for connecting high-level AI algorithms with hardware implementations while tackling the aforementioned issues. Unlike other approaches that address individual aspects of the AI deployment flow, we investigate solutions across multiple layers of the design stack. Through our work we develop efficient hardware, map AI algorithms to hardware while simultaneously ensuring security and reliability. Furthermore, we leverage AI-techniques to assist with Electronic Design Automation (EDA) workflows for design optimization, verification and implementation. Our open source approach aims to reduce entry barriers, promote transparency and education, and spark innovation. This paper presents the current state of the DI-EDAI project at midterm, highlighting our latest contributions, identifying limitations in existing state-of-the-art approaches, and outlining ongoing work to address these gaps.
Georgios Sotiropoulos, Felix Frombach, Julian Höfer, Tanja Harbaum, Jürgen Becker 0001, Henrik Iver Thorøe, Vincent Meyers, Mehdi Baradaran Tahoori, Zeynep Demirdag, Mohammed Bakr Sikal, Hassan Nassar, Heba Khdr, Jörg Henkel, Christopher Wolters, Philipp van Kempen, Johannes Geier, Ulf Schlichtmann, Batuhan Sesli, Muhammad Sabih, Jakob Wittmann, Frank Hannig, Jürgen Teich, Lukas Steiner, Norbert Wehn, Mohamed Shelkamy Ali, Philipp Schmitz, Wolfgang Kunz, Stefan Koegler, Georg Sigl
DATE19
2025 Special Session - Hardware-Software Co-Design for Machine Learning Systems Made Open-Source
abstract
Chip technologies are crucial for the digital transformation of industry and society. Machine Learning (ML) and Artificial Intelligence (AI) are increasingly shaping both daily life and industrial applications, with AI hardware playing a vital role in enabling efficient and scalable ML deployment. However, significant challenges remain in bridging the gap between ML algorithm development and hardware implementation, particularly for edge ML applications where efficiency, power constraints, and adaptability are critical. In such resource-constrained environments, hardware-software co-design becomes essential to achieve the necessary trade-offs between performance, energy efficiency, and system responsiveness. One of the key bottlenecks in ML hardware development is the lack of seamless integration between ML toolchains and electronic design automation (EDA) tools for hardware synthesis and mapping. Current solutions often require extensive manual optimization and costly proprietary software, limiting accessibility and innovation. Open-source tools can play a transformative role in democratizing ML hardware design, fostering collaboration, and addressing the growing shortage of skilled professionals. This paper covers key aspects of hardware-software co-design for ML systems, such as ML algorithms, hardware design, compiler technologies and system security, with a focus on open-source solutions. We highlight the critical need for open-source toolchains that connect ML model development with hardware synthesis and optimization and present solutions for custom hardware, as well as FPGA accelerators.
Mehdi Baradaran Tahoori, Vincent Meyers, Mahboobe Sadeghipourrudsari, Huashuangyang Xu, Jürgen Becker 0001, Tanja Harbaum, Felix Frombach, Julian Höfer, Georgios Sotiropoulos, Jörg Henkel, Zeynep Demirdag, Heba Khdr, Hassan Nassar, Ulf Schlichtmann, Johannes Geier, Philipp van Kempen, Georg Sigl, Stefan Koegler, Matthias Probst, Jürgen Teich, Frank Hannig, Muhammad Sabih, Batuhan Sesli, Norbert Wehn, Lukas Steiner, Wolfgang Kunz, Mohamed Shelkamy Ali
CODES+ISSS22
2025 Design of Machine Learning Accelerators as RISC-V Extensions using an Open Source Tool Flow
abstract
The fast-evolving nature of machine learning applications demands agile hardware development to keep pace with innovation. Harnessing the customizability of the open-source RISC-V instruction set architecture (ISA), we present an automated methodology to accelerate activation functions in quantized long short-term memory (LSTM) networks through custom functional units as hardware extensions. One common approach to accelerate the computation of activation functions is to use approximation techniques, such as function tables, which enable a fast lookup but are, in turn, memory-intensive. To address this challenge, we analyze execution profiles of LSTM networks and present a flexible method for generating and exploring table-based function approximation. In order to reflect memory constraints, we propose splitting the input range of computationally expensive activation functions, such as sigmoid and hyperbolic tangent, into intervals using different quantization granularities for each subinterval. Our open-source design flow includes simulation-based verification, synthesis, placement, routing, and compilation. The results highlight the potential of table-based acceleration by addressing trade-offs between memory demand and accuracy, and provide an efficient hardware/software co-design solution for AI applications.
Batuhan Sesli, Muhammad Sabih, Frank Hannig, Jürgen Teich
ICCAD2
2024 Accelerating DNNs Using Weight Clustering on RISC-V Custom Functional Units
abstract
Weight clustering is typically used to compress a Deep Neural Network (DNN) by reducing the number of unique weight values, which can be encoded using a few bits. However, using weight clustering for acceleration remains an unexplored area. In this work, we propose a design using Custom Functional Units (CFUs) to accelerate DNNs with weight clustering on a RISC-V-based$\mathbf{SoC}$. We evaluate our accelerator on resource-constrained ML use cases and are able to report considerable speedups of up to 8 times with minimal overhead in the utilization of FPGA resources.
Muhammad Sabih, Batuhan Sesli, Frank Hannig, Jürgen Teich
DATE1
2022 Deception detection on social media: A source-based perspective
Khubaib Ahmed Qureshi, Rauf Ahmed Shams Malick, Muhammad Sabih, Hocine Cherifi
Knowl. Based Syst.3
2020 Clustering-Based Scenario-Aware LTE Grant Prediction
abstract
Reducing the energy consumption of mobile phones is a crucial design goal for cellular modem solutions for LTE and 5G standards. Recent approaches for dynamic power management incorporate traffic prediction to power down components of the modem as often as possible. These predictive approaches have been shown to still provide substantial energy savings, even if trained purely on-line. However, a higher prediction accuracy could be achieved when performing predictor training off-line. Additionally, having pre-trained predictors opens up the ability to successfully employ predictive techniques also in less favorable situations such as short intervals of stable traffic patterns. For this purpose, we introduce a notion of similarity, based on which a clustering is performed to identify similar traffic patterns. For each resulting cluster, i.e., an identified traffic scenario, one predictor is designed and trained off-line. At run time, the system selects the pre-trained predictor with the lowest average short-term false negative rate allowing for energy-efficient and highly accurate on-line prediction. Through experiments, it is shown that the presented mixed static/dynamic approach is able to improve the prediction accuracy and energy savings compared to a state-of-the-art approach by factors of up to 2 and up to 1.9, respectively.
Peter Brand, Muhammad Sabih, Joachim Falk, Jonathan Ah Sue, Jürgen Teich
WCNC2