Taha Soliman

dblp:252/2274 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-9421-9489ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Neoscope: How Resilient Is My SoC to Workload Churn?
abstract
The lifetime of hardware is increasing, but the lifetime of software is not.This leads to devices that, while performant when released, have fall-off due to changing workload suitability.To ensure that performance is maintained, computer architects must begin considering the effects of evolving workloads early in the design process.The ways that workloads can change, i.e., churn, over time profoundly affect hardware's ability to maintain performance.To better understand churn, we introduce the concepts Magnitude and Disruption, which enable how workloads change to be quantitatively described.By leveraging these terms, we present churn as a span, where the churn of a given workload can be categorized as either Minimal, Perturbing, Escalating, or Volatile.To account for how churn affects performance, we propose Neoscope, the first multi-objective pre-silicon design space exploration tool for investigating System-on-Chip (SoC) architectures that are resilient to workload churn.Neoscope uses integer linear programming and concepts from job-shop scheduling to construct near-optimal SoCs for a given workload.Unlike previous methods, Neoscope approaches (and often finds) the globally optimal SoC with a single invocation, instead of needing to parameter sweep.Neoscope is also multi-objective, i.e., it can optimize for many kinds of metrics other than absolute performance, including area, energy, cost, and carbon efficiencies.Using Neoscope, we explore near-optimal SoCs for many workload-churn configurations, and investigate how changing the objective function changes the ideal SoC.Neoscope shows that small SoCs need high levels of specialization, but that this is risky, as it increases susceptibility to churn.
Joseph Rogers, Lieven Eeckhout, Taha Soliman, Magnus Jahre
ISCA3
2025 A Lightweight PUF-Based Weights Obfuscation Technique for Secure In-Memory AI Inference
abstract
In-Memory Computing (IMC) has introduced a novel computational approach that substantially improves emerging embedded AI accelerators’ latency and power consumption efficiency. Despite the numerous advantages, IMC architectures also introduce new security vulnerabilities that may compromise the confidentiality of the deployed Neural Network (NN) algorithms. In this work, following an analysis of the potential threats, we present a novel lightweight security countermeasure for IMC accelerators. This methodology can be employed to de-obfuscate the pre-trained weights of NN architectures whose bits’ significance has been reordered prior to the deployment phase onto the IMC crossbar. The proposed solution is based on the coordinated action of a Ferroelectric Field-Effect Transistor (FeFET) based Physical Unclonable Function (PUF) design and shifting registers. These components perform custom arithmetic shift operations on the values calculated by the IMC device at runtime to obtain a coherent inference computation. Furthermore, a design-space exploration method is proposed to investigate the trade-off between area overhead and the level of security provided by the implementation. The results show that with less than 3% of area overhead our design is robust against all the tested attack strategies.
Luca Parrini, Anirban Kar, Benjamin Hettwer, Taha Soliman, Yogesh Singh Chauhan, Hussam Amrouch, Norbert Wehn
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 CoNAX: Towards Comprehensive Co-Design Neural Architecture Search Using HW Abstractions
abstract
HW-aware neural architecture search (HW-NAS) aims to yield high-accuracy neural network (NN) architectures by automatically exploring multiple architectural parameters of potential network candidates. In most HW-NAS approaches, the HW parameter search space is limited. Hence, HW awareness is tied to only a few degrees of design freedom, leading to the following sub-optimalities - First, it restricts exploration of HW parameters, which can potentially lead to better network candidates; Second, HW-NAS is still entirely a software-centric process where HW-awareness is taken care by an HW function exposed to the NAS process and is oblivious to the actual deployment. To tackle the above challenges, this paper proposes a Co-Design Neural Architecture Search (Co-NAS) approach that simultaneously explores hardware and neural architecture variations, thus allowing for full system optimization. By connecting the mutual impact of variable neural networks and HW parameters on the network's prediction accuracy and on-device efficiency in a shared optimization loop, Co-Nas finds designs of optimum performance and enables HW/SW Co-Design. This work aims to enable more diverse HW search spaces (higher degrees of design freedom) for ML accelerators (such as using a virtual prototype) and efficient exploration by integrating abstract ML accelerator and NN architecture modeling into a comprehensive Co-Nas environment. In our experiments, we explore hardware variations of a baseline accelerator architecture to demonstrate how our work can help find designs with better hardware latency and comparable network accuracy. Designs yielded by our framework provide a speedup of$1.4\times$compared to the baseline on a restricted SW search space at the same HW resources.
Yannick Braatz, Taha Soliman, Shubham Rai, Dennis Rieber, Oliver Bringmann 0001
ASAP2
2024 Error Detection and Correction Codes for Safe In-Memory Computations
abstract
In-Memory Computing (IMC) introduces a new paradigm of computation that offers high efficiency in terms of latency and power consumption for AI accelerators. However, the non-idealities and defects of emerging technologies used in advanced IMC can severely degrade the accuracy of inferred Neural Networks (NN) and lead to malfunctions in safety-critical applications. In this paper, we investigate an architectural-level mitigation technique based on the coordinated action of multiple checksum codes, to detect and correct errors at run-time. This implementation demonstrates higher efficiency in recovering accuracy across different AI algorithms and technologies compared to more traditional methods such as Triple Modular Redundancy (TMR). The results show that several configurations of our implementation recover more than 91% of the original accuracy with less than half of the area required by TMR and less than 40% of latency overhead.
Luca Parrini, Taha Soliman, Benjamin Hettwer, Jan Micha Borrmann, Simranjeet Singh, Ankit Bende, Vikas Rana, Farhad Merchant, Norbert Wehn
ETS2
2024 AIO: An Abstraction for Performance Analysis Across Diverse Accelerator Architectures
abstract
Specialization is the key approach for continued performance growth beyond the end of Dennard scaling. Academics and industry are hence continuously proposing new accelerator architectures, including conventional Domain-Specific Accelerators (DSAs) and emerging Processing in Memory (PIM) accelerators. We are thus fast approaching an era in which earlystage accelerator analysis is critical for maintaining the productivity of software developers, system software designers, and computer architects — to ensure that they focus time-consuming implementation and optimization efforts on the most favorable class of accelerators for the problem at hand. Unfortunately, existing approaches fall short because they either adopt a level of abstraction that is too high — and therefore are unable to account for key performance phenomena — or too low — because they focus on details that do not generalize across diverse accelerators. Our Architecture-Independent Operation (AIO) abstraction addresses this issue by leveraging that accelerators typically focus on data-level parallelism, and an AIO is hence a key piece of algorithm-level data-parallel work that remains the same across diverse accelerators. To demonstrate that the AIO abstraction can be accurate and useful, we create the AccMe performance model which predicts kernel performance by estimating the number of clock cycles spent on compute, memory, and invocation overhead while accounting for overlap between compute and memory cycles as well as finite memory bandwidth. We demonstrate that AccMe can be accurate, i.e., it yields an average performance prediction error of $5.6 \%$ across our diverse kernels and accelerators. This is a significant improvement over the $\mathbf{2 0. 6 \%}$ average error of curve-fitted Roofline which provides the best-case accuracy of Roofline’s operational intensity abstraction. We further demonstrate that AccMe is useful through three case studies that illustrate (i) how developers can use AccMe for accelerator selection under uncertainty; (ii) how system software can use AccMe for scheduling — and thereby improve throughput by $2.8 \times$ on average compared to Roofline-driven scheduling; and (iii) how computer architects can use AccMe for architectural exploration.
Joseph Rogers, Taha Soliman, Magnus Jahre
ISCA2
2023 SimPyler: A Compiler-Based Simulation Framework for Machine Learning Accelerators
abstract
Co-optimization of hardware and software in modern deep neural network (DNN) systems can be performed using design space exploration (DSE) tools. Leveraging estimation models to predict design decisions' impact on the system's performance allows a fast evaluation of up to billions of architectural choices. In this work, we propose SimPyler, an end-to-end framework for latency estimations of DNN workloads on machine learning (ML) accelerators. SimPyler represents DNN kernels as graphs executed on abstract accelerator models to simulate the system's latency. By generating the entire simulation infrastructure automated from the DNN operator description, the framework can flexibly adjust to changes at the hardware or algorithmic level, enabling the usage in DSE applications. A key enabler in this automation process is a machine-learning compiler. The framework is implemented in Python, using only open-source software. We demonstrate and validate the proposed methodology by modeling different single-core and multi-core hardware architectures and DNNs, comparable to state-of-the-art. Our experiments show we can estimate the end-to-end latency with an average error of 4.12%
Yannick Braatz, Dennis Rieber, Taha Soliman, Oliver Bringmann 0001
ASAP3
2022 FeFET versus DRAM based PIM Architectures: A Comparative Study
abstract
The throughput and energy efficiency of compute-centric architectures for memory intensive Deep Neural Networks (DNN) applications are limited by memory bound issues like high data-access energy, long latencies, and limited bandwidth. Processing-in-Memory (PIM) is a very promising approach to address these challenges and bridge the memory-computation gap. PIM places computational logic inside the memory to exploit minimum data movement and massive internal data parallelism. There are currently two PIM trends: 1) Use of emerging non-volatile memories to perform highly parallel analog computation of MAC operations and implicit storage of weights within the memory arrays, and 2) exploiting mature memory technologies that are enhanced by additional logic to enable efficient computation of MAC operations near the memory arrays. In this paper, we will compare both trends from an architectural perspective. Our study mainly emphasizes on FeFET memories (an emerging memory candidate) and DRAM memories (a mature memory candidate). We will highlight the major architectural constraints of these memory candidates that impact the PIM designs and their overall performance. Finally, we will assess feasible choice of candidate for different computations or DNN task types.
Chirag Sudarshan, Taha Soliman, Thomas Kämpfe, Christian Weis, Norbert Wehn
VLSI-SoC2
2022 FELIX: A Ferroelectric FET Based Low Power Mixed-Signal In-Memory Architecture for DNN Acceleration
abstract
Today, a large number of applications depend on deep neural networks (DNN) to process data and perform complicated tasks at restricted power and latency specifications. Therefore, processing-in-memory (PIM) platforms are actively explored as a promising approach to improve the throughput and the energy efficiency of DNN computing systems. Several PIM architectures adopt resistive non-volatile memories as their main unit to build crossbar-based accelerators for DNN inference. However, these structures suffer from several drawbacks such as reliability, low accuracy, large ADCs/DACs power consumption and area, high write energy, and so on. In this article, we present a new mixed-signal in-memory architecture based on the bit-decomposition of the multiply and accumulate (MAC) operations. Our in-memory inference architecture uses a single FeFET as a non-volatile memory cell. Compared to the prior work, this system architecture provides a high level of parallelism while using only 3-bit ADCs. Also, it eliminates the need for any DAC. In addition, we provide flexibility and a very high utilization efficiency even for varying tasks and loads. Simulations demonstrate that we outperform state-of-the-art efficiencies with 36.5 TOPS/W and can pack 2.05 TOPS with 8-bit activation and 4-bit weight precision in an area of 4.9 mm 2 using 22 nm FDSOI technology. Employing binary operation, we obtain 1169 TOPS/W and over 261 TOPS/W/mm 2 on system level.
Taha Soliman, Nellie Laleni, Tobias Kirchner, Franz Müller 0001, Thomas Kämpfe, Andre Guntoro, Norbert Wehn
ACM Trans. Embed. Comput. Syst.1
2021 A Novel DRAM-Based Process-in-Memory Architecture and its Implementation for CNNs
abstract
Processing-in-Memory (PIM) is an emerging approach to bridge the memory-computation gap. One of the key challenges of PIM architectures in the scope of neural network inference is the deployment of traditional area-intensive arithmetic multipliers in memory technology, especially for DRAM-based PIM architectures. Hence, existing DRAM PIM architectures are either confined to binary networks or exploit the analog property of the sub-array bitlines to perform bulk bit-wise logic operations. The former reduces the accuracy of predictions, i.e. Quality-of-results, while the latter increases overall latency and power consumption.
Chirag Sudarshan, Taha Soliman, Cecilia De la Parra, Christian Weis, Leonardo Ecco, Matthias Jung 0001, Norbert Wehn, Andre Guntoro
ASP-DAC2
2021 Exploiting Resiliency for Kernel-Wise CNN Approximation Enabled by Adaptive Hardware Design
abstract
Efficient low-power accelerators for Convolutional Neural Networks (CNNs) largely benefit from quantization and approximation, which are typically applied layer-wise for efficient hardware implementation. In this work, we present a novel strategy for efficient combination of these concepts at a deeper level, which is at each channel or kernel. We first apply layer-wise, low bit-width, linear quantization and truncation-based approximate multipliers to the CNN computation. Then, based on a state-of-the-art resiliency analysis, we are able to apply a kernel-wise approximation and quantization scheme with negligible accuracy losses, without further retraining. Our proposed strategy is implemented in a specialized framework for fast design space exploration. This optimization leads to a boost in estimated power savings of up to 34% in residual CNN architectures for image classification, compared to the base quantized architecture.
Cecilia De la Parra, Ahmed El-Yamany, Taha Soliman, Akash Kumar 0001, Norbert Wehn, Andre Guntoro
ISCAS3
2020 Efficient FeFET Crossbar Accelerator for Binary Neural Networks
abstract
This paper presents a novel ferroelectric field-effect transistor (FeFET) in-memory computing architecture dedicated to accelerate Binary Neural Networks (BNNs). We present in-memory convolution, batch normalization and dense layer processing through a grid of small crossbars with reduced unit size, which enables multiple bit operation and value accumulation. Additionally, we explore the possible operations parallelization for maximized computational performance. Simulation results show that our new architecture achieves a computing performance up to 2.46 TOPS while achieving a high power efficiency reaching 111.8 TOPS/Watt and an area of 0.026 mm2in 22nm FDSOI technology.
Taha Soliman, Ricardo Olivo, Tobias Kirchner, Cecilia De la Parra, Maximilian Lederer, Thomas Kämpfe, Andre Guntoro, Norbert Wehn
ASAP1