Hassan Ghasemzadeh Mohammadi

dblp:62/6023-2 · also Hassan Ghasemzadeh 0002 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 DP-MCTS: Deep Playout-Driven MCTS for Approximate Accelerator Design
abstract
Approximate computing offers substantial efficiency gains for modern accelerators, yet register-transfer level design space exploration (DSE) remains challenging due to the exponential growth of approximation choices. Although recent use of machine learning (ML)-based accuracy estimators has reduced exploration runtime, it introduces a new challenge: estimator inaccuracy, which becomes increasingly problematic for larger circuits. In this paper, we propose DP-MCTS, an enhanced Monte Carlo Tree Search framework that rethinks estimator usage by treating ML models as lightweight guidance tools rather than decision makers. DP-MCTS employs deep playouts to provide fast, informative look-ahead along candidate paths, coupled with a promotion mechanism that selectively validates promising nodes through simulation-based accuracy evaluation. This synergy preserves broad design space coverage while mitigating estimator-driven drift, significantly improving exploration robustness and solution quality. Experiments across diverse accelerator benchmarks show that DP-MCTS closely matches or even outperforms the solution quality of an MCTS-based search, achieving up to 6.0% additional area reduction and an order of magnitude lower runtime cost, thereby delivering a substantially improved quality-runtime tradeoff over existing search-based frameworks.
Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Sayed Morteza Jawadi, Marco Platzner
DDECS2
2025 Swift Synthesis of Approximate Hardware Accelerators Using Generative Adversarial Networks
abstract
Deploying modern applications with significant resource demands is often challenging, but approximate designs offer a promising alternative by delivering high performance with minimal compromises in output quality. Traditionally, approximate hardware accelerators have been developed through search-based iterative frameworks, which suffer from long runtimes due to the exponential growth of the design space. A significant portion of the runtime is consumed by either invalid nodes or valid nodes that offer minimal improvements in performance metrics, such as runtime or power consumption. This severely limits the thorough exploration of the design space. In this paper, we introduce a novel approach for synthesizing approximate accelerators that leverages sparsity to reduce the complexity of the design space exploration problem. Our method employs a generative adversarial network (GAN) to rapidly generate a diverse set of high-quality design nodes, eliminating the need for costly node evaluations. This enables the swift creation of approximate accelerators generated for any given error threshold in a fraction of time as compared to a simulation-based framework. We conducted experiments on a suite of benchmarks from real-world domains, demonstrating that our methodology can generate approximate hardware designs with significant area and power savings, comparable to state-of-the-art search-based approaches. In a comparative evaluation against two leading methods, our approach achieved equal or better quality results for two out of four benchmarks while reaching up to 55% area savings, thus effectively demonstrating a new avenue for automated generation of approximate accelerators.
Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner
VLSI-SoC2
2025 Design Space Exploration for Approximate Circuits via Checkpointing and DNN-Based Estimators
abstract
Approximate computing (AC) is a design paradigm that trades in reductions in hardware area, power, and delay of digital circuits for an increased error. Workflows for synthesizing approximate accelerators are often implemented as search-approximate–evaluate cycles that build a search tree by iteratively approximating components of the original circuit. Design space exploration (DSE) for approximate accelerators is challenging due to the sheer size of the design space and the computational costs for validating the error bound during the search. While early work employed rather slow simulation-based methods, more recent approaches use machine learning (ML) to create fast estimators for validating error bounds. For larger accelerators and design spaces, however, the mispredictions of ML-based methods can steer the search process to explore invalid regions of the design space. In this article, we introduce the frameworkDeepApproxfor approximate accelerator synthesis. The novelty of this framework is its combination of fast ML-based quality estimators with occasional simulations to correct possible mispredictions. The simulations are performed at the so-called checkpoints during the search, and we propose two methods: static and dynamic checkpointing (DC). We present theDeepApproxflow and elaborate on the ML quality estimators and the checkpointing mechanisms. Using a set of smaller and larger benchmarks, we compareDeepApproxwith two state-of-the-art flows: a simulation-based method leveraging Monte Carlo tree search (MCTS) and an ML-based method with fast estimators. Our experiments demonstrate thatDeepApproxcan achieve similar reductions in hardware area and power consumption compared to simulation-based methods, albeit at much lower runtimes, and higher reductions in these metrics compared to methods relying exclusively on ML-based estimators. Thus, by combining ML-based techniques with simulations,DeepApproxprovides a scalable DSE for approximate accelerators.
Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner
IEEE Trans. Very Large Scale Integr. Syst.2
2021 LDAX: A Learning-based Fast Design Space Exploration Framework for Approximate Circuit Synthesis
abstract
The majority of existing frameworks for automated synthesis of Approximate Circuits (AxCs) employ a search-based Design Space Exploration (DSE) approach. This includes an iterative process where approximate circuit instances are created and evaluated in terms of quality and performance metrics. The quality evaluation of each instance via either verification or testing results in extremely long computation times, imposing a practical limit on the number of nodes that can be explored in the design space. To overcome this problem, we exploit Random Forests (RFs) to develop extremely fast estimators to evaluate the quality and performance of AxC instances while avoiding time-consuming computations. Utilizing these estimators we build LDAX, an efficient design space exploration framework that attempts to improve the runtime of the AxCs synthesis process. LDAX is based on the fact that, in an iterative search space exploration, a large number of validations can be skipped by leveraging a high-accuracy predictor. To efficiently explore the design space, which is usually represented by a tree and each node denotes an AxC, we propose a learning-based search technique that can quickly analyze a large set of nodes even very deep nodes in the tree. Our experimental results reveal that LDAX can achieve average speed-up of 41 × on a set of practical benchmarks from different application domains in comparison with the most competitive state-of-the-art framework.
Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner
ACM Great Lakes Symposium on VLSI2
2020 DeepWind: An Accurate Wind Turbine Condition Monitoring Framework via Deep Learning on Embedded Platforms
abstract
Condition monitoring of critical components in wind turbines is of utmost importance to enable low cost and preventive maintenance and to minimize downtime. In this paper we propose DEEPWIND, an end-to-end condition monitoring and fault detection framework that can be implemented on resource-constrained embedded platforms. DEEPWIND exploits multi-channel convolutional neural networks to automatically extract features from sensor data without any need of human feature engineering, and utilizes the features to classify faults occurring in rotor blades of wind turbines. Experiments on a real-world dataset provided by Weidmüller Monitoring Systems GmbH reveal an average F1-score of 0.94 and thus underline the suitability of the approach for fault detection in wind turbines.
Hassan Ghasemzadeh Mohammadi, Rahil Arshad, Sneha Rautmare, Suraj Manjunatha, Maurice Kuschel, Felix Paul Jentzsch, Marco Platzner, Alexander Boschmann, Dirk Schollbach
ETFA1
2020 A Hybrid Synthesis Methodology for Approximate Circuits
abstract
Automated synthesis of approximate circuits via functional approximations is of prominent importance to provide efficiency in energy, runtime, and chip area required to execute an application. Approximate circuits are usually obtained either through analytical approximation methods leveraging approximate transformations such as bit-width scaling or via iterative search-based optimization methods when a library of approximate components, e.g., approximate adders and multipliers, is available. For the latter, exploring the extremely large design space is challenging in terms of both computations and quality of results. While the combination of both methods can create more room for further approximations, theDesign Space Exploration ~(DSE) becomes a crucial issue. In this paper, we present such a hybrid synthesis methodology that applies a low-cost analytical method followed by parallel stochastic search-based optimization. We address the DSE challenge through efficient pruning of the design space and skipping unnecessary expensive testing and/or verification steps. The experimental results reveal up to 10.57x area savings in comparison with both purely analytical or search-based approaches.
Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner
ACM Great Lakes Symposium on VLSI2
2019 Jump Search: A Fast Technique for the Synthesis of Approximate Circuits
abstract
State-of-the-art frameworks for generating approximate circuits automatically explore the search space in an iterative process - often greedily. Synthesis and verification processes are invoked in each iteration to evaluate the found solutions and to guide the search algorithm. As a result, a large number of approximate circuits is subjected to analysis - leading to long runtimes - but only a few approximate circuits might form an acceptable solution.
Linus Witschen, Hassan Ghasemzadeh Mohammadi, Matthias Artmann, Marco Platzner
ACM Great Lakes Symposium on VLSI2
2018 An MCTS-based Framework for Synthesis of Approximate Circuits
abstract
Approximate computing has become a very popular design strategy that exploits error resilient computations to achieve higher performance and energy efficiency. Automated synthesis of approximate circuits is performed via functional approximation, in which various parts of the target circuit are extensively examined with a library of approximate components/transformations to trade off the functional accuracy and computational budget (i.e., power). However, as the number of possible approximate transformations increases, traditional search techniques suffer from a combinatorial explosion due to the large branching factor. In this work, we present a comprehensive framework for automated synthesis of approximate circuits from either structural or behavioral descriptions. We adapt the Monte Carlo Tree Search (MCTS), as a stochastic search technique, to deal with the large design space exploration, which enables a broader range of potential possible approximations through lightweight random simulations. The proposed framework is able to recognize the design Pareto set even with low computational budgets. Experimental results highlight the capabilities of the proposed synthesis framework by resulting in up to 61.69% energy saving while maintaining the predefined quality constraints.
Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner
VLSI-SoC2
2016 A Fault-Tolerant Ripple-Carry Adder with Controllable-Polarity Transistors
abstract
This article first explores the effects of faults on circuits implemented with controllable-polarity transistors. We propose a new fault model that suits the characteristics of these devices, and we report the results of a SPICE-based analysis of the effects of faults on the behavior of some basic gates implemented with them. Hence, we show that the considered devices are able to intrinsically tolerate a rather high number of faults. We finally exploit this property to build a robust and scalable adder whose area, performance, and leakage power characteristics are improved by 15%, 18%, and 12%;, respectively, when compared to an equivalent FinFET solution at 22nm technology node.
Hassan Ghasemzadeh Mohammadi, Pierre-Emmanuel Gaillardon, Jian Zhang 0067, Giovanni De Micheli, Ernesto Sánchez 0001, Matteo Sonza Reorda
ACM J. Emerg. Technol. Comput. Syst.1
2016 Efficient Statistical Parameter Selection for Nonlinear Modeling of Process/Performance Variation
abstract
With the growing number of process variation (PV) sources in deeply nano-scaled technologies, parameterized device and circuit modeling is becoming very important for chip design and verification. However, the high dimensionality of parameter space, for PV analysis, is a serious modeling challenge for emerging VLSI technologies. These parameters correspond to various interdie and intradie variations, and considerably increase the difficulties of design validation. Today's response surface models and most commonly used parameter reduction methods, such as principal component analysis and independent component analysis, limit parameter reduction to linear or quadratic form and they do not address the higher order of nonlinearity among process and performance parameters. In this paper, we propose and validate a feature selection method to reduce the circuit modeling complexity associated with high parameter dimensionality. This method relies on a learning-based nonlinear sparse regression, and performs a parameter selection in the input space rather than creating a new space. This method is capable of dealing with mixed Gaussian and non-Gaussian parameters and results in a more precise parameter selection considering statistical nonlinear dependencies among input and output parameters. The application of this method is demonstrated in digital circuit timing analysis in both FinFET and Silicon Nanowire technologies. The results confirm the efficiency of this method to significantly reduce the number of required simulations while keeping estimation error small.
Hassan Ghasemzadeh Mohammadi, Pierre-Emmanuel Gaillardon, Giovanni De Micheli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2015 Fault modeling in controllable polarity silicon nanowire circuits
Hassan Ghasemzadeh Mohammadi, Pierre-Emmanuel Gaillardon, Giovanni De Micheli
DATE1
2010 Sub-threshold charge recovery circuits
abstract
Embedded systems account for wide range of applications. However, the design of such systems is faced with a diverse spectrum of criteria. The energy consumption, performance, and demanding security concerns are some of the most significant challenges in designing of such systems. With these challenges, the design process can be managed more easily if a flexible logic circuit with the ability of satisfying the above-mentioned concerns is taken into account. To achieve such a logic circuit, in this paper we have combined the sub-threshold operation and charge recovery techniques. Using our technique, lower power consumption, ability of operating at higher frequencies, and more security (to side channel attacks) than the existing logic circuits are achieved. This paper also presents an analytical proof about how sub-threshold charge recovery circuits can meet these characterizations. we have also confirmed our analytical discussions by examining our technique for full adder, and 8 × 8 carry-save multiplier in different frequencies, supply voltages, and CMOS technologies. Detailed SPICE simulations show significant improvements as compared to its existing counterparts in all simulated frequencies and supply voltages.
Mehrdad Khatir, Hassan Ghasemzadeh Mohammadi, Alireza Ejlali
ICCD2