VLDB 2026 Research / reviewers in the wild / expert
Muhammad Awais 0009
dblp:80/3639-9
· DBLP profile ↗
9ranked-venue papers
7as first author
7since 2021 · last 2026
0000-0003-4148-2969ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 7 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DP-MCTS: Deep Playout-Driven MCTS for Approximate Accelerator DesignabstractApproximate computing offers substantial efficiency gains for modern accelerators, yet register-transfer level design space exploration (DSE) remains challenging due to the exponential growth of approximation choices. Although recent use of machine learning (ML)-based accuracy estimators has reduced exploration runtime, it introduces a new challenge: estimator inaccuracy, which becomes increasingly problematic for larger circuits. In this paper, we propose DP-MCTS, an enhanced Monte Carlo Tree Search framework that rethinks estimator usage by treating ML models as lightweight guidance tools rather than decision makers. DP-MCTS employs deep playouts to provide fast, informative look-ahead along candidate paths, coupled with a promotion mechanism that selectively validates promising nodes through simulation-based accuracy evaluation. This synergy preserves broad design space coverage while mitigating estimator-driven drift, significantly improving exploration robustness and solution quality. Experiments across diverse accelerator benchmarks show that DP-MCTS closely matches or even outperforms the solution quality of an MCTS-based search, achieving up to 6.0% additional area reduction and an order of magnitude lower runtime cost, thereby delivering a substantially improved quality-runtime tradeoff over existing search-based frameworks. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Sayed Morteza Jawadi, Marco Platzner |
DDECS | 1 |
| 2026 | Reliability Assessment in Approximate Accelerator SynthesisabstractWhile optimizing for core hardware performancerelated target metrics, frameworks for approximate accelerators often overlook the reliability aspect. Approximated implementations obtained by these frameworks can potentially differ in terms of reliability and may impact the reliability of the overall system. In particular, approximation changes the data profiles transmitted between system modules, which can trigger crosstalk on interconnect lines and aggravate electromigration. We propose a two-stage process that performs a reliability assessment of the circuit interconnects after the approximate accelerator synthesis. Our approach aims to find the most reliable solutions from the approximate candidate circuits generated by an automated approximation flow. We then leverage Pareto-filtering to strike a balance between area, reliability, and accuracy. Notably, the selected designs achieve up to a 178% improvement in mission time compared to the original accelerator, and a 68% improvement over designs optimized solely for area. In addition, our methodology allows custom priority settings to be adaptable to a user's preference, thereby leading to circuits that meet diverse design constraints. Our experimental results show the effectiveness of our methodology in achieving superior trade-offs between area, reliability, and accuracy, hence uncovering a new dimension for approximate accelerator design methodologies. Somayeh Sadeghi Kohan, Muhammad Awais 0009, Qazi Arbab Ahmed, Marco Platzner, Sybille Hellebrand, Thorsten Jungeblut, Hans-Joachim Wunderlich |
DDECS | 2 |
| 2025 | A Two-Stage Approximation Methodology for Efficient DNN Hardware ImplementationabstractDeploying high-performance Deep Neural Networks (DNNs) on embedded hardware platforms, such as FPGAs, not only promises low latency and high energy-efficiency implementations, but also poses significant challenges due to the resource constraints of such systems. Recent efforts have focused on approximation techniques at the algorithmic level of DNNs to reduce resource requirements, such as network pruning and quantization of weights and activations. Other approximation techniques that approximate register transfer level (RTL) designs by substituting single components with approximated ones are not yet applied to DNNs, mainly because these techniques do not scale well to the typically huge RTL representations of DNNs.This paper introduces the idea of a two-stage approximation methodology for DNNs that combines algorithmic level with RTL approximation to achieve reduced resource requirements at acceptable accuracy. As a novel instance of that idea, we leverage the LogicNets approach that performs a low-bit quantization to map DNN neurons to LUT representations and the CIRCA framework that provides a search-based approximation flow for RTL designs. To balance hardware efficiency and accuracy, and to achieve acceptable runtimes, we limit the RTL approximation to a subset of neurons identified as resilient using a systematic sampling approach. Experimental results demonstrate the potential of the methodology with up to 15.20% further decrease in resource utilization with a mere 2.15% accuracy drop. Amir Hossein Hadipour, Atousa Jafari, Muhammad Awais 0009, Marco Platzner |
DDECS | 3 |
| 2025 | Swift Synthesis of Approximate Hardware Accelerators Using Generative Adversarial NetworksabstractDeploying modern applications with significant resource demands is often challenging, but approximate designs offer a promising alternative by delivering high performance with minimal compromises in output quality. Traditionally, approximate hardware accelerators have been developed through search-based iterative frameworks, which suffer from long runtimes due to the exponential growth of the design space. A significant portion of the runtime is consumed by either invalid nodes or valid nodes that offer minimal improvements in performance metrics, such as runtime or power consumption. This severely limits the thorough exploration of the design space. In this paper, we introduce a novel approach for synthesizing approximate accelerators that leverages sparsity to reduce the complexity of the design space exploration problem. Our method employs a generative adversarial network (GAN) to rapidly generate a diverse set of high-quality design nodes, eliminating the need for costly node evaluations. This enables the swift creation of approximate accelerators generated for any given error threshold in a fraction of time as compared to a simulation-based framework. We conducted experiments on a suite of benchmarks from real-world domains, demonstrating that our methodology can generate approximate hardware designs with significant area and power savings, comparable to state-of-the-art search-based approaches. In a comparative evaluation against two leading methods, our approach achieved equal or better quality results for two out of four benchmarks while reaching up to 55% area savings, thus effectively demonstrating a new avenue for automated generation of approximate accelerators. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
VLSI-SoC | 1 |
| 2025 | Design Space Exploration for Approximate Circuits via Checkpointing and DNN-Based EstimatorsabstractApproximate computing (AC) is a design paradigm that trades in reductions in hardware area, power, and delay of digital circuits for an increased error. Workflows for synthesizing approximate accelerators are often implemented as search-approximate–evaluate cycles that build a search tree by iteratively approximating components of the original circuit. Design space exploration (DSE) for approximate accelerators is challenging due to the sheer size of the design space and the computational costs for validating the error bound during the search. While early work employed rather slow simulation-based methods, more recent approaches use machine learning (ML) to create fast estimators for validating error bounds. For larger accelerators and design spaces, however, the mispredictions of ML-based methods can steer the search process to explore invalid regions of the design space. In this article, we introduce the frameworkDeepApproxfor approximate accelerator synthesis. The novelty of this framework is its combination of fast ML-based quality estimators with occasional simulations to correct possible mispredictions. The simulations are performed at the so-called checkpoints during the search, and we propose two methods: static and dynamic checkpointing (DC). We present theDeepApproxflow and elaborate on the ML quality estimators and the checkpointing mechanisms. Using a set of smaller and larger benchmarks, we compareDeepApproxwith two state-of-the-art flows: a simulation-based method leveraging Monte Carlo tree search (MCTS) and an ML-based method with fast estimators. Our experiments demonstrate thatDeepApproxcan achieve similar reductions in hardware area and power consumption compared to simulation-based methods, albeit at much lower runtimes, and higher reductions in these metrics compared to methods relying exclusively on ML-based estimators. Thus, by combining ML-based techniques with simulations,DeepApproxprovides a scalable DSE for approximate accelerators. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2022 | Automated Framework for Fast Synthesis of Approximate Hardware AcceleratorsabstractGenerating approximate accelerators automatically via libraries of functional units faces combinatorial explosion due to extremely large design space. Moreover, long verification times become bottleneck in the process making it nearly impossible to find best suited approximate instances with exhaustive search. Multiple works try to explore the design space with a greedy-based approach that reduces the complexity of the process, albeit overlooking promising combinations and ultimately producing inferior solutions. This paper proposes an automated framework that handles the design space exploration problem with an learning-based search algorithm i.e., MCTS. Furthermore, by leveraging fast machine learning quality estimators, the framework offers up to 47× reduction of runtime when compared to a simulation-based framework. Muhammad Awais 0009, Marco Platzner |
VLSI-SoC | 1 |
| 2021 | LDAX: A Learning-based Fast Design Space Exploration Framework for Approximate Circuit SynthesisabstractThe majority of existing frameworks for automated synthesis of Approximate Circuits (AxCs) employ a search-based Design Space Exploration (DSE) approach. This includes an iterative process where approximate circuit instances are created and evaluated in terms of quality and performance metrics. The quality evaluation of each instance via either verification or testing results in extremely long computation times, imposing a practical limit on the number of nodes that can be explored in the design space. To overcome this problem, we exploit Random Forests (RFs) to develop extremely fast estimators to evaluate the quality and performance of AxC instances while avoiding time-consuming computations. Utilizing these estimators we build LDAX, an efficient design space exploration framework that attempts to improve the runtime of the AxCs synthesis process. LDAX is based on the fact that, in an iterative search space exploration, a large number of validations can be skipped by leveraging a high-accuracy predictor. To efficiently explore the design space, which is usually represented by a tree and each node denotes an AxC, we propose a learning-based search technique that can quickly analyze a large set of nodes even very deep nodes in the tree. Our experimental results reveal that LDAX can achieve average speed-up of 41 × on a set of practical benchmarks from different application domains in comparison with the most competitive state-of-the-art framework. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | A Hybrid Synthesis Methodology for Approximate CircuitsabstractAutomated synthesis of approximate circuits via functional approximations is of prominent importance to provide efficiency in energy, runtime, and chip area required to execute an application. Approximate circuits are usually obtained either through analytical approximation methods leveraging approximate transformations such as bit-width scaling or via iterative search-based optimization methods when a library of approximate components, e.g., approximate adders and multipliers, is available. For the latter, exploring the extremely large design space is challenging in terms of both computations and quality of results. While the combination of both methods can create more room for further approximations, theDesign Space Exploration ~(DSE) becomes a crucial issue. In this paper, we present such a hybrid synthesis methodology that applies a low-cost analytical method followed by parallel stochastic search-based optimization. We address the DSE challenge through efficient pruning of the design space and skipping unnecessary expensive testing and/or verification steps. The experimental results reveal up to 10.57x area savings in comparison with both purely analytical or search-based approaches. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
ACM Great Lakes Symposium on VLSI | 1 |
| 2018 | An MCTS-based Framework for Synthesis of Approximate CircuitsabstractApproximate computing has become a very popular design strategy that exploits error resilient computations to achieve higher performance and energy efficiency. Automated synthesis of approximate circuits is performed via functional approximation, in which various parts of the target circuit are extensively examined with a library of approximate components/transformations to trade off the functional accuracy and computational budget (i.e., power). However, as the number of possible approximate transformations increases, traditional search techniques suffer from a combinatorial explosion due to the large branching factor. In this work, we present a comprehensive framework for automated synthesis of approximate circuits from either structural or behavioral descriptions. We adapt the Monte Carlo Tree Search (MCTS), as a stochastic search technique, to deal with the large design space exploration, which enables a broader range of potential possible approximations through lightweight random simulations. The proposed framework is able to recognize the design Pareto set even with low computational budgets. Experimental results highlight the capabilities of the proposed synthesis framework by resulting in up to 61.69% energy saving while maintaining the predefined quality constraints. Muhammad Awais 0009, Hassan Ghasemzadeh Mohammadi, Marco Platzner |
VLSI-SoC | 1 |