EDBT 2026 Demo / reviewers in the wild / expert
Salim Ullah
dblp:218/1191
· DBLP profile ↗
32ranked-venue papers
12as first author
24since 2021 · last 2026
0000-0002-9774-9522ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 12 first-author · 24 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RESIST : Structured Regularization from Weight Similarity to Weight Diversity for Improving Error Resilience of Neural Networks
Maryam Eslami, Salim Ullah, Akash Kumar 0001 |
IOLTS | 2 |
| 2025 | Special Sessions - Emerging Scope and Design Challenges for Approximate Computing: Optimizing Accuracy-PPA trade-offs and BeyondabstractThe rapid growth of AI workloads is driving interest in Approximate Computing (AxC) as a means to enable low-cost, energy-efficient inference in resource-constrained systems. By introducing controlled inaccuracies, AxC can deliver substantial gains in power, performance, and area (PPA) while leveraging the inherent error tolerance of many AI models. Achieving this potential requires adapting existing frameworks to support the design and optimization of neural networks with approximate operators. Modern AxC research extends beyond accuracy-PPA trade-offs to address reliability and security, reducing redundancy overheads and exploring the distinctive side-channel implications of approximation. Application-aware approaches, such as those for spiking neural networks, show that tailoring approximation to workload-specific error behavior can surpass generic strategies. This article examines AI-guided design methods and the interplay between efficiency, reliability, and security, highlighting how these interconnected facets can advance embedded and high-performance computing. Siva Satyendra Sahoo, Bastien Deveautour, Marcello Traiola, Chongyan Gu, Yun Wu 0003, Aditya Japa, Salim Ullah, Akash Kumar 0001 |
CASES | 7 |
| 2025 | Design, Model, and Explore Approximate Arithmetic Operators with AI/ML: A TutorialabstractApproximate Computing (AxC) is being actively explored to meet the energy and performance requirements of resource-constrained embedded systems. Approximate arithmetic operators (AxOs), for instance, let edge-AI systems trade tiny, bounded errors for big wins in power, performance, and area. This tutorial demystifies AxO design, modeling, and exploration: from platform-aware operator synthesis (e.g., selective LUT pruning) to application-specific DSE that uses AI/ML to navigate massive trade-off spaces. We contrast selection (library) vs. synthesis (generate-and-optimize) flows, show when FPGA-aware adders/multipliers outperform ASIC-ported designs, and connect operator-level error to task-level metrics (e.g., Conv2D, MLP). The tutorial includes hands-on Jupyter notebooks, ready-to-reuse operator models, and a practical recipe for building Pareto-optimal AxOs under accuracy constraints - plus a peek at AxOSyn, an open-source framework that unifies selection/synthesis, surrogate fitness, and search using evolutionary algorithms. Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001 |
CASES | 1 |
| 2025 | BiKA: Binarized KAN-inspired Neural Network for Efficient Hardware Accelerator DesignsabstractThe continuously growing size of Neural Network (NN) models makes the design of lightweight neural network accelerators for edge devices an emerging subject in recent research. Previous works explored different lightweight technologies or even emerging neural network structures, such as quantization, approximate computing, neuromorphic computing, etc., to reduce hardware resource consumption in accelerator designs. This inspired our interest in exploring the potential of other emerging network structures in hardware accelerator designs. Kolmogorov-Arnold Network (KAN) [1] is a recently proposed novel neural network structure by replacing the multiplication and activation function in Artificial Neural Networks (ANN) with learnable nonlinear functions, which has the potential to transform the paradigm of neural network design. However, considering the complexity of the nonlinear function on hardware, the design of the lightweight hardware accelerator of KAN lacks thoroughly related research. Salim Ullah, Akash Kumar 0001 |
FCCM | 2 |
| 2025 | Invited Paper: Circuit and Architecture Design with Emerging Computing ParadigmsabstractAs emerging computing paradigms push beyond the limitations of traditional CMOS-based computing using Von Neumann architectures, there is a growing need to rethink and extend Electronic Design Automation (EDA) methodologies to support their unique characteristics. These paradigms—including Approximate Computing, In-Memory Computing, Reconfigurable Field-Effect Transistors (RFETs), and Photonic Computing—represent diverse and promising directions beyond conventional digital design. Collectively, they offer transformative potential for achieving significant improvements in energy efficiency, computational speed, and architectural scalability. For example, application-specific approximate computing enables the design of custom arithmetic circuits that exploit application-level error resilience, allowing for optimized accuracy–power–performance–area (PPA) trade-offs in error-tolerant applications. Similarly, processing-in-non-volatile memories, such as those based on Ferroelectric Field-effect Transistors (FeFETs), enhances energy efficiency by enabling analog computation—particularly for operations like matrix multiplication—directly within the memory arrays. The intrinsic polymorphism of RFETs supports compact, multifunctional logic gates and introduces new opportunities for circuit-level obfuscation and security-aware design. Likewise, photonic analog wavefront computing offers substantial gains in latency and energy efficiency by encoding and processing information in the analog optical domain, leveraging phenomena such as diffraction and interference to perform computation at the speed of light. However, they also introduce a host of new challenges in circuit and architecture design, such as vast and irregular design spaces, analog and non-Boolean behavior, and new device-level constraints that existing EDA tools are not capable of handling. To this end, the current article focuses on the development of efficient and robust EDA frameworks that can enable the practical realization of circuits and architectures in these emerging domains. Salim Ullah, Siva Satyendra Sahoo, Can Li 0024, Chao Li 0065, Liu Liu 0023, Tomas Sousa Pereira, Xunzhao Yin, Armin Darjani, Nima Kavand, Chakravarthy Bodla, Rupa Yashaswi Panduga, Aniruddh Holemadlu, Johannes Maly, Jonathan Förste, Samarth Vadia, Xiaobo Sharon Hu, Akash Kumar 0001 |
ICCAD | 1 |
| 2024 | Enabling Energy-efficient AI Computing: Leveraging Application-specific Approximations : (Education Class)abstractThe widespread adoption of Artificial intelligence and Machine Learning (AI/ML) models across various fields, such as healthcare, autonomous vehicles, smart agriculture, and industrial automation, has led to a growing demand for efficient and scalable AI/ML solutions. However, as AI/ML algorithms grow more complex, their substantial memory requirements and high energy consumption pose significant challenges for deployment on resource-constrained embedded systems, such as wearable health monitors and IoT devices. To this end, various techniques, such as model pruning, knowledge distillation, quantization of model parameters, and employing approximate arithmetic operators, are commonly explored to overcome these challenges [1] . Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001 |
CASES | 1 |
| 2024 | LeQC-At: Learning Quantization Configurations During Adversarial Training for Robust Deep Neural NetworksabstractDue to the high feature learning capability of Deep Neural Networks (DNNs), they are widely used in state-of-the-art machine learning tasks such as image and text recognition and natural language processing. However, in the recent developments of deep learning models, it has been observed that specially crafted inputs (adversarial samples) can deceive DNNs and result in incorrect predictions with high confidence. Such incorrect predictions by adversarially attacked DNNs can have devastating results in safety-critical applications. This situation can be further exacerbated in quantized DNNs, which employ reduced precision numbers to reduce the overall computational complexity of DNNs. To this end, this work proposes a framework for the joint optimization of DNNs to improve the natural accuracy of quantized DNNs and increase the robustness of the quantized models against adversarial attacks. In particular, the proposed framework employs quantization step size aware adversarial training of DNNs. Our proposed framework is generic and can be utilized with any quantization scheme that allows learning of quantization configurations during training. Furthermore, we present a novel loss function for adversarial training to improve the quantized networks' accuracy. For example, our 3-bit quantized adversarial training of ResNet-18, ResNet-34, and WideResNet shows up to 21.63%, 24.49%, and 15.08% higher inference accuracy with attacked data, respectively, when compared to 3-bit vanilla quantized adversarial training on benchmark datasets. Siddharth Gupta 0004, Salim Ullah, Akash Kumar 0001 |
DSD | 2 |
| 2024 | BitSys: Bitwise Systolic Array Architecture for Multi-precision Quantized Hardware AcceleratorsabstractQuantized Neural Networks (QNN) have been widely applied in hardware accelerator designs for edge. Because lower precision in quantization leads to higher accuracy loss, the mixed-precision scheme has been explored by using different precision in different layers to trade off resource consumption and inference accuracy. Because regular multiplier designs do not support the reconfiguration for multi-precision, we explored a runtime reconfigurable multi-precision bitwise systolic array design, BitSys, for mixed-precision multiplication in QNN accelerators. The popular design in previous works, such as [1], divides the inputs of multipliers as two parts for four sub-multipliers and preset left shifting, achieving the reconfigurable multiplication by disabling two of the submultipliers. Our design is inspired by the Bitshifter architecture from the works of Liu et al. [2], [1]. We convert the$n\times n$- bit multiplication as$A \times B=\sum_{i=0}^{n-1} \sum_{j=0}^{n-1} 2^{i+j} a_i b_j$. As shown in Figure 1, partial products$P_{i+j}$is the sum of subpartial products$a_{i}b_{j}$with left shifting value,$i+j$. The subpartial product masks select the desired$a_{i}b_{j}$to configure the multi-channel according to the precision. One mask square represents one channel. Therefore, we can implement a bitwise systolic array as Figure 2 (left). The sub-partial product computation and mask are fused in one LUT primitive as one processing element in Figure 2 (right up). The sums of elements, considered the sign-bits, in the diagonal with the same left shifting value shown in Figure 2 (left) are the inputs,$D_{k}$, of the output-generate pipeline of Figure 2 (right down). Systolic array sequentially generates the$D_{k}$, and the final multi-precision output is the sum of all$D_{k}$. We implemented our BitSys multiplier as a systolic array for mixed-precision QNN acceleration. The comparison with the works of Liu et al. [1] is shown in Table I. Our systolic array accelerator consumes more hardware resources than the three single-layer accelerator instances of Liu et al. [1]. However, because of our bitwise processing design, BitSys instance can support 250MHz clock frequency because of the low critical path delay of our multiplier and achieves 188.5-274.7% speed-up in the evaluation of one four-layer 1/2/4/8-bit mixed-precision quantized MLP. Furthermore, our design does not change the input/out width when configuring to different precision, which can be easily integrated into existing accelerator designs. Salim Ullah, Akash Kumar 0001 |
FCCM | 2 |
| 2024 | AxOSpike: Spiking Neural Networks-Driven Approximate Operator DesignabstractApproximate computing (AxC) is being widely researched as a viable approach to deploying compute-intensive artificial intelligence (AI) applications on resource-constrained embedded systems. In general, AxC aims to provide disproportionate gains in system-level power-performance-area (PPA) by leveraging the implicit error tolerance of an application. One of the more widely used methods in AxC involves circuit pruning of arithmetic operators used to process AI workloads. However, most related works adopt an application-agnostic approach to operator modeling for the design space exploration (DSE) of Approximate Operators (AxOs). To this end, we propose an application-driven approach to designing AxOs. Specifically, we use spiking neural network (SNN)-based inference to present an application-driven operator model resulting in AxOs with better-PPA-accuracy tradeoffs compared to traditional circuit pruning. Additionally, we present a novel FPGA-specific operator model to improve the quality of AxOs that can be obtained using circuit pruning. With the proposed methods, we report designs with up to 26.5% lower PDPxLUTs with similar application-level accuracy. Further, we report a considerably better set of design points than related works with up to 51% better-Pareto front hypervolume. Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | AxOCS: Scaling FPGA-Based Approximate Operators Using Configuration SupersamplingabstractThe rising usage of AI/ML-based processing across application domains has exacerbated the need for low-cost ML implementation, specifically for resource-constrained embedded systems. To this end, approximate computing, an approach that explores the power, performance, area (PPA), and behavioral accuracy (BEHAV) trade-offs, has emerged as a possible solution for implementing embedded machine learning. Due to the predominance of MAC operations in ML, designing platform-specific approximate arithmetic operators forms one of the major research problems in approximate computing. Recently, there has been a rising usage of AI/ML-based design space exploration techniques for implementing approximate operators. However, most of these approaches are limited to using ML-based surrogate functions for predicting the PPA and BEHAV impact of a set of related design decisions. While this approach leverages the regression capabilities of ML methods, it does not exploit the more advanced approaches in ML. To this end, we propose, a methodology for designing approximate arithmetic operators through ML-based supersampling. Specifically, we present a method to leverage the correlation of PPA and BEHAV metrics across operators of varying bit-widths for generating larger bit-width operators. The proposed approach involves traversing the relatively smaller design space of smaller bit-width operators and employing its associatedDesign-PPA-BEHAVrelationship to generate initial solutions for metaheuristics-based optimization for larger operators. The experimental evaluation of for FPGA-optimized approximate operators shows that the proposed approach significantly improves the quality—resulting hypervolume for multi-objective optimization—of$8\times8$signed approximate multipliers. Siva Satyendra Sahoo, Salim Ullah, Soumyo Bhattacharjee, Akash Kumar 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | AxOMaP: Designing FPGA-based Approximate Arithmetic Operators using Mathematical ProgrammingabstractWith the increasing application of machine learning (ML) algorithms in embedded systems, there is a rising necessity to design low-cost computer arithmetic for these resource-constrained systems. As a result, emerging models of computation, such as approximate and stochastic computing, that leverage the inherent error-resilience of such algorithms are being actively explored for implementing ML inference on resource-constrained systems. Approximate computing (AxC) aims to provide disproportionate gains in the power, performance, and area (PPA) of an application by allowing some level of reduction in its behavioral accuracy (BEHAV). Using approximate operators (AxOs) for computer arithmetic forms one of the more prevalent methods of implementing AxC. AxOs provide the additional scope for finer granularity of optimization, compared to only precision scaling of computer arithmetic. To this end, the design of platform-specific and cost-efficient approximate operators forms an important research goal. Recently, multiple works have reported the use of AI/ML-based approaches for synthesizing novel FPGA-based AxOs. However, most of such works limit the use of AI/ML to designing ML-based surrogate functions that are used during iterative optimization processes. To this end, we propose a novel data analysis-driven mathematical programming-based approach to synthesizing approximate operators for FPGAs. Specifically, we formulate mixed integer quadratically constrained programs based on the results of correlation analysis of the characterization data and use the solutions to enable a more directed search approach for evolutionary optimization algorithms. Compared to traditional evolutionary algorithms-based optimization, we report up to 21% improvement in the hypervolume, for joint optimization of PPA and BEHAV, in the design of signed 8-bit multipliers. Further, we report up to 27% better hypervolume than other state-of-the-art approaches to DSE for FPGA-based application-specific AxOs. Siva Satyendra Sahoo, Salim Ullah, Akash Kumar 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2023 | SyFAxO-GeN: Synthesizing FPGA-Based Approximate Operators with Generative NetworksabstractWith rising trends of moving AI inference to the edge, due to communication and privacy challenges, there has been a growing focus on designing low-cost Edge-AI. Given the diversity of application areas at the edge, FPGA-based systems are increasingly used for high-performance inference. Similarly, approximate computing has emerged as a viable approach to achieve disproportionate resource gains by utilizing the applications' inherent robustness. However, most related research has focused on selecting the appropriate approximate operators for an application from a set of ASIC-based designs. This approach fails to leverage the FPGA's architectural benefits and limits the scope of approximation to already existing generic designs. To this end, we propose an AI-based approach to synthesizing novel approximate operators for FPGA's Look-up-table-based structure. Specifically, we use state-of-the-art generative networks to search for constraint-aware arithmetic operator designs optimized for FPGA-based implementation. With the proposed GANs, we report up to 49% faster training, with negligible accuracy degradation, than related generative networks. Similarly, we report improved hypervolume and increased pareto-front design points compared to state-of-the-art approaches to synthesizing approximate multipliers. Rohit Ranjan, Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001 |
ASP-DAC | 2 |
| 2023 | CoOAx: Correlation-aware Synthesis of FPGA-based Approximate OperatorsabstractThe run-time reconfigurability and high parallelism offered by Field Programmable Gate Arrays (FPGAs) make them an attractive choice for implementing hardware accelerators for Machine Learning (ML) algorithms. In the quest for designing efficient FPGA-based hard-ware accelerators for ML algorithms, the inherent error-resilience of ML algorithms can be exploited to implement approximate hard-ware accelerators to trade the output accuracy with better over-all performance. As multiplication and addition are the two main arithmetic operations in ML algorithms, most state-of-the-art approximate accelerators have considered approximate architectures for these operations. However, these works have mainly considered the exploration and selection of approximate operators from an existing set of operators. To this end, we provide an efficient methodology for synthesizing and implementing novel approximate operators. Specifically, we propose a novel operator synthesis approach that supports multiple operator algorithms to provide new approximate multiplier and adder designs for AI inference applications. We report up to 27% and 25% lower power than state-of-the-art approximate designs, with equivalent error behavior, for 8-bit unsigned adders and 4-bit signed multipliers respectively. Further, we propose a correlation-aware Design Space Exploration (DSE) method that can improve the efficacy of randomized search algorithms in synthesizing novel approximate operators. Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2023 | AxOTreeS: A Tree Search Approach to Synthesizing FPGA-based Approximate OperatorsabstractApproximate computing (AxC) provides the scope for achieving disproportionate gains in a system’s power, performance, and area (PPA) metrics by leveraging an application’s inherent error-resilient behavior (BEHAV). Trading computational accuracy for performance gains makes AxC an attractive proposition for implementing computationally complex AI/ML-based applications on resource-constrained embedded systems. The growing diversity of application domains using AI/ML has also led to the increasing usage of FPGA-based embedded systems. However, implementing AxC for FPGAs has primarily been limited to the post-processing of ASIC-optimized approximate operators (AxOs). This approach usually involves selecting from a set of AxOs that have been optimized for a gate-based implementation in an ASIC. While such an approach does allow leveraging existing knowledge of ASIC-based AxO design, it limits the scope for considering the challenges and opportunities associated with FPGA’s LUT-based computation structures. Similarly, the few works considering the LUT-based computing for AxO design use generic optimization approaches that do not allow integrating problem-specific prior knowledge—empirical and/or statistical. To this end, we propose a novel tree search-based approach to AxO synthesis for FPGAs. Specifically, we present a design methodology using Monte Carlo Tree Search (MCTS)-based search tree traversal that allows the designer to integrate statistical data, such as correlation, into the AxOs optimization. With the proposed methods, we report improvements over standard MCTS algorithm-based results as well as improved hypervolume for both operator-level and application-specific DSE, compared to state-of-the-art design methodologies. Siva Satyendra Sahoo, Salim Ullah, Akash Kumar 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | Multi-Precision Deep Neural Network Acceleration on FPGAsabstractQuantization is a promising approach to reduce the computational load of neural networks. The minimum bit-width that preserves the original accuracy varies significantly across different neural networks and even across different layers of a single neural network. Most existing designs over-provision neural network accelerators with sufficient bit-width to preserve the required accuracy across a wide range of neural networks. In this paper, we present mpDNN, a multi-precision multiplier with dynamically adjustable bit-width for deep neural network acceleration. The design supports run-time splitting an arithmetic operator into multiple independent operators with smaller bit-width, effectively increasing throughput when lower precision is required. The proposed architecture is designed for FPGAs, in that the multipliers and bit-width adjustment mechanism are optimized for the LUT-based structure of FPGAs. Experimental results show that by enabling run-time precision adjustment, mpDNN can offer 3-15x improvement in throughput. Negar Neda, Salim Ullah, Azam Ghanbari, Hoda Mahdiani, Mehdi Modarressi, Akash Kumar 0001 |
ASP-DAC | 2 |
| 2022 | PosAx-O: Exploring Operator-level Approximations for Posit Arithmetic in Embedded AI/MLabstractThe quest for low-cost embedded AI/ML applications has motivated innovations across multiple abstractions of the computation stack. Novel approaches for arithmetic operations have primarily involved quantization, precision-scaling, approximations, and modified data representation. In this context, Posit has emerged as an alternative to the IEEE-754 standard as it offers multiple benefits, primarily due to its dynamic range and tapered precision. However, the implementation of Posit arithmetic operations tends to result in high resource utilization and power dissipation. Consequently, recent works have delved into the idea of exploiting the error resilience of machine learning algorithms by using low-precision Posit arithmetic. However, limiting the exploration to precision-scaling limits the scope for application-specific optimizations for embedded AI/ML applications. To this end, we explore operator-level optimizations and approximations for low-precision Posit numbers. Specifically, we identify and eliminate redundant operations in state-of-the-art Posit arithmetic operator designs and provide a modular framework for exploring approximations in various stages of the computation. We also present a novel framework for behaviorally testing the corresponding Posit approximate designs in Artificial Neural Networks. The proposed optimizations and approximations exhibit considerable resource improvements with a small error in many cases. For instance, a Posit-based multiplier with 1-bit reduced precision shows a 33% improvement in power and utilization, with only a 0.2% degradation in overall accuracy. Amritha Immaneni, Salim Ullah, Suresh Nambi, Siva Satyendra Sahoo, Akash Kumar 0001 |
DSD | 2 |
| 2022 | ERMES: Efficient Racetrack Memory Emulation System based on FPGAabstractWith the scaling of CMOS technology almost over, non-volatile memories based on emerging technologies are gaining considerable popularity. Particularly, spintronic-based Racetrack memories (RTMs) exhibit unprecedented storage capacity, as well as reduced energy per operation and high write endurance, which make them promising candidates to revolutionize the architecture of memory sub-systems. However, since RTM exploits shifting of magnetic domains to align the required data with the access port, its read/write latency is not constant. Due to this behaviour, several performance optimizations related to the target application may be introduced either on memory architecture or data placement or both. To this purpose, specific tools able to emulate the timing characteristics of RTMs are highly desired. Unfortunately, existing software-based simulators show poor flexibility and run-time. To address such limitations, this paper presents a new emulation system for RTMs based on heterogeneous FPGA-CPU Systems-on-Chips (SoCs). Thanks to its high flexibility, the proposed emulator can be easily configured to evaluate different memory architectures. In addition, the CPU can be used to stimulate the RTM architecture under test with appropriate benchmarks, thus providing a fast self-contained evaluation environment. As case study, ERMES has been implemented within the Xilinx Zynq Ultrascale XCUZ9EG SoC to evaluate performances of several memory configurations when running benchmark applications from the MiBench suite, experiencing a speed-up higher than × 146 over software-based simulators. Fanny Spagnolo, Salim Ullah, Pasquale Corsonello, Akash Kumar 0001 |
FPL | 2 |
| 2022 | NetPU: Prototyping a Generic Reconfigurable Neural Network Accelerator ArchitectureabstractFPGA-based Neural Network (NN) accelerator is a rapidly advancing subject in recent research. Related works can be classified as two hardware architectures: i) Heterogeneous Streaming Dataflow (HSD) architecture and ii) Processing Element Matrix (PEM) architecture. HSD architecture explores the reconfigurability of FPGAs to support the customization and optimization of hardware design to implement a complete network on FPGA for one given trained model. PEM architecture achieves relatively generic support for different network models, essentially implementing the neuron processing modules on the FPGA scheduled by the runtime software environment. In summary, the HSD architecture requires more resources with simplified runtime software control. The PEM architecture consumes fewer resources than the HSD architecture. However, the runtime software environment can be a heavy payload for lightweight systems, such as the low-power microcontroller of IoT or edge devices. Shubham Rai, Salim Ullah, Akash Kumar 0001 |
FPT | 3 |
| 2022 | High-Performance Accurate and Approximate Multipliers for FPGA-Based Hardware AcceleratorsabstractMultiplication is one of the widely used arithmetic operations in a variety of applications, such as image/video processing and machine learning. FPGA vendors provide high-performance multipliers in the form of DSP blocks. These multipliers are not only limited in number and have fixed locations on FPGAs but can also create additional routing delays and may prove inefficient for smaller bit-width multiplications. Therefore, FPGA vendors additionally provide optimized soft IP cores for multiplication. However, in this work, we advocate that these soft multiplier IP cores for FPGAs still need better designs to provide high-performance and resource efficiency. Toward this, we present generic area-optimized, low-latency accurate, and approximate softcore multiplier architectures, which exploit the underlying architectural features of FPGAs, i.e., lookup table (LUT) structures and fast-carry chains to reduce the overall critical path delay (CPD) and resource utilization of multipliers. Compared to Xilinx multiplier LogiCORE IP, our proposed unsigned and signed accurate architecture provides up to 25% and 53% reduction in LUT utilization, respectively, for different sizes of multipliers. Moreover, with our unsigned approximate multiplier architectures, a reduction of up to 51% in the CPD can be achieved with an insignificant loss in output accuracy when compared with the LogiCORE IP. For illustration, we have deployed the proposed multiplier architecture in accelerators used in image and video applications, and evaluated them for area and performance gains. Our library of accurate and approximate multipliers is opensource and available online athttps://cfaed.tu-dresden.de/pd-downloadsto fuel further research and development in this area, facilitate reproducible research, and thereby enabling a new research direction for the FPGA community. Salim Ullah, Semeen Rehman, Muhammad Shafique 0001, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | AppAxO: Designing Application-specific Approximate Operators for FPGA-based Embedded SystemsabstractApproximate arithmetic operators, such as adders and multipliers, are increasingly used to satisfy the energy and performance requirements of resource-constrained embedded systems. However, most of the available approximate operators have an application-agnostic design methodology, and the efficacy of these operators can only be evaluated by employing them in the applications. Furthermore, the various available libraries of approximate operators do not share any standard approximation-induction policy to design new operators according to an application’s accuracy and performance constraints. These limitations also hinder the utilization of machine learning models to explore and determine approximate operators according to an application’s requirements. In this work, we present a generic design methodology for implementing FPGA-based application-specific approximate arithmetic operators. Our proposed technique utilizes lookup tables and carry-chains of FPGAs to implement approximate operators according to the input configurations. For instance, for an \( \text{M}\times \text{N} \) accurate multiplier utilizing K lookup tables, our methodology utilizes K -bit configurations to design \( 2^K \) approximate multipliers. We then utilize various machine learning models to evaluate and select configurations satisfying application accuracy and performance constraints. We have evaluated our proposed methodology for three benchmark applications, i.e., biomedical signal processing, image processing, and ANNs. We report more non-dominated approximate multipliers with better hypervolume contribution than state-of-the-art designs for these benchmark applications with the proposed design methodology. Salim Ullah, Siva Satyendra Sahoo, Nemath Ahmed, Debabrata Chaudhury, Akash Kumar 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2021 | CLAppED: A Design Framework for Implementing Cross-Layer Approximation in FPGA-based Embedded SystemsabstractWith the rising variation and complexity of embedded work-loads, FPGA-based systems are being increasingly used for many applications. The reconfigurability and high parallelism offered by FPGAs are used to enhance the overall performance of these applications. However, the resource constraints of embedded platforms can limit the performance in multiple ways. In recent years, Approximate Computing has emerged as a viable tool for improving the performance by utilizing reduced precision data structures and resource-optimized high-performance arithmetic operators. However, most of the related state-of-the-art research has mainly focused on utilizing approximate computing principles individually on different layers of the computing stack. Nonetheless, approximations across different layers of computing stack can substantially enhance the system’s performance. To this end, we present a framework to enable the intelligent exploration and highly accurate identification of the feasible design points in the large design space enabled by cross-layer approximations. Our framework proposes a novel polynomial regression-based method to model approximate arithmetic operators. The proposed method enables machine learning models to better correlate approximate operators with their impact on an application’s output quality. We use a 2D convolution operator as a test case and present the results for FPGA- based approximate hardware accelerators. Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001 |
DAC | 1 |
| 2021 | MemOReL: A Memory-oriented Optimization Approach to Reinforcement Learning on FPGA-based Embedded SystemsabstractReinforcement Learning (RL) represents the machine learning method that has come closest to showing human-like learning. While Deep RL is becoming increasingly popular for complex applications such as AI-based gaming, it has a high implementation cost in terms of both power and latency. Q-Learning, on the other hand, is a much simpler method that makes it more feasible for implementation on resource-constrained embedded systems for control and navigation. However, the optimal policy search in Q-Learning is a compute-intensive and inherently sequential process and a software-only implementation may not be able to satisfy the latency and throughput constraints of such applications. To this end, we propose a novel accelerator design with multiple design trade-offs for implementing Q-Learning on FPGA-based SoCs. Specifically, we analyze the various stages of the Epsilon-Greedy algorithm for RL and propose a novel microarchitecture that reduces the latency by optimizing the memory access during each iteration. Consequently, we present multiple designs that provide varying trade-offs between performance, power dissipation, and resource utilization of the accelerator. With the proposed approach, we report considerable improvement in throughput with lower resource utilization over state-of-the-art design implementations. Siva Satyendra Sahoo, Akhil Raj Baranwal, Salim Ullah, Akash Kumar 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2021 | Area-Optimized Accurate and Approximate Softcore Signed Multiplier ArchitecturesabstractMultiplication is one of the most extensively used arithmetic operations in a wide range of applications. In order to provide resource-efficient and high-performance multipliers, previous works have proposed different designs of accurate and approximate multipliers-mainly for ASIC-based systems. However, the architectural differences between ASICs- and FPGA-based systems limit the effectiveness of these multipliers for FPGA-based systems. Moreover, most of these multiplier designs are valid only for unsigned numbers. To bridge this gap, we propose a novel implementation technique for designing resource-efficient and low-power accurate and approximate signed multipliers which are optimized for FPGA-based systems. Compared to Vivado's area-optimized multiplier IPs, the designs obtained using our proposed technique occupy 47 to 63 percent less area (Lookup Tables). To accelerate further research in this direction and reproduce the presented results, the RTL and behavioral models of our proposed methodology are available as an open-source library.11.Online. [Available]: https://cfaed.tu-dresden.de/pd-downloads. Salim Ullah, Hendrik Schmidl, Siva Satyendra Sahoo, Semeen Rehman, Akash Kumar 0001 |
IEEE Trans. Computers | 1 |
| 2021 | ReLAccS: A Multilevel Approach to Accelerator Design for Reinforcement Learning on FPGA-Based SystemsabstractReinforcement learning (RL), specifically Q-learning, with human-like learning abilities to learn from experience without any a priori data, is being increasingly used in embedded systems in the field of control and navigation. However, finding the optimal policy in this approach can be highly compute-intensive, and a software-only implementation may not satisfy the application's timing constraints. To this end, we propose optimization methods at multiple levels of accelerator design for RL. Specifically, at the architecture-level, we exploit the instruction-level parallelism and the spatial parallelism in FPGAs to improve the throughput over state-of-the-art designs by up to 34%. Further, we propose lookup table-level optimizations to reduce the resource utilization and power dissipation of the accelerator. Finally, we propose algorithm-level approximation that can be used for acceleration of Q-learning problems with more states and for reducing the peak power dissipation. We report up to 10× reduction in power dissipation with marginal degradation in quality of results. Akhil Raj Baranwal, Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | LeAp: Leading-one Detection-based Softcore Approximate Multipliers with Tunable AccuracyabstractApproximate multipliers are ubiquitously used in diverse applications by exploiting circuit simplification, mainly specialized for Application-Specific Integrated Circuit (ASIC) platforms. However, the intrinsic architectural specifications of Field-Programmable Gate Arrays (FPGAs) prohibited comparable resource gains when directly applying these techniques. LeAp is an area-, throughput-, and energy-efficient approximate multiplier for FPGAs which efficiently utilizes 6-input Look-up Tables (6-LUTs) and fast carry chains in its novel approximate log calculator to implement Mitchell's algorithm. Moreover, three novel error-refinement schemes with negligible area overhead and independent from multiplier-size, have boosted accuracy to>99%. Experimental results obtained from Vivado, Artificial Neural Network (ANN) and image processing applications indicate superiority of proposed multiplier over accurate and state-of-the-art approximate counterparts. In particular, LeAp outperforms the 32x32 accurate multiplier by achieving 69.7%, 14.7%, 42.1%, and 37.1% improvement in area, throughput, power, and energy, respectively. The library of RTL and behavioral implementations will be open-sourced at https://cfaed.tu-dresden.de/pd-downloads. Zahra Ebrahimi, Salim Ullah, Akash Kumar 0001 |
ASP-DAC | 2 |
| 2020 | L2L: A Highly Accurate Log_2_Lead Quantization of Pre-trained Neural NetworksabstractDeep Neural Networks are one of the machine learning techniques which are increasingly used in a variety of applications. However, the significantly high memory and computation demands of deep neural networks often limit their deployment on embedded systems. Many recent works have considered this problem by proposing different types of data quantization schemes. However, most of these techniques either require post-quantization retraining of deep neural networks or bear a significant loss in output accuracy. In this paper, we propose a novel quantization technique for parameters of pre-trained deep neural networks. Our technique significantly maintains the accuracy of the parameters and does not require retraining of the networks. Compared to the single-precision floating-point numbers-based implementation, our proposed 8-bit quantization technique generates only ~1% and the ~0.4%, loss in top-1 and top-5 accuracies respectively for VGG16 network using ImageNet dataset. Salim Ullah, Siddharth Gupta 0004, Kapil Ahuja, Aruna Tiwari, Akash Kumar 0001 |
DATE | 1 |
| 2020 | SIMDive: Approximate SIMD Soft Multiplier-Divider for FPGAs with Tunable AccuracyabstractThe ever-increasing quest for data-level parallelism and variable precision in ubiquitous multimedia and Deep Neural Network (DNN) applications has motivated the use of Single Instruction, Multiple Data (SIMD) architectures. To alleviate energy as their main resource constraint, approximate computing has re-emerged, albeit mainly specialized for their Application-Specific Integrated Circuit (ASIC) implementations. This paper, presents for the first time, an SIMD architecture based on novel multiplier and divider with tunable accuracy, targeted for Field-Programmable Gate Arrays (FPGAs). The proposed hybrid architecture implements Mitchell's algorithms and supports precision variability from 8 to 32 bits. Experimental results obtained from Vivado, multimedia and DNN applications indicate superiority of proposed architecture (both in SISD and SIMD) over accurate and state-of-the-art approximate counterparts. In particular, the proposed SISD divider outperforms the accurate Intellectual Property (IP) divider provided by Xilinx with 4x higher speed and 4.6x less energy and tolerating only 0.8% error. Moreover, the proposed SIMD multiplier-divider supersede accurate SIMD multiplier by achieving up to 26%, 45%, 36%, and 56% improvement in area, throughput, power, and energy, respectively. Zahra Ebrahimi, Salim Ullah, Akash Kumar 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | A Comparative Analysis of Neural Networks and Enhancement of ELM for Short Term Load Forecasting
Rahim Ullah, Nadeem Javaid, Ghulam Hafeez, Salim Ullah, Fahad Ahmad, Ashraf Ullah |
CISIS | 4 |
| 2019 | Design Methodology for Embedded Approximate Artificial Neural NetworksabstractArtificial neural networks (ANNs) have demonstrated significant promise while implementing recognition and classification applications. The implementation of pre-trained ANNs on embedded systems requires representation of data and design parameters in low-precision fixed-point formats; which often requires retraining of the network. For such implementations, the multiply-accumulate operation is the main reason for resultant high resource and energy requirements. To address these challenges, we present Rox-ANN, a design methodology for implementing ANNs using processing elements (PEs) designed with low-precision fixed-point numbers and high performance and reduced-area approximate multipliers on FPGAs. The trained design parameters of the ANN are analyzed and clustered to optimize the total number of approximate multipliers required in the design. With our methodology, we achieve insignificant loss in application accuracy. We evaluated the design using a LeNet based implementation of the MNIST digit recognition application. The results show a 65.6%, 55.1% and 18.9% reduction in area, energy consumption and latency for a PE using 8-bit precision weights and activations and approximate arithmetic units, when compared to 16-bit full precision, accurate arithmetic PEs. Adarsha Balaji, Salim Ullah, Anup Das 0001, Akash Kumar 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | SMApproxlib: library of FPGA-based approximate multipliersabstractThe main focus of the existing approximate arithmetic circuits has been on ASIC-based designs. However, due to the architectural differences between ASICs and FPGAs, comparable performance gains cannot be achieved for FPGA-based systems by using the approximations defined, particularly for ASIC-based systems. This paper exploits the structure of the 6-input lookup tables and associated carry chains of modern FPGAs to define a methodology for designing approximate multipliers optimized for FPGA-based systems. Using our presented methodology, we present SMApproxLib, an open source library of approximate multipliers with different bit-widths, output accuracies and performance gains. Being the first open source library of FPGA-based approximate multipliers, SMAp-proxLib can serve as a benchmark for designing and comparing future FPGA-based approximate arithmetic circuits. Salim Ullah, Sanjeev Sripadraj Murthy, Akash Kumar 0001 |
DAC | 1 |
| 2018 | Area-optimized low-latency approximate multipliers for FPGA-based hardware acceleratorsabstractThe architectural differences between ASICs and FPGAs limit the effective performance gains achievable by the application of ASIC-based approximation principles for FPGA-based reconfigurable computing systems. This paper presents a novel approximate multiplier architecture customized towards the FPGA-based fabrics, an efficient design methodology, and an open-source library. Our designs provide higher area, latency and energy gains along with better output accuracy than those offered by the state-of-the-art ASIC-based approximate multipliers. Moreover, compared to the multiplier IP offered by the Xilinx Vivado, our proposed design achieves up to 30%, 53%, and 67% gains in terms of area, latency, and energy, respectively, while incurring an insignificant accuracy loss (on average, below 1% average relative error). Our library of approximate multipliers is open-source and available online at https://cfaed.tudresden.de/pd-downloads to fuel further research and development in this area, and thereby enabling a new research direction for the FPGA community. Salim Ullah, Semeen Rehman, Bharath Srinivas Prabakaran, Florian Kriebel, Muhammad Abdullah Hanif, Muhammad Shafique 0001, Akash Kumar 0001 |
DAC | 1 |
| 2018 | DeMAS: An efficient design methodology for building approximate adders for FPGA-based systemsabstractThe current state-of-the-art approximate adders are mostly ASIC-based, i.e., they focus solely on gate and/or transistor level approximations (e.g., through circuit simplification or truncation) to achieve area, latency, power and/or energy savings at the cost of accuracy loss. However, when these designs are synthesized for FPGA-based systems, they do not offer similar reductions in area, latency and power/energy due to the underlying architectural differences between ASICs and FPGAs. In this paper, we present a novel generic design methodology to synthesize and implement approximate adders for any FPGA-based system by considering the underlying resources and architectural differences. Using our methodology, we have designed, analyzed and presented eight different multi-bit adder architectures. Compared to the 16-bit accurate adder, our designs are successful in achieving area, latency and power-delay product gains of 50%, 38%, and 53%, respectively. We also compare our approximate adders to state-of-the-art approximate adders specialized for ASIC and FPGA fabrics and demonstrate the benefits of our approach. We will make the RTL and behavioral models of our and state-of-the-art designs open-source at https://sourceforge.net/projects/approxfpgas/ to further fuel the research and development in the FPGA community and to ensure reproducible research. Bharath Srinivas Prabakaran, Semeen Rehman, Muhammad Abdullah Hanif, Salim Ullah, Ghazal Mazaheri, Akash Kumar 0001, Muhammad Shafique 0001 |
DATE | 4 |