Akash Kumar 0001

dblp:29/414 · DBLP profile ↗
← Back
211ranked-venue papers
8as first author
89since 2021 · last 2026
0000-0001-7125-1737ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 191 · 8 first-author · 84 since 2021Software engineering, systems software and programming languages · 49 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Security and privacy · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 RETRO: Mitigating Power Side-Channel Attacks with Reconfigurable RFET-based Ring Oscillators
abstract
Power side-channel attacks are among the most effective physical attacks, threatening the security of circuits such as cryptographic circuits by exploiting information leakage from their physical implementation. Among various masking and hiding countermeasures that have been proposed, Ring Oscillator (RO)-based solutions are considered low-overhead circuitry addons that can be integrated into different circuits to hide the data dependency of power consumption by adding noise to their power signatures. The Three-Independent-Gate Reconfigurable Field-Effect Transistor (TIG-RFET) is an emerging technology that offers runtime reconfigurability between N-type and P-type operation, supports both low-VTand high-VTmodes, and provides an internal wired-AND function, making it a strong candidate for efficient implementation of various hardware security methods. In this paper, we propose a novel reconfigurable RFET-based RO that provides controllable frequency through RFET-based inverters with reconfigurable delay. Using these ROs, we introduce a countermeasure called RETRO, which can generate noise by varying both the amplitude and frequency of power consumption. To evaluate the efficacy of RETRO, we applied it to the Piccolo S-box, a lightweight cryptographic circuit, and simulation results demonstrate that it effectively enhances resilience against Correlation Power Analysis (CPA). Furthermore, we show that reconfigurable frequency broadens the noise spectrum, making filtering considerably more difficult.
Nima Kavand, Tushar Niranjan, Armin Darjani, Akash Kumar 0001
DATE4
2026 Focus Session: Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification
abstract
The design of embedded safety-critical systems such as those used in next-generation automotive and autonomous platforms, is increasingly challenged by escalating system complexity, hardware–software heterogeneity, and the integration of intelligent, data-driven components. Ensuring dependability in such systems requires a holistic approach that spans multiple abstraction layers and encompasses both design- and run-time assurance. Traditional methods for reliability, safety, and security management often fall short in addressing the dynamic and uncertain behaviors introduced by Artificial Intelligence (AI) and Machine Learning (ML) components, especially under stringent real-time, power, and safety constraints. While AI and ML offer powerful predictive, adaptive, and self-optimizing capabilities that can enhance system dependability, their inherent non-determinism, data-dependence, and lack of formal guarantees introduce new challenges for verification, validation, and certification. This paper explores emerging methodologies, architectures, and frameworks for designing dependable autonomous and embedded systems in the era of AI. It highlight advances in reliability modeling, secure system design, and certification approaches that account for imperfect, learning-enabled components, aiming to bridge the gap between AI innovation and certifiable system-level dependability.
Behnaz Ranjbar, Kirankumar Raveendiran, Sudeep Pasricha, Samarjit Chakraborty, Cecilia Carbonelli, Akash Kumar 0001
DATE6
2026 EAGEL: Explainable And Generalized Structural Exploitation Against Logic Locking
abstract
Logic locking protects hardware intellectual property (IP) against piracy, but recent machine learning (ML)-based attacks have shown highly accurate key prediction against this technique. However, existing attacks have three main shortcomings: (1) they lack a statistical analysis regarding the size of leakage locality and they lack a formal leakage model, leading to a back-and-forth cycle of attacks and defenses; (2) they lack circuit-level analysis for choosing suitable ML models; and (3) they limit attackers to the same toolset available to security designers leaving a gap in understanding tool-agnostic vulnerabilities. In this paper, we analyze gate-based obfuscation and provide statistical evidence that leakage is concentrated within a 2–3 hop key locality. We propose a first-of-its-kind information-theoretic formalization of structural leakage, showing that locking remains vulnerable when it introduces independent, localized effects after re-synthesis, regardless of the locking approach. Building on this, we design compact structural and simulation-based functional features that capture local leakage around key gates. Using these features, we introduce EAGEL, an explainable and generalized security assessment framework that removes the attacker’s dependence on the designer’s synthesis toolchain and empirically validates our leakage formalization. EAGEL achieves up to \(98\%\) key prediction accuracy and \(99.5\%\) prediction certainty, outperforming complex state-of-the-art GNN-based attacks by up to \(11\%\) and on average by \(8.58\%\), with a 3.4 × speedup.
Armin Darjani, Palaniappan Ramasamy, Nima Kavand, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI4
2026 RESIST : Structured Regularization from Weight Similarity to Weight Diversity for Improving Error Resilience of Neural Networks
Maryam Eslami, Salim Ullah, Akash Kumar 0001
IOLTS3
2026 Structurally Secure Obfuscation: Assessing and Mitigating Structural Vulnerabilities in Circuits Obfuscation
abstract
Because of the globalization of IC manufacturing and to protect IP integrity and confidentiality, circuit obfuscation techniques have been developed. These methods secure the circuit through obfuscation approaches. Recently, advanced machine learning (ML)-based structural attacks have been introduced that employ the structure of the circuit to reverse the obfuscation mechanism. These attacks use ML-based approaches to analyze and neutralize obfuscation schemes without requiring unlocked functional circuits, posing a significant challenge to IP security. To counter ML-based attacks, in this article, we first analyze the sources of structural leakages of the interconnect obfuscation technique, one of the most robust IP protection mechanisms. We conduct a first-of-its-kind analysis of the circuit’s netlist graph, obfuscated using interconnect obfuscation, to evaluate its robustness against link prediction techniques. Based on our analysis, we introduce a security assessment tool that evaluates the strength of the obfuscation technique in omitting structural leakages that lead to the success of ML-based attacks. Our assessment tool reveals that previous obfuscation methods fall short of achieving their intended security levels. This leads to our second contribution, which is proposing ML-SafeConnect, an interconnect obfuscation technique that protects the obfuscated substructures by completely eliminating distance-based structural leakages. Using our assessment tool and state-of-the-art ML-based attack, we demonstrate that our obfuscation mechanism surpasses previous interconnect obfuscation techniques in preventing structural leakages. We show that ML-SafeConnect completely thwarts ML-based attacks for all benchmark circuits by decreasing the accuracy of state-of-the-art attacks to below 50%.
Armin Darjani, Nima Kavand, Zhentao Han, Akash Kumar 0001
ACM Trans. Design Autom. Electr. Syst.4
2025 Special Sessions - Emerging Scope and Design Challenges for Approximate Computing: Optimizing Accuracy-PPA trade-offs and Beyond
abstract
The rapid growth of AI workloads is driving interest in Approximate Computing (AxC) as a means to enable low-cost, energy-efficient inference in resource-constrained systems. By introducing controlled inaccuracies, AxC can deliver substantial gains in power, performance, and area (PPA) while leveraging the inherent error tolerance of many AI models. Achieving this potential requires adapting existing frameworks to support the design and optimization of neural networks with approximate operators. Modern AxC research extends beyond accuracy-PPA trade-offs to address reliability and security, reducing redundancy overheads and exploring the distinctive side-channel implications of approximation. Application-aware approaches, such as those for spiking neural networks, show that tailoring approximation to workload-specific error behavior can surpass generic strategies. This article examines AI-guided design methods and the interplay between efficiency, reliability, and security, highlighting how these interconnected facets can advance embedded and high-performance computing.
Siva Satyendra Sahoo, Bastien Deveautour, Marcello Traiola, Chongyan Gu, Yun Wu 0003, Aditya Japa, Salim Ullah, Akash Kumar 0001
CASES8
2025 Design, Model, and Explore Approximate Arithmetic Operators with AI/ML: A Tutorial
abstract
Approximate Computing (AxC) is being actively explored to meet the energy and performance requirements of resource-constrained embedded systems. Approximate arithmetic operators (AxOs), for instance, let edge-AI systems trade tiny, bounded errors for big wins in power, performance, and area. This tutorial demystifies AxO design, modeling, and exploration: from platform-aware operator synthesis (e.g., selective LUT pruning) to application-specific DSE that uses AI/ML to navigate massive trade-off spaces. We contrast selection (library) vs. synthesis (generate-and-optimize) flows, show when FPGA-aware adders/multipliers outperform ASIC-ported designs, and connect operator-level error to task-level metrics (e.g., Conv2D, MLP). The tutorial includes hands-on Jupyter notebooks, ready-to-reuse operator models, and a practical recipe for building Pareto-optimal AxOs under accuracy constraints - plus a peek at AxOSyn, an open-source framework that unifies selection/synthesis, surrogate fitness, and search using evolutionary algorithms.
Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001
CASES3
2025 BiKA: Binarized KAN-inspired Neural Network for Efficient Hardware Accelerator Designs
abstract
The continuously growing size of Neural Network (NN) models makes the design of lightweight neural network accelerators for edge devices an emerging subject in recent research. Previous works explored different lightweight technologies or even emerging neural network structures, such as quantization, approximate computing, neuromorphic computing, etc., to reduce hardware resource consumption in accelerator designs. This inspired our interest in exploring the potential of other emerging network structures in hardware accelerator designs. Kolmogorov-Arnold Network (KAN) [1] is a recently proposed novel neural network structure by replacing the multiplication and activation function in Artificial Neural Networks (ANN) with learnable nonlinear functions, which has the potential to transform the paradigm of neural network design. However, considering the complexity of the nonlinear function on hardware, the design of the lightweight hardware accelerator of KAN lacks thoroughly related research.
Salim Ullah, Akash Kumar 0001
FCCM3
2025 Flip-Break: Breaking Flip-flop-based Logic Locking in Sequential Circuits
Armin Darjani, Nima Kavand, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI3
2025 Invited Paper: Circuit and Architecture Design with Emerging Computing Paradigms
abstract
As emerging computing paradigms push beyond the limitations of traditional CMOS-based computing using Von Neumann architectures, there is a growing need to rethink and extend Electronic Design Automation (EDA) methodologies to support their unique characteristics. These paradigms—including Approximate Computing, In-Memory Computing, Reconfigurable Field-Effect Transistors (RFETs), and Photonic Computing—represent diverse and promising directions beyond conventional digital design. Collectively, they offer transformative potential for achieving significant improvements in energy efficiency, computational speed, and architectural scalability. For example, application-specific approximate computing enables the design of custom arithmetic circuits that exploit application-level error resilience, allowing for optimized accuracy–power–performance–area (PPA) trade-offs in error-tolerant applications. Similarly, processing-in-non-volatile memories, such as those based on Ferroelectric Field-effect Transistors (FeFETs), enhances energy efficiency by enabling analog computation—particularly for operations like matrix multiplication—directly within the memory arrays. The intrinsic polymorphism of RFETs supports compact, multifunctional logic gates and introduces new opportunities for circuit-level obfuscation and security-aware design. Likewise, photonic analog wavefront computing offers substantial gains in latency and energy efficiency by encoding and processing information in the analog optical domain, leveraging phenomena such as diffraction and interference to perform computation at the speed of light. However, they also introduce a host of new challenges in circuit and architecture design, such as vast and irregular design spaces, analog and non-Boolean behavior, and new device-level constraints that existing EDA tools are not capable of handling. To this end, the current article focuses on the development of efficient and robust EDA frameworks that can enable the practical realization of circuits and architectures in these emerging domains.
Salim Ullah, Siva Satyendra Sahoo, Can Li 0024, Chao Li 0065, Liu Liu 0023, Tomas Sousa Pereira, Xunzhao Yin, Armin Darjani, Nima Kavand, Chakravarthy Bodla, Rupa Yashaswi Panduga, Aniruddh Holemadlu, Johannes Maly, Jonathan Förste, Samarth Vadia, Xiaobo Sharon Hu, Akash Kumar 0001
ICCAD18
2025 Mixa-Q: Revisiting Activation Sparsity for Vision Transformers From a Mixed-Precision Quantization Perspective
abstract
In this paper, we propose MixA-Q, a mixed-precision activation quantization framework that leverages intra-layer activation sparsity (a concept widely explored in activation pruning methods) for efficient inference of quantized window-based vision transformers. For a given uniform-bit quantization configuration, MixA-Q separates the batched window computations within Swin blocks and assigns a lower bit width to the activations of less important windows, improving the trade-off between model performance and efficiency. We introduce a Two-Branch Swin Block that processes activations separately in high- and low-bit precision, enabling seamless integration of our method with most quantization-aware training (QAT) and post-training quantization (PTQ) methods, or with simple modifications. Our experimental evaluations over the COCO dataset demonstrate that MixA-Q achieves a training-free 1.35x computational speedup without accuracy loss in PTQ configuration. With QAT, MixA-Q achieves a lossless 1.25x speedup and a 1.53x speedup with only a 1% mAP drop by incorporating activation pruning. Notably, by reducing the quantization error in important regions, our sparsity-aware quantization adaptation improves the mAP of the quantized W4A4 model (with both weights and activations in 4-bit precision) by 0.7%, reducing quantization degradation by 24%.
Weitian Wang, Shubham Rai, Cecilia De la Parra, Akash Kumar 0001
ICCV4
2025 X-DINC: Toward Cross-Layer ApproXimation for theDistributed and In-Network ACceleration of Multi-Kernel Applications
abstract
With the rapid evolution of programmable network devices and the urge for energy-efficient and sustainable computing, network infrastructures are mutating toward a computing pipeline, providing In-Network Computing (INC) capability. Despite the initial success in offloading single/small kernels to the network devices, deploying multi-kernel applications remains challenging due to limited memory, computing resources, and lack of support for Floating Point (FP) and complex operations. To tackle these challenges, we present a cross-layer approximation and distribution methodology (X-DINC) that exploits the error resilience of applications. X-DINC utilizes a chain of techniques to facilitate kernel deployment and distribution across heterogeneous devices in INC environments. First, we identify approximation and optimization opportunities in data acquisition and computation phases of multi-kernel applications. Second, we simplify complex arithmetic operations to cope with the computation limitations of the programmable network switches. Third, we perform application-level sensitivity analysis to measure the trade-off between performance gain and Quality of Results (QoR) loss when approximating individual kernels via various techniques. Finally, a greedy heuristic swiftly generates Pareto/near-Pareto mixed-precision configurations that maximize the performance gain while maintaining the user-defined QoR. X-DINC is prototyped on a Virtex-7 Field Programmable Gate Array (FPGA) and evaluated using the Blind Source Separation (BSS) application on industrial audio dataset. Results show that X-DINC performs separation up to 35% faster with up to 88% lower Area-Delay Product (ADP) compared to an Accurate-Centralized approach, when distributed across 2 to 7 network nodes, while maintaining audio quality within an acceptable range of 15–20 dB.
Zahra Ebrahimi, Maryam Eslami, Xun Xiao, Akash Kumar 0001
Future Gener. Comput. Syst.4
2025 Fast Retraining of Approximate CNNs for High Accuracy
abstract
One technique for approximating neural networks (NNs) when deploying to resource-constrained systems is the use of approximate multiplications. Giving up full mathematical accuracy opens new opportunities for more efficient hardware implementations. Modeling the effects of inaccurate hardware already in the training stage improves performance but significantly slows down the training due to expensive type conversions and memory access operations. We propose a method to speed up the simulation of inaccurate hardware by using a composition of floating-point functions. Both an analytical and a data-driven method for finding these functions are provided. We further provide a study and implementation of per-channel quantization, a scheme that enhances the granularity of converting NN parameters to integers. This helps boost the application’s accuracy. In our evaluation, our floating-point models achieve up to a$4 \times $speed-up over the commonly used lookup table implementation, while providing a high-fidelity simulation of the target function. Extending quantization with per-channel granularity yields a median accuracy improvement of 0.87 p.p. for ResNet8/CIFAR10 with 4-bit weight quantization in combination with hardware using approximate multipliers (AMs). Our extended software toolkit for the study of AMs in PyTorch is publicly available and provides a variety of building blocks for applying inaccurate product functions to NNs.
Elias Trommer, Bernd Waschneck, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 Analytical Uncertainty Propagation in Neural Networks
abstract
The usage of machine-learning techniques, such as neural networks, is common in a large variety of domains. Estimating the certainty of a predicted value is important when precise information is gained. Nevertheless, the forward propagation of uncertainty in machine-learning models is hardly understood. In general, providing error bars for measurements (measurement uncertainty) is crucial when high precision is needed for decision-making. The objective of this work is the development of an analytical method for aleatoric uncertainty forward propagation in neural networks, based on analytical uncertainty propagation well known from physics and engineering. With that, the method gives provable correct results. A benefit is that the method does not require a different training procedure, but only needs the weights and biases of the neural network and is computationally inexpensive. The analytical method is applied to real-world examples from the semiconductor industry (regression and image classification). Its usefulness is demonstrated by the provided examples, which show how meaningful error bars are when machine learning may be used for decision-making.
Paul Jungmann, Julia Poray, Akash Kumar 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Enabling Energy-efficient AI Computing: Leveraging Application-specific Approximations : (Education Class)
abstract
The widespread adoption of Artificial intelligence and Machine Learning (AI/ML) models across various fields, such as healthcare, autonomous vehicles, smart agriculture, and industrial automation, has led to a growing demand for efficient and scalable AI/ML solutions. However, as AI/ML algorithms grow more complex, their substantial memory requirements and high energy consumption pose significant challenges for deployment on resource-constrained embedded systems, such as wearable health monitors and IoT devices. To this end, various techniques, such as model pruning, knowledge distillation, quantization of model parameters, and employing approximate arithmetic operators, are commonly explored to overcome these challenges [1] .
Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001
CASES3
2024 Dynamic Reconfigurable Security Cells Based on Emerging Devices Integrable in FDSOI Technology
abstract
While a number of measures have been proposed to protect the integrity of COS hardware, there are some inherent limitations from classical CMOS methods. Those already existing security methods, like logic locking can be improved with emerging technologies such as Reconfigurable Field Effect Transistors (RFETs). RFETs are a special type of doping-free, Schottky transistors which can work as a PFET or NFET as a function of biasing across its gates. In the present study we developed standard cell layouts for dynamic reconfigurable security cells based on three-independent-gated RFETs (TIG-RFETs). They layouts are compatible to an industrial 22nm FDSOI technology, feature the minimum pitch of the baseline technology, and obey all design rules necessary for co-integration. The designs enable a fair area comparison for RFET based digital application for the first time. Based on the sizing constraints from the layouts, a TCAD model of such a TIG-RFET is developed in Sentaurus TCAD to illustrate two biasing schemes for the application of TIGRFETs in this platform: reconfigurability with individual body-bias per transistor and reconfigurability at globally fixed body-bias. Due to the different operation options three variants of reconfigurable 2-XOR-XNOR and 2-NAND-NOR logic cells exhibiting different level of utility are designed. While the smallest dynamic 2-NAND-NOR gate needs roughly double the area of a CMOS 2-NAND gate from the reference library, the smallest 2-XOR-XNOR gate is only 20% larger than a CMOS 2-XOR. To quantify the area overhead for hardware security applications we calculated the number of logic locking gates that can be added per area overhead for a given circuit, here the ISCAS-85 C6288 benchmark circuit, as an example. Dynamic replacement based logic locking with TIG-RFETs shows to allow up to double the number of keys compared to classical CMOS logic locking per area overhead. Therefore, this work allows a realistic view on the application of RFETs in hardware security and its co-integrability along with some design constraints from an industrial PDK.
Niladri Bhattacharjee, Viktor Havel, Suruchi Kumari, Nima Kavand, Jorge Navarro Quijada, Akash Kumar 0001, Thomas Mikolajick, Jens Trommer
DATE6
2024 REDCAP: Reconfigurable RFET-Based Circuits Against Power Side-Channel Attacks
abstract
Power attacks are effective side-channel attacks (SCAs) that exploit weaknesses in the physical implementation of a cryptographic circuit to extract its secret information like encryption key. In recent years, emerging technologies have unlocked new possibilities in designing effective SCA countermeasures with less overhead. Reconfigurable Field-Effect Transistors (RFETs) are a type of beyond-CMOS technology that can be configured at run-time to act as an NFET or PFET transistor and provide two or more independent gates. These features make RFETs potent candidates for implementing hardware security techniques like logic locking and SCA countermeasures. In this paper, we propose REDCAP, a method to add randomness to the power traces of a circuit, employing compact reconfigurable RFET-based gates to make the design resilient against power SCAs. First, we explain the construction and control of reconfigurable blocks with isofunctional configurations inside the circuit. Then, we provide an algorithm to efficiently compose the reconfigurable blocks with other circuit parts to minimize the overhead and enable designers to determine the granularity of the reconfiguration. To evaluate our approach, we performed a Correlation Power Attack (CPA) on the S-box of the Piccolo and PRESENT, two lightweight cryptographic circuits, and the results show that REDCAP can highly enhance the resilience of the circuit against power SCAs.
Nima Kavand, Armin Darjani, Giulio Galderisi, Jens Trommer, Thomas Mikolajick, Akash Kumar 0001
DATE6
2024 Motivating the Use of Machine-Learning for Improving Timing Behaviour of Embedded Mixed-Criticality Systems
abstract
In Mixed-Criticality (MC) systems, due to encoun-tering multiple Worst-Case Execution Times (WCETs) for each task corresponding to the system operation modes, estimating appropriate WCETs for tasks in lower-criticality (LO) modes is essential to improve the system's timing behavior. While numerous studies focus on determining WCET in the high-criticality mode, determining the appropriate WCET in the LO mode poses significant challenges and has been addressed in a few research works due to its inherent complexity. This article introduces a novel scheme to obtain appropriate WCET for LO modes. We propose an ML-based approach for WCET estimation based on the application's source code analysis and the model training using a comprehensive data set. The experimental results show a significant improvement in utilization by up to 23.3 % for the ML-based approach, while mode switching probability is bounded by 7.19 % in the worst-case scenario.
Behnaz Ranjbar, Akash Kumar 0001
DATE3
2024 LeQC-At: Learning Quantization Configurations During Adversarial Training for Robust Deep Neural Networks
abstract
Due to the high feature learning capability of Deep Neural Networks (DNNs), they are widely used in state-of-the-art machine learning tasks such as image and text recognition and natural language processing. However, in the recent developments of deep learning models, it has been observed that specially crafted inputs (adversarial samples) can deceive DNNs and result in incorrect predictions with high confidence. Such incorrect predictions by adversarially attacked DNNs can have devastating results in safety-critical applications. This situation can be further exacerbated in quantized DNNs, which employ reduced precision numbers to reduce the overall computational complexity of DNNs. To this end, this work proposes a framework for the joint optimization of DNNs to improve the natural accuracy of quantized DNNs and increase the robustness of the quantized models against adversarial attacks. In particular, the proposed framework employs quantization step size aware adversarial training of DNNs. Our proposed framework is generic and can be utilized with any quantization scheme that allows learning of quantization configurations during training. Furthermore, we present a novel loss function for adversarial training to improve the quantized networks' accuracy. For example, our 3-bit quantized adversarial training of ResNet-18, ResNet-34, and WideResNet shows up to 21.63%, 24.49%, and 15.08% higher inference accuracy with attacked data, respectively, when compared to 3-bit vanilla quantized adversarial training on benchmark datasets.
Siddharth Gupta 0004, Salim Ullah, Akash Kumar 0001
DSD3
2024 BitSys: Bitwise Systolic Array Architecture for Multi-precision Quantized Hardware Accelerators
abstract
Quantized Neural Networks (QNN) have been widely applied in hardware accelerator designs for edge. Because lower precision in quantization leads to higher accuracy loss, the mixed-precision scheme has been explored by using different precision in different layers to trade off resource consumption and inference accuracy. Because regular multiplier designs do not support the reconfiguration for multi-precision, we explored a runtime reconfigurable multi-precision bitwise systolic array design, BitSys, for mixed-precision multiplication in QNN accelerators. The popular design in previous works, such as [1], divides the inputs of multipliers as two parts for four sub-multipliers and preset left shifting, achieving the reconfigurable multiplication by disabling two of the submultipliers. Our design is inspired by the Bitshifter architecture from the works of Liu et al. [2], [1]. We convert the$n\times n$- bit multiplication as$A \times B=\sum_{i=0}^{n-1} \sum_{j=0}^{n-1} 2^{i+j} a_i b_j$. As shown in Figure 1, partial products$P_{i+j}$is the sum of subpartial products$a_{i}b_{j}$with left shifting value,$i+j$. The subpartial product masks select the desired$a_{i}b_{j}$to configure the multi-channel according to the precision. One mask square represents one channel. Therefore, we can implement a bitwise systolic array as Figure 2 (left). The sub-partial product computation and mask are fused in one LUT primitive as one processing element in Figure 2 (right up). The sums of elements, considered the sign-bits, in the diagonal with the same left shifting value shown in Figure 2 (left) are the inputs,$D_{k}$, of the output-generate pipeline of Figure 2 (right down). Systolic array sequentially generates the$D_{k}$, and the final multi-precision output is the sum of all$D_{k}$. We implemented our BitSys multiplier as a systolic array for mixed-precision QNN acceleration. The comparison with the works of Liu et al. [1] is shown in Table I. Our systolic array accelerator consumes more hardware resources than the three single-layer accelerator instances of Liu et al. [1]. However, because of our bitwise processing design, BitSys instance can support 250MHz clock frequency because of the low critical path delay of our multiplier and achieves 188.5-274.7% speed-up in the evaluation of one four-layer 1/2/4/8-bit mixed-precision quantized MLP. Furthermore, our design does not change the input/out width when configuring to different precision, which can be easily integrated into existing accelerator designs.
Salim Ullah, Akash Kumar 0001
FCCM3
2024 Flip-Lock: A Flip-Flop-Based Logic Locking Technique for Thwarting ML-based and Algorithmic Structural Attacks
abstract
Machine learning (ML) and algorithmic structural attacks have highlighted the possibility of utilizing structural leakages of an obfuscated circuit to reverse engineer the locking mechanism. These structural leakages are rooted in security-agnostic synthesis tools that lead to discernible patterns within the vicinity of the locking substructures. This paper has two contributions. Firstly, we present the innovative Flip-lock, a novel approach that utilizes flip-flops along logic gates to prevent synthesis tools’ structural leakages. As this is the first work that incorporates flip-flops in locking structures, our second contribution is the development of a comprehensive analysis tool designed to identify and assess potential structural leakages in the vicinity of flip-flops within circuits. We name our tool Flip-attack. We employ the Flip-attack analysis to strengthen Flip-lock’s resilience against potential future attacks. Our findings demonstrate that Flip-lock possesses the capability to neutralize all existing state-of-the-art structural attacks effectively. Furthermore, by enhancing Flip-lock through the incorporation of our analysis tool, we establish a robust defense mechanism that can safeguard the circuit’s security against potential future threats.
Armin Darjani, Nima Kavand, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI3
2024 GREEN: An Approximate SIMD/MIMD CGRA for Energy-Efficient Processing at the Edge
abstract
The rapid evolution of compute-intensive programs from bio-signal to image-, and video-processing has motivated moving toward Coarse Grained Reconfigurable Architectures (CGRAs), having high parallelism capability with post-fabrication datapath versatility. To enhance energy-efficiency of such error-resilient applications, State-of-the-Art (SoA) CGRAs exploit approximation techniques, while maintaining an acceptable accuracy for the final Quality of Result (QoR). However, such CGRAs suffer from overheads of utilizing separate Add/Mul/Div units. We propose GREEN as an energy-efficient CGRA, which enables synergistic effects of a chain of approximation and optimization techniques in various levels of abstraction, from application-, to architecture-, to circuit-level, in a cross-layer hierarchy. Enabling this, GREEN offers different levels of energy-accuracy trade-offs through the flexibility of its small Processing Elements (PEs), each of which can support various functionalities and precision-adaptability in a Single Instruction, Multiple Data (SIMD) or Multiple Instruction, Multiple Data (MIMD) manner. Experimental results obtained with Synopsys Design Compiler and Cadence Innovus at 45 nm CMOS technology node demonstrate the efficiency of the proposed SISD/SIMD/MIMD CGRA over the accurate and SoA counterparts. In particular, the MIMD mode of GREEN enables up to 6.6× higher throughput while dissipating 21% less energy than the accurate counterpart. Moreover, the end-to-end evaluation of GREEN variants on eight single-and multi-kernel applications from classification, bio-signal (ECG/EEG), and image/video processing domains demonstrates significant performance improvement, compared to the accurate CGRA. In particular, GREEN-MIMD not only speed-ups the ECG QRS detection by 49% and consumes 43% less area and 66% less energy than the accurate CGRA, but also maintains the heartbeat detection accuracy at 100%. GREEN implementations is available at https://cfaed.tu-dresden.de/pd-downloads.
Zahra Ebrahimi, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Smaller Together: Groupwise Encoding of Sparse Neural Networks
abstract
With the drive towards ever more intelligent devices, neural networks are deployed on smaller and smaller systems. For these embedded microcontrollers, memory consumption becomes a significant challenge. We propose multiple encoding schemes that convert the decrease in parameter counts, achieved through unstructured pruning, into tangible memory savings. We first discuss a sparse encoding scheme for arbitrary sparse matrices that is based on encoding offsets from a predicted even spacing of elements in a row. The compression rate of this scheme is improved further by identifying groups of elements which can be encoded with even lower overhead. Both methods are combined into a hybrid scheme which encodes arbitrary sparse matrices with low overhead, while allowing for parallel access to multiple elements in a row at once—an important feature for using the scheme on the latest generation of microcontrollers with parallel SIMD capabilities. Our scheme compresses sparse models to below the size of their dense counterparts for sparsities as low as 30% and reduces model size by 32.4% and 26.4% at less than one percentage point of accuracy loss for two convolutional neural network tasks in our evaluation.
Elias Trommer, Bernd Waschneck, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 AxOSpike: Spiking Neural Networks-Driven Approximate Operator Design
abstract
Approximate computing (AxC) is being widely researched as a viable approach to deploying compute-intensive artificial intelligence (AI) applications on resource-constrained embedded systems. In general, AxC aims to provide disproportionate gains in system-level power-performance-area (PPA) by leveraging the implicit error tolerance of an application. One of the more widely used methods in AxC involves circuit pruning of arithmetic operators used to process AI workloads. However, most related works adopt an application-agnostic approach to operator modeling for the design space exploration (DSE) of Approximate Operators (AxOs). To this end, we propose an application-driven approach to designing AxOs. Specifically, we use spiking neural network (SNN)-based inference to present an application-driven operator model resulting in AxOs with better-PPA-accuracy tradeoffs compared to traditional circuit pruning. Additionally, we present a novel FPGA-specific operator model to improve the quality of AxOs that can be obtained using circuit pruning. With the proposed methods, we report designs with up to 26.5% lower PDPxLUTs with similar application-level accuracy. Further, we report a considerably better set of design points than related works with up to 51% better-Pareto front hypervolume.
Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 AxOCS: Scaling FPGA-Based Approximate Operators Using Configuration Supersampling
abstract
The rising usage of AI/ML-based processing across application domains has exacerbated the need for low-cost ML implementation, specifically for resource-constrained embedded systems. To this end, approximate computing, an approach that explores the power, performance, area (PPA), and behavioral accuracy (BEHAV) trade-offs, has emerged as a possible solution for implementing embedded machine learning. Due to the predominance of MAC operations in ML, designing platform-specific approximate arithmetic operators forms one of the major research problems in approximate computing. Recently, there has been a rising usage of AI/ML-based design space exploration techniques for implementing approximate operators. However, most of these approaches are limited to using ML-based surrogate functions for predicting the PPA and BEHAV impact of a set of related design decisions. While this approach leverages the regression capabilities of ML methods, it does not exploit the more advanced approaches in ML. To this end, we propose, a methodology for designing approximate arithmetic operators through ML-based supersampling. Specifically, we present a method to leverage the correlation of PPA and BEHAV metrics across operators of varying bit-widths for generating larger bit-width operators. The proposed approach involves traversing the relatively smaller design space of smaller bit-width operators and employing its associatedDesign-PPA-BEHAVrelationship to generate initial solutions for metaheuristics-based optimization for larger operators. The experimental evaluation of for FPGA-optimized approximate operators shows that the proposed approach significantly improves the quality—resulting hypervolume for multi-objective optimization—of$8\times8$signed approximate multipliers.
Siva Satyendra Sahoo, Salim Ullah, Soumyo Bhattacharjee, Akash Kumar 0001
IEEE Trans. Circuits Syst. I Regul. Pap.4
2024 Thwarting GNN-Based Attacks Against Logic Locking
abstract
The globalization of the IC manufacturing flow has exposed intellectual property (IP) to many untrustworthy entities. As a result, security should be considered a new paradigm in designing circuits to protect the integrity and confidentiality of the IP. Logic locking is a holistic design-for-trust (DFT) technique that can protect circuits against IP piracy and reverse engineering. However, a large body of recent research has demonstrated successful methods of recovering the secret key and restoring the original functionality of existing locking systems. Although SAT attack has been a de facto technique to break the logic locking, the threat model and efficiency of this attack have been questioned recently. To overcome these shortcomings, researchers have proposed powerful structural attacks that break the locked circuits without the need for functionally unlocked circuits (Oracle). Among structural attacks, machine learning (ML)-based attacks are the most potent attacks as they harness the power of neural networks to learn traces of the locking structures and use this knowledge to reverse back and neutralize the locking scheme. Among ML approaches, GNN (graph neural networks)-based attacks are shown to be the most capable tools that attackers can employ as they exploit graph structures inherent to a circuit’s netlist. In this paper,(1)We discuss the inherent structural weaknesses of the logic locking techniques.(2)Knowing these weaknesses, we investigate the challenges of protecting circuits against GNN-based attacks.(3)We propose GNN-resilient Interconnect-based obfuscation (GRIN) and GNN-resilient Gate-based Obfuscation (GREGO) logic locking schemes with learning resilient structures. We evaluate our secure schemes using ISCAS-85 and ITC-99 benchmarks and provide comprehensive security and overhead analysis of our proposed schemes.
Armin Darjani, Nima Kavand, Shubham Rai, Akash Kumar 0001
IEEE Trans. Inf. Forensics Secur.4
2024 Introduction to the FPL 2021 Special Section
abstract
The International Conference on Field-Programmable Logic and Applications (FPL) was the first and remains the largest conference covering the rapidly growing area of field-programmable logic and reconfigurable computing.During the past 30 years, many of the advances in reconfigurable system architectures, applications, embedded processors, and design automation methods and tools were first published in the proceedings of the FPL conference series.The conference objective is to bring together researchers and practitioners from both academia and industry and from around the world.The 31st edition of the FPL (2021) took place from August 30 till September 3, 2021.It is the second FPL conference that had to be organized as a virtual event due to the COVID-19 pandemic.The purpose of this Special Section is to provide an insight into current research and development in aspects related to Field-Programmable Gate Array (FPGA) applications, FPGA technology, and FPGA programming models and tools.This Special Section includes three papers that were presented in the 2021 edition of the conference.The three articles were appropriately selected (based on their quality) to cover various topics of the conference.
Diana Göhringer, Georgios Keramidas, Akash Kumar 0001
ACM Trans. Reconfigurable Technol. Syst.3
2024 AxOMaP: Designing FPGA-based Approximate Arithmetic Operators using Mathematical Programming
abstract
With the increasing application of machine learning (ML) algorithms in embedded systems, there is a rising necessity to design low-cost computer arithmetic for these resource-constrained systems. As a result, emerging models of computation, such as approximate and stochastic computing, that leverage the inherent error-resilience of such algorithms are being actively explored for implementing ML inference on resource-constrained systems. Approximate computing (AxC) aims to provide disproportionate gains in the power, performance, and area (PPA) of an application by allowing some level of reduction in its behavioral accuracy (BEHAV). Using approximate operators (AxOs) for computer arithmetic forms one of the more prevalent methods of implementing AxC. AxOs provide the additional scope for finer granularity of optimization, compared to only precision scaling of computer arithmetic. To this end, the design of platform-specific and cost-efficient approximate operators forms an important research goal. Recently, multiple works have reported the use of AI/ML-based approaches for synthesizing novel FPGA-based AxOs. However, most of such works limit the use of AI/ML to designing ML-based surrogate functions that are used during iterative optimization processes. To this end, we propose a novel data analysis-driven mathematical programming-based approach to synthesizing approximate operators for FPGAs. Specifically, we formulate mixed integer quadratically constrained programs based on the results of correlation analysis of the characterization data and use the solutions to enable a more directed search approach for evolutionary optimization algorithms. Compared to traditional evolutionary algorithms-based optimization, we report up to 21% improvement in the hypervolume, for joint optimization of PPA and BEHAV, in the design of signed 8-bit multipliers. Further, we report up to 27% better hypervolume than other state-of-the-art approaches to DSE for FPGA-based application-specific AxOs.
Siva Satyendra Sahoo, Salim Ullah, Akash Kumar 0001
ACM Trans. Reconfigurable Technol. Syst.3
2023 SyFAxO-GeN: Synthesizing FPGA-Based Approximate Operators with Generative Networks
abstract
With rising trends of moving AI inference to the edge, due to communication and privacy challenges, there has been a growing focus on designing low-cost Edge-AI. Given the diversity of application areas at the edge, FPGA-based systems are increasingly used for high-performance inference. Similarly, approximate computing has emerged as a viable approach to achieve disproportionate resource gains by utilizing the applications' inherent robustness. However, most related research has focused on selecting the appropriate approximate operators for an application from a set of ASIC-based designs. This approach fails to leverage the FPGA's architectural benefits and limits the scope of approximation to already existing generic designs. To this end, we propose an AI-based approach to synthesizing novel approximate operators for FPGA's Look-up-table-based structure. Specifically, we use state-of-the-art generative networks to search for constraint-aware arithmetic operator designs optimized for FPGA-based implementation. With the proposed GANs, we report up to 49% faster training, with negligible accuracy degradation, than related generative networks. Similarly, we report improved hypervolume and increased pareto-front design points compared to state-of-the-art approaches to synthesizing approximate multipliers.
Rohit Ranjan, Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001
ASP-DAC4
2023 Special Session: Mitigating Side-Channel Attacks Through Circuit to Application Layer Approaches
abstract
Side-Channel Attacks (SCAs), which are always considered a severe threat to the security of the cryptographic circuits, today can also be employed to extract IP secrets and neural network models. Hence, developing novel security solutions at different design levels is crucial. In this paper, we explore recent countermeasures at the circuit, algorithmic, and microarchitecture levels. First, we explain how Reconfigurable Field-Effect Transistor (RFET), as a beyond CMOS technology, enables us to provide both IP and data protection against SCAs at the circuit level. Second, we investigate an automated method for generating masked circuits as an algorithmic solution, and then we review machine learning-based SCA detection mechanisms at the microarchitecture level. Finally, we discuss emerging threats of SCAs from the industrial point of view.
Nima Kavand, Armin Darjani, Jens Trommer, Giulio Galderisi, Thomas Mikolajick, Nicolai Müller, Amir Moradi 0001, Chongzhou Fang, Ning Miao, Han Wang 0020, Sai Manoj Pudukotai Dinakarrao, Houman Homayoun, Benjamin Hettwer, Luca Parrini, Akash Kumar 0001
CODES+ISSS15
2023 Discerning Limitations of GNN-based Attacks on Logic Locking
abstract
Machine learning (ML)-based attacks have revealed the possibility of utilizing neural networks to break locked circuits without needing functional chips (Oracle). Among ML approaches, GNN (graph neural networks)-based attacks are the most potent tools that attackers can employ as they exploit graph structures inherent to a circuit’s netlist. Although promising, in this paper, we reveal that GNNs have some impediments in attacking locked circuits. We investigate the limits of the state-of-the-art GNN-based attacks against logic locking and show that we can drastically decrease the accuracy of these attacks by utilizing these limitations in the locking process.
Armin Darjani, Nima Kavand, Shubham Rai, Akash Kumar 0001
DAC4
2023 ADAPTIVE: Agent-Based Learning for Bounding Time in Mixed-Criticality Systems
abstract
In Mixed-Criticality (MC) systems, the high Worst-Case Execution Time (WCET) of a task is a pessimistic bound, the maximum execution time of the task under all circumstances, while the low WCET should be close to the actual execution time of most instances of the task to improve utilization and Quality-of-Service (QoS). Most MC systems consider a static low WCET for each task which cannot adapt to dynamism at run-time. In this regard, we consider the run-time behavior of tasks and propose a learning-based approach that dynamically monitors the tasks’ execution times and adapts the low WCETs to determine the ideal trade-off between mode-switches, utilization, and QoS. Based on our observations on running embedded real-time benchmarks on a real platform, the proposed scheme improves the QoS by 16.4% on average while reducing the utilization waste by 17.7%, on average, compared to state-of-the-art works.
Behnaz Ranjbar, Ali Hosseinghorban, Akash Kumar 0001
DAC3
2023 KeRRaS: Sort-Based Database Query Processing on Wide Tables Using FPGAs
abstract
Sorting is an important operation in database query processing. Complex pipeline-breaking operators (e.g., aggregation and equi-join) become single-pass algorithms on sorted tables. Therefore, sort-based query processing is a popular method for FPGA-based database system acceleration. However, most accelerators have a limit on the table width or the number of columns they can sort. This limit is often set by the width of the data path or the amount of BRAM present on the FPGA. In this paper we propose KeRRaS, an abstract sorting algorithm that enables existing sort-based query processors to support arbitrarily wide tables while offering scalability, preserving modularity, and having low resource overhead. Moreover, we present an implementation of KeRRaS based on morphing sort-merge, a resource-efficient FPGA-based query accelerator. The implementation behaves similarly to morphing sort-merge on narrow tables, and scales well as the number of key columns increases.
Mehdi Moghaddamfar, Christian Färber, Wolfgang Lehner, Akash Kumar 0001
DaMoN4
2023 Motivating Agent-Based Learning for Bounding Time in Mixed-Criticality Systems
abstract
In Mixed-Criticality (MC) systems, the high Worst-Case Execution Time (WCET) of a task is a pessimistic bound, the maximum execution time of the task under all circumstances, while the low WCET should be close to the actual execution time of most instances of the task to improve utilization and Quality-of-Service (QoS). Most MC systems consider a static low WCET for each task which cannot adapt to dynamism at run-time. In this regard, we consider the run-time behavior of tasks and motivate to propose a learning-based approach that dynamically monitors the tasks' execution times and adapts the low WCETs to determine the ideal trade-off between mode-switches, utilization, and QoS. Based on our observations on running embedded real-time benchmarks on a real platform, the proposed scheme reduces the utilization waste by 47.2%, on average, compared to state-of-the-art works.
Behnaz Ranjbar, Ali Hosseinghorban, Akash Kumar 0001
DATE3
2023 Learning-Oriented Reliability Improvement of Computing Systems From Transistor to Application Level
abstract
Due to technology scaling in modern computing platforms, the safety and reliability issues have increased tremendously, which often accelerate aging, lead to permanent faults, and cause unreliable execution of applications. Failure in some computing systems like avionics may cause catastrophic consequences. Therefore, managing reliability under all circumstances of stress and environmental changes is crucial in all abstraction layers, from application to transistor levels. Machine learning techniques are recently being employed for dynamic reliability estimation and optimization. They can adapt to varying workloads and system conditions. This paper presents reliability improvement approaches from multiple perspectives-from transistor-level to application-level-and discusses their effectiveness and limitations as well as open challenges.
Behnaz Ranjbar, Florian Klemme, Paul R. Genssler, Hussam Amrouch, Jinhyo Jung, Shail Dave, Hwisoo So, Kyongwoo Lee, Aviral Shrivastava, Ji-Yung Lin, Pieter Weckx, Subrat Mishra, Francky Catthoor, Dwaipayan Biswas, Akash Kumar 0001
DATE15
2023 Design Enablement Flow for Circuits with Inherent Obfuscation based on Reconfigurable Transistors
abstract
Reconfigurable transistors are a new emerging type of device, which offer the promise to improve the resistance of electronic components against know-how theft. In order to enable a product development of such an emerging device, a cross-layer design enablement strategy is needed, as emerging technologies are not necessarily compatible withstandard tools used in the industry. In ‘CirroStrato’, we aim on the development of such a complete flow enabling CMOS co-integration of reconfigurable transistors, ranging from process adjustments, device modeling, library characterization, physical and logical synthesis up towards sophisticated hardware security tests. In this multi-partner-project (MPP) paper, our aim is to elucidate the overall design enablement flow, as well as current research challenges on the individual stages.
Jens Trommer, Niladri Bhattacharjee, Thomas Mikolajick, Sebastian Huhn 0001, Marcel Merten, Mohammed E. Djeridane, Muhammad Hassan 0002, Rolf Drechsler, Shubham Rai, Nima Kavand, Armin Darjani, Akash Kumar 0001, Violetta Sessi, M. Drescher, S. Kolodinski, M. Wiatr
DATE12
2023 High-Throughput Approximate Multiplication Models in PyTorch
abstract
Approximate multipliers can reduce the resource consumption of neural network accelerators. To study their effects on an application, they need to be simulated during network training. We develop simulation models for a common class of approximate multipliers. Our models speed up execution by replacing time-consuming type conversions and memory accesses with fast floating-point arithmetic. Across six different neural network architectures, these models increase throughput by 2.7× over the commonly used array lookup while recreating behavioral simulation with high fidelity.
Elias Trommer, Bernd Waschneck, Akash Kumar 0001
DDECS3
2023 A Study of Early Aggregation in Database Query Processing on FPGAs
abstract
In database query processing, aggregation is an operator by which data with a common property is grouped and expressed in a summary form. Early aggregation is a popular method for improving the performance of the aggregation operator. In this paper, we study early aggregation algorithms in the context of query processing acceleration in database systems on FPGAs. The comparative study leads us to set-associative caches with a low inter-reference recency set (LIRS) replacement policy. They show both great performance and modest implementation complexity compared to some of the most prominent early aggregation algorithms. We also present a novel application-specific architecture for implementing set-associative caches. Benchmarks of our implementation show speedups of up to 3x for end-to-end aggregation compared to a state-of-the-art FPGA-based query engine.
Mehdi Moghaddamfar, Norman May, Christian Färber, Wolfgang Lehner, Akash Kumar 0001
FPGA5
2023 CoOAx: Correlation-aware Synthesis of FPGA-based Approximate Operators
abstract
The run-time reconfigurability and high parallelism offered by Field Programmable Gate Arrays (FPGAs) make them an attractive choice for implementing hardware accelerators for Machine Learning (ML) algorithms. In the quest for designing efficient FPGA-based hard-ware accelerators for ML algorithms, the inherent error-resilience of ML algorithms can be exploited to implement approximate hard-ware accelerators to trade the output accuracy with better over-all performance. As multiplication and addition are the two main arithmetic operations in ML algorithms, most state-of-the-art approximate accelerators have considered approximate architectures for these operations. However, these works have mainly considered the exploration and selection of approximate operators from an existing set of operators. To this end, we provide an efficient methodology for synthesizing and implementing novel approximate operators. Specifically, we propose a novel operator synthesis approach that supports multiple operator algorithms to provide new approximate multiplier and adder designs for AI inference applications. We report up to 27% and 25% lower power than state-of-the-art approximate designs, with equivalent error behavior, for 8-bit unsigned adders and 4-bit signed multipliers respectively. Further, we propose a correlation-aware Design Space Exploration (DSE) method that can improve the efficacy of randomized search algorithms in synthesizing novel approximate operators.
Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI3
2023 Reconfigurable FET Approximate Computing-based Accelerator for Deep Learning Applications
abstract
Reconfigurable nanotechnologies such as Silicon Nanowire Field Effect Transistors (FETs) serve as a promising technology that not only facilitates lower power consumption but also supports multi-functionality through reconfigurability. It enables reconfigurability and supports multiple functionalities per computational unit. These features motivate us to design a novel state-of-the-art energy-efficient hardware accelerator for implementing memory-intensive applications including convolutional neural networks (CNNs) and deep neural networks (DNNs). To accelerate the computations, we design Multiply and Accumulate (MAC) units to perform the computations. For the design of MACs, we employ Silicon nanowire reconfigurable FETs (RFETs). The use of RFETs leads to nearly 70% power reduction compared to the traditional CMOS implementation and also reduced latency in performing the computations. Further to optimize the overheads and improve memory efficiency, we introduce a novel approximation technique for RFETs. The RFET-based approximate adders lead to reduced power, area, and delay while having a minimal impact on the accuracy of the DNN/CNN. In addition, we carry out a detailed study of varied combinations of architectures involving CMOS, RFETs, accurate adders, and approximate adders to demonstrate the benefits of the proposed RFET-based approximate acclerator. The proposed RFET-based accelerator achieves an accuracy of 94% on MNIST datasets with 93% and 73%reduction in the area, power and delay metrics respectively compared to the state-of-the-art hardware accelerator architectures.
Raghul Saravanan, Sathwika Bavikadi, Shubham Rai, Akash Kumar 0001, Sai Manoj Pudukotai Dinakarrao
ISCAS4
2023 RAPID: Approximate Pipelined Soft Multipliers and Dividers for High Throughput and Energy Efficiency
abstract
The rapid updates in error-resilient applications along with their quest for high throughput has motivated designing fast approximate functional units for field-programmable gate arrays (FPGAs). Studies have proposed various imprecise functional techniques, albeit posed with three shortcomings: first, most existing inexact multipliers and dividers are specialized for application-specific integrated circuit (ASIC) platforms. Therefore, due to the architectural differences of underlying building blocks in FPGA and ASIC, ASIC-customized designs have not yielded comparable improvements when directly synthesized and ported to FPGAs. Second, state-of-the-art (SoA) approximate units are substituted, mostly in a single kernel of a multikernel application. Moreover, the end-to-end assessment is adopted on the quality of results (QoR), but not on the overall gained performance. Finally, the existing imprecise components are not designed to support a pipelined approach, which could boost the operating frequency/throughput of, e.g., division-included applications. In this article, we propose RAPID, the first pipelined approximate multiplier and divider architectures, customized for FPGAs. The proposed units efficiently utilize 6-input look-up tables (6-LUTs) and fast carry chains to implement Mitchell’s approximate algorithms. Our novel error-refinement scheme not only has negligible overhead over the baseline Mitchell’s approach but also boosts its accuracy to 99.4% for arbitrary size of multiplication and division. Experimental results obtained with Xilinx Vivado demonstrate the efficiency of the proposed pipelined and nonpipelined RAPID multipliers and dividers over accurate counterparts. In particular, the 4-stage pipelined architecture of a 32-bit RAPID multiplier (divider) enables$3.3\times $($5.1\times $) higher throughput,$2.3\times $($6.8\times $) higher throughput/Watt, and 52% (31%) savings of look-up tables (LUTs), over their 4-stage pipelined, accurate Intellectual Property (IP) counterparts. Moreover, the end-to-end evaluations of nonpipelined RAPID, deployed in three multikernel applications in the domains of biosignal processing, image processing, and moving object tracking for unmanned aerial vehicles (UAVs) indicate up to 35%, 33%, and 45% improvements in area, latency, and area-delay-product (ADP), respectively, over accurate kernels, with negligible loss in QoR. To springboard future research in reconfigurable and approximate computing communities, our implementations will be available and opensourced athttps://cfaed.tu-dresden.de/pd-downloads.
Zahra Ebrahimi, Muhammad Zaid, Mark Wijtvliet, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 SeqL+: Secure Scan-Obfuscation With Theoretical and Empirical Validation
abstract
Scan-obfuscation is a powerful methodology to protect Silicon-based intellectual property from theft. Prior work on scan-obfuscation in the context of logic-locking have unique limitations, which are addressed by our previous work, SeqL, which looks at functional output corruption to obfuscate scan-chains, but is unable to resist removal attacks on circuits with inadequate number of flip-flops without feedback. To address this issue, we propose to scramble flip-flops with feedback to increase key length without introducing further vulnerabilities. This study reveals the first formulation and complexity analysis of Boolean satisfiability (SAT)-based attack on scan-scrambling. We formulate the attack as a conjunctive normal form (CNF) using a worst-case$\mathcal {O}(n^{3})$reduction in terms of scramble-graph size$n$. In order to defeat SAT-based attack, we propose an iterative swapping-based scan-cell scrambling algorithm that has$\mathcal {O}(n)$implementation time-complexity and$\mathcal {O}(2^{\lfloor ({\alpha.n+1}/{3}) \rfloor })$SAT-decryption time-complexity in terms of a user-configurable cost constraint$\alpha ~(0 < \alpha \le 1)$.
Seetal Potluri, Shamik Kundu, Akash Kumar 0001, Kanad Basu, Aydin Aysu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Utilizing XMG-Based Synthesis to Preserve Self-Duality for RFET-Based Circuits
abstract
Individual transistors based on emerging reconfigurable nanotechnologies exhibit electrical conduction for both types of charge carriers. These transistors [referred to as reconfigurable field-effect transistors (RFETs)] enable dynamic reconfiguration to demonstrate either a p- or an n-type functionality. This duality of functionality at the transistor level is efficiently abstracted as a self-dual Boolean logic, that can be physically realized with fewer RFET transistors compared to the contemporary CMOS technology. Consequently, to achieve better area reduction for RFET-based circuits, the self-duality of a given circuit should be preserved during logic optimization and technology mapping. In this article, we specifically aim to preserve self-duality by using Xor-majority graphs (XMGs) as the logic representation during logic synthesis and technology mapping. We propose a synthesis flow that uses new restructuring techniques, called rewriting and resubstitution for XMGs to preserve self-duality during technology-independent logic synthesis. For technology mapping, we use a novel open-source and a logic-representation agnostic mapping tool. Using the above-proposed XMG-based flow, we demonstrate its benefits by comparing post-mapping areas for synthetic and cryptographic benchmarks with three different synthesis flows: 1) AIG-based optimization and AIG-based mapping; 2) XMG-based optimization with AIG-based mapping; and 3) AIG-based optimization with logic-representation agnostic mapping. Our experiments show that the proposed XMG-based flow efficiently preserves self-duality and achieves the best area results for RFET-based circuits (up to 12.36% area reduction) with respect to the baseline.
Shubham Rai, Alessandro Tempia Calvino, Heinz Riener, Giovanni De Micheli, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 AxOTreeS: A Tree Search Approach to Synthesizing FPGA-based Approximate Operators
abstract
Approximate computing (AxC) provides the scope for achieving disproportionate gains in a system’s power, performance, and area (PPA) metrics by leveraging an application’s inherent error-resilient behavior (BEHAV). Trading computational accuracy for performance gains makes AxC an attractive proposition for implementing computationally complex AI/ML-based applications on resource-constrained embedded systems. The growing diversity of application domains using AI/ML has also led to the increasing usage of FPGA-based embedded systems. However, implementing AxC for FPGAs has primarily been limited to the post-processing of ASIC-optimized approximate operators (AxOs). This approach usually involves selecting from a set of AxOs that have been optimized for a gate-based implementation in an ASIC. While such an approach does allow leveraging existing knowledge of ASIC-based AxO design, it limits the scope for considering the challenges and opportunities associated with FPGA’s LUT-based computation structures. Similarly, the few works considering the LUT-based computing for AxO design use generic optimization approaches that do not allow integrating problem-specific prior knowledge—empirical and/or statistical. To this end, we propose a novel tree search-based approach to AxO synthesis for FPGAs. Specifically, we present a design methodology using Monte Carlo Tree Search (MCTS)-based search tree traversal that allows the designer to integrate statistical data, such as correlation, into the AxOs optimization. With the proposed methods, we report improvements over standard MCTS algorithm-based results as well as improved hypervolume for both operator-level and application-specific DSE, compared to state-of-the-art design methodologies.
Siva Satyendra Sahoo, Salim Ullah, Akash Kumar 0001
ACM Trans. Embed. Comput. Syst.3
2023 ACM TECS Special Issue on Embedded System Security Tutorials
abstract
No abstract available.
Aviral Shrivastava, Jian-Jia Chen, Akash Kumar 0001, Anup Das 0001
ACM Trans. Embed. Comput. Syst.3
2022 A Versatile Mapping Approach for Technology Mapping and Graph Optimization
abstract
This paper proposes a versatile mapping approach that has three objectives: i) it can map from one technology-independent graph representation to another; ii) it can map to a cell library; iii) it supports logic rewriting. The method is cut-based, mitigates logic-sharing issues of previous graph mapping approaches, and exploits structural hashing. The mapper is the first one of its kind to support remapping among various graph representations, thus enabling specialized mapping to emerging technologies (such as AQFP) and for security applications (such as XAG-based design). We show that mapping to MIGs improves area by 10% as compared to the state of the art, and that technology mapping is 18% faster than ABC with slightly better results.
Alessandro Tempia Calvino, Heinz Riener, Shubham Rai, Akash Kumar 0001, Giovanni De Micheli
ASP-DAC4
2022 Multi-Precision Deep Neural Network Acceleration on FPGAs
abstract
Quantization is a promising approach to reduce the computational load of neural networks. The minimum bit-width that preserves the original accuracy varies significantly across different neural networks and even across different layers of a single neural network. Most existing designs over-provision neural network accelerators with sufficient bit-width to preserve the required accuracy across a wide range of neural networks. In this paper, we present mpDNN, a multi-precision multiplier with dynamically adjustable bit-width for deep neural network acceleration. The design supports run-time splitting an arithmetic operator into multiple independent operators with smaller bit-width, effectively increasing throughput when lower precision is required. The proposed architecture is designed for FPGAs, in that the multipliers and bit-width adjustment mechanism are optimized for the LUT-based structure of FPGAs. Experimental results show that by enabling run-time precision adjustment, mpDNN can offer 3-15x improvement in throughput.
Negar Neda, Salim Ullah, Azam Ghanbari, Hoda Mahdiani, Mehdi Modarressi, Akash Kumar 0001
ASP-DAC6
2022 DELTA: DEsigning a stealthy trigger mechanism for analog hardware trojans and its detection analysis
abstract
This paper presents a stealthy triggering mechanism that reduces the dependencies of analog hardware Trojans on the frequent toggling of the software-controlled rare nets. The trigger to activate the Trojan is generated by using a glitch generation circuit and a clock signal, which increases the selectivity and feasibility of the trigger signal. The proposed trigger is able to evade the state-of-the-art run-time detection (R2D2) and Built-In Acceleration Structure (BIAS) schemes. Furthermore, the simulation results show that the proposed trigger circuit incurs a minimal overhead in side-channel footprints in terms of area (29 transistors), delay (less than 1ps in the clock cycle), and power (1μW).
Mohil Sandip Desai, Mark Wijtvliet, Shubham Rai, Akash Kumar 0001
DAC5
2022 Exploring Standard-Cell Designs for Reconfigurable Nanotechnologies: A Formal Approach
abstract
Standard-cell design has always been a craft, and common field-effect transistors span only a small design space. This has changed with reconfigurable transistors. Boolean functions that exhibit multiple dual product-terms in their sum-of-product form yield various beneficial circuit implementations with recon-figurable transistors. In this work, we present an approach to automatically generate these implementations through a formal modeling approach. Using the 3-input XOR function as an example, we discuss the variations and show how to quantify properties like worst-case delay and power dissipation, as well as averages of delay and energy consumption per operation over different scenarios. The quantification runs fully automated on charge transport network models employing probabilistic model checking. This yields exact results instead of approximations obtained from experiments and sampling. The highlight of our work is that the proposed approach provides a comprehensive early technology evaluation flow.
Michael Raitza, Steffen Märcker, Shubham Rai, Akash Kumar 0001
DATE4
2022 Improving Technology Mapping for And-Inverter-Cones
abstract
AND-inverter-cones (AICs), proposed in 2012, offer a suitable alternative to Look-Up-Tables (LUTs) as the basic building block for FPGAs. They support tapping of multiple side outputs and are intrinsically fracturable which favours reduction of logic duplication. Unlike${k-inputs}$LUTs, their area scales linearly with the number of inputs. Technology mapping is one of the crucial tasks to realize the full power of AIC-based FPGAs. However, the current state-of-the-art implementations suffers two main drawbacks as they do not account for the AIC properties fully: (i) The required time set for each node is suboptimal in the context of AIC and that impairs the mapping quality; (ii) they rely on priority cuts, which are unnecessarily runtime-intensive in the context of AIC mapping. To improve the mapping quality, we propose and proof a new method to calculate the maximal required time for each node purely based on its graph depth and height. We propose an asymptotically runtime-optimal in-memory direct cut selection method which leads to similar area numbers (~ 1% area overhead) as our reference priority cut implementation. Combining these improvements with a second area recovery round leads to a final area reduction of 16.4% and 3% for the MCNC and VTR benchmarks respectively as compared to our reference implementation of the latest known technology mapper, while leaving the delay unaltered.
Martin Thümmler, Shubham Rai, Akash Kumar 0001
DATE3
2022 A Hybrid Scheduling Mechanism for Multi-programming in Mixed-Criticality Systems
abstract
In the last decade, the rapid evolution of the Commercial-Off-The-Shelf (COTS) platforms led safety-critical systems towards integrating tasks and applications with different criticality levels in a shared hardware platform, i.e., Mixed-Criticality Systems (MCS)s. Therefore, several scheduling algorithms and approaches have been proposed upon a commonly used model, i.e., Vestal's model. However, consolidating software functions onto shared processors cannot be implemented directly in real-life applications and industrial systems while complying with certification requirements. The existing scheduling approaches do not provide a simple solution for eliminating the interference effect among the tasks with different criticality levels on the shared processing resources. Moreover, the system mode switch guarantees the timing constraints of the high-criticality tasks throw the termination of the low-criticality tasks. In this paper, we developed a new scheduling algorithm that addresses these challenges based on the round-robin technique, which improves the overall schedulability. We compared the proposed algorithm against existing scheduling algorithms in both academia and industry using extensive experiments to evaluate it. Our results show improvements in the schedulability from 0.8% to 14.0% and from 2.7% to 10.7% compared to the conventional Earliest Deadline First with Virtual Deadline (EDF-VD) and Fixed Priority Preemptive (FPP) scheduling approaches, respectively.
Mohammad Bawatna, Behnaz Ranjbar, Akash Kumar 0001
DSD3
2022 PosAx-O: Exploring Operator-level Approximations for Posit Arithmetic in Embedded AI/ML
abstract
The quest for low-cost embedded AI/ML applications has motivated innovations across multiple abstractions of the computation stack. Novel approaches for arithmetic operations have primarily involved quantization, precision-scaling, approximations, and modified data representation. In this context, Posit has emerged as an alternative to the IEEE-754 standard as it offers multiple benefits, primarily due to its dynamic range and tapered precision. However, the implementation of Posit arithmetic operations tends to result in high resource utilization and power dissipation. Consequently, recent works have delved into the idea of exploiting the error resilience of machine learning algorithms by using low-precision Posit arithmetic. However, limiting the exploration to precision-scaling limits the scope for application-specific optimizations for embedded AI/ML applications. To this end, we explore operator-level optimizations and approximations for low-precision Posit numbers. Specifically, we identify and eliminate redundant operations in state-of-the-art Posit arithmetic operator designs and provide a modular framework for exploring approximations in various stages of the computation. We also present a novel framework for behaviorally testing the corresponding Posit approximate designs in Artificial Neural Networks. The proposed optimizations and approximations exhibit considerable resource improvements with a small error in many cases. For instance, a Posit-based multiplier with 1-bit reduced precision shows a 33% improvement in power and utilization, with only a 0.2% degradation in overall accuracy.
Amritha Immaneni, Salim Ullah, Suresh Nambi, Siva Satyendra Sahoo, Akash Kumar 0001
DSD5
2022 FPGA-Based Database Query Processing on Arbitrarily Wide Tables
abstract
Thanks to the flexibility of FPGAs and their widespread adoption in the cloud, they have become attractive solutions for the acceleration of resource- and memory-intensive database workloads. Complex pipeline-breaking operators (e.g., aggregation, join) often constitute most of the execution time of the queries involved in these workloads. A popular approach in processing these operators is by pre-sorting the input, as they become single-pass algorithms on sorted tables [1] .
Mehdi Moghaddamfar, Christian Färber, Norman May, Wolfgang Lehner, Akash Kumar 0001
FCCM5
2022 ERMES: Efficient Racetrack Memory Emulation System based on FPGA
abstract
With the scaling of CMOS technology almost over, non-volatile memories based on emerging technologies are gaining considerable popularity. Particularly, spintronic-based Racetrack memories (RTMs) exhibit unprecedented storage capacity, as well as reduced energy per operation and high write endurance, which make them promising candidates to revolutionize the architecture of memory sub-systems. However, since RTM exploits shifting of magnetic domains to align the required data with the access port, its read/write latency is not constant. Due to this behaviour, several performance optimizations related to the target application may be introduced either on memory architecture or data placement or both. To this purpose, specific tools able to emulate the timing characteristics of RTMs are highly desired. Unfortunately, existing software-based simulators show poor flexibility and run-time. To address such limitations, this paper presents a new emulation system for RTMs based on heterogeneous FPGA-CPU Systems-on-Chips (SoCs). Thanks to its high flexibility, the proposed emulator can be easily configured to evaluate different memory architectures. In addition, the CPU can be used to stimulate the RTM architecture under test with appropriate benchmarks, thus providing a fast self-contained evaluation environment. As case study, ERMES has been implemented within the Xilinx Zynq Ultrascale XCUZ9EG SoC to evaluate performances of several memory configurations when running benchmark applications from the MiBench suite, experiencing a speed-up higher than × 146 over software-based simulators.
Fanny Spagnolo, Salim Ullah, Pasquale Corsonello, Akash Kumar 0001
FPL4
2022 NetPU: Prototyping a Generic Reconfigurable Neural Network Accelerator Architecture
abstract
FPGA-based Neural Network (NN) accelerator is a rapidly advancing subject in recent research. Related works can be classified as two hardware architectures: i) Heterogeneous Streaming Dataflow (HSD) architecture and ii) Processing Element Matrix (PEM) architecture. HSD architecture explores the reconfigurability of FPGAs to support the customization and optimization of hardware design to implement a complete network on FPGA for one given trained model. PEM architecture achieves relatively generic support for different network models, essentially implementing the neuron processing modules on the FPGA scheduled by the runtime software environment. In summary, the HSD architecture requires more resources with simplified runtime software control. The PEM architecture consumes fewer resources than the HSD architecture. However, the runtime software environment can be a heavy payload for lightweight systems, such as the low-power microcontroller of IoT or edge devices.
Shubham Rai, Salim Ullah, Akash Kumar 0001
FPT4
2022 ENTANGLE: An Enhanced Logic-locking Technique for Thwarting SAT and Structural Attacks
abstract
Among the SAT-resilient logic locking techniques, the Stripped-Functionality-Logic-Locking (SFLL) is the most promising solution which can guard the intellectual property against approximate, sensitization, SAT, and structural attacks which target Point-function techniques. However, even the SFLL technique has been shown to be vulnerable to a recent class of structural attacks that identify the perturbation logic. In this paper, we first categorize all possible classes of attacks on SFLL. Then we propose ENTANGLE a novel logic locking technique built upon SFLL that can resist all of these attacks, including the emerging ML-Based attacks. We test our technique against publicly available SFLL attacks. The implementation results show that ENTANGLE can secure large-sized industrial circuits with an average overhead of 11.6 percent and 9.1 percent for area and power, respectively.
Armin Darjani, Nima Kavand, Shubham Rai, Mark Wijtvliet, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI5
2022 Securing Hardware through Reconfigurable Nano-Structures
abstract
Hardware security has been an ever-growing concern of the integrated circuit (IC) designers. Through different stages in the IC design and life cycle, an adversary can extract sensitive design information and private data stored in the circuit using logical, physical, and structural weaknesses. Besides, in recent times, ML-based attacks have become the new de facto standard in hardware security community. Contemporary defense strategies are often facing unforeseen challenges to cope up with these attack schemes. Additionally, the high overhead of the CMOS-based secure addon circuitry and intrinsic limitations of these devices indicate the need for new nano-electronics. Emerging reconfigurable devices like Reconfigurable Field Effect transistors (RFETs) provide unique features to fortify the design against various threats at different stages in the IC design and life cycle. In this manuscript, we investigate the applications of the RFETs for securing the design against traditional and machine learning (ML)-based intellectual property (IP) piracy techniques and side-channel attacks (SCAs).
Nima Kavand, Armin Darjani, Shubham Rai, Akash Kumar 0001
ICCAD4
2022 Combining Gradients and Probabilities for Heterogeneous Approximation of Neural Networks
abstract
This work explores the search for heterogeneous approximate multiplier configurations for neural networks that produce high accuracy and low energy consumption. We discuss the validity of additive Gaussian noise added to accurate neural network computations as a surrogate model for behavioral simulation of approximate multipliers. The continuous and differentiable properties of the solution space spanned by the additive Gaussian noise model are used as a heuristic that generates meaningful estimates of layer robustness without the need for combinatorial optimization techniques. Instead, the amount of noise injected into the accurate computations is learned during network training using backpropagation. A probabilistic model of the multiplier error is presented to bridge the gap between the domains; the model estimates the standard deviation of the approximate multiplier error, connecting solutions in the additive Gaussian noise space to actual hardware instances. Our experiments show that the combination of heterogeneous approximation and neural network retraining reduces the energy consumption for multiplications by 70% to 79% for different ResNet variants on the CIFAR-10 dataset with a Top-1 accuracy loss below one percentage point. For the more complex Tiny ImageNet task, our VGG16 model achieves a 53 % reduction in energy consumption with a drop in Top-5 accuracy of 0.5 percentage points. We further demonstrate that our error model can predict the parameters of an approximate multiplier in the context of the commonly used additive Gaussian noise (AGN) model with high accuracy. Our software implementation is available under https://github.com/etrommer/agn-approx.
Elias Trommer, Bernd Waschneck, Akash Kumar 0001
ICCAD3
2022 Toward the Design of Fault-Tolerance-Aware and Peak-Power-Aware Multicore Mixed-Criticality Systems
abstract
Mixed-criticality (MC) systems have recently been devised to address the requirements of real-time systems in industrial applications, where the system runs tasks with different criticality levels on a single platform. In some workloads, a high-critically task might overrun and overload the system, or a fault can occur during the execution. However, these systems must be fault tolerant and guarantee the correct execution of all high-criticality (HC) tasks by their deadlines to avoid catastrophic consequences, in any situation. Furthermore, in these MC systems, the peak-power consumption of the system may increase, especially in an overload situation and exceed the processor thermal design power (TDP) constraint. This may cause generating heat beyond the cooling capacity, resulting the system stop to avoid excessive heat and halting the processor. In this article, we propose a technique for dependent dual-criticality tasks in fault-tolerant multicore MC systems to manage peak-power consumption and temperature. The technique develops a tree of possible task mapping and scheduling at design-time to cover all possible scenarios and reduce the low-criticality task drop rate in the HC mode. At the runtime, the system exploits the tree to select a proper schedule according to fault occurrences and criticality mode changes. Experimental results show that the average task schedulability is 74.14% on average for the proposed method, while the peak-power consumption and maximum temperature are improved by 16.65% and 14.9 °C on average, respectively, compared to a recent work. In addition, for a real-life application, our method reduces the peak power and maximum temperature by up to 20.06% and 5 °C, respectively, compared to a state-of-the-art approach.
Behnaz Ranjbar, Ali Hosseinghorban, Alireza Ejlali, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 BOT-MICS: Bounding Time Using Analytics in Mixed-Criticality Systems
abstract
An increasing trend for reducing cost, space, and weight leads to modern embedded systems that execute multiple tasks with different criticality levels on a common hardware platform while guaranteeing a safe operation. In such mixed-criticality (MC) systems, multiple worst case execution times (WCETs) are defined for each task, corresponding to the system operation mode to improve the MC system’s timing behavior at runtime. Determining the appropriate WCETs for lower criticality (LC) modes is nontrivial. On the one hand, considering a very low WCET for tasks can improve the processor utilization by scheduling more tasks in that mode, on the other hand, using a larger WCET ensures that the mode switches (which causes by task overrunning) are minimized, thereby improving the quality of service for all tasks, albeit at the cost of processor utilization. Hitherto, no analytical solutions are proposed to determine WCETs in LC modes. In this regard, we propose a scheme to determine WCETs by the Chebyshev theorem, to make a tradeoff between the number of scheduled tasks at design-time and the number of dropped low-criticality tasks at runtime as a result of frequent mode switches. To have a tight bound of execution times and mode switching probability, we also propose a distribution analytics-based scheme, in which the mode switching probability is obtained based on the cumulative distribution function. Our experimental results show that our scheme improves the utilization of state-of-the-art MC systems by up to 72.27%, while maintaining 24.28% mode switching probability in the worst case scenario. Besides, the results of running embedded real-time benchmarks on a real platform show that the distribution-based scheme can improve the utilization by 7.30% while bounding the mode switching probability by 4.85% more, compared to the Chebyshev-based scheme.
Behnaz Ranjbar, Ali Hosseinghorban, Siva Satyendra Sahoo, Alireza Ejlali, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 High-Performance Accurate and Approximate Multipliers for FPGA-Based Hardware Accelerators
abstract
Multiplication is one of the widely used arithmetic operations in a variety of applications, such as image/video processing and machine learning. FPGA vendors provide high-performance multipliers in the form of DSP blocks. These multipliers are not only limited in number and have fixed locations on FPGAs but can also create additional routing delays and may prove inefficient for smaller bit-width multiplications. Therefore, FPGA vendors additionally provide optimized soft IP cores for multiplication. However, in this work, we advocate that these soft multiplier IP cores for FPGAs still need better designs to provide high-performance and resource efficiency. Toward this, we present generic area-optimized, low-latency accurate, and approximate softcore multiplier architectures, which exploit the underlying architectural features of FPGAs, i.e., lookup table (LUT) structures and fast-carry chains to reduce the overall critical path delay (CPD) and resource utilization of multipliers. Compared to Xilinx multiplier LogiCORE IP, our proposed unsigned and signed accurate architecture provides up to 25% and 53% reduction in LUT utilization, respectively, for different sizes of multipliers. Moreover, with our unsigned approximate multiplier architectures, a reduction of up to 51% in the CPD can be achieved with an insignificant loss in output accuracy when compared with the LogiCORE IP. For illustration, we have deployed the proposed multiplier architecture in accelerators used in image and video applications, and evaluated them for area and performance gains. Our library of accurate and approximate multipliers is opensource and available online athttps://cfaed.tu-dresden.de/pd-downloadsto fuel further research and development in this area, facilitate reproducible research, and thereby enabling a new research direction for the FPGA community.
Salim Ullah, Semeen Rehman, Muhammad Shafique 0001, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Blocks: Challenging SIMDs and VLIWs With a Reconfigurable Architecture
abstract
Demand for coarse grain reconfigurable architectures (CGRAs) has significantly increased in recent years as architectures need to be both energy efficient and flexible. However, most CGRAs are optimized for performance instead of energy efficiency. In this work, a novel paradigm for reconfigurable architectures, Blocks, is presented. Blocks uses two separate circuit-switched networks, one for control and one for the data path. This enables the runtime construction of energy-efficient application-specific VLIW-SIMD processors on a reconfigurable fabric. Its energy efficiency is demonstrated by comparing Blocks to four reference architectures, a VLIW, an SIMD, a commercial low-power microprocessor, and a traditional CGRA. All comparisons are based on commercial low-power 40-nm CMOS layout, including memories. Results show that Blocks can achieve a mean total energy reduction of$2.05\times $,$1.84\times $,$8.01\times $, and$1.22\times $over a VLIW, an SIMD, an energy-efficient microprocessor and a traditional CGRA, respectively. At the same time, Blocks delivers equal or higher performance per area due to its ability to adapt to applications by reconfiguration.
Mark Wijtvliet, Akash Kumar 0001, Henk Corporaal
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 AppAxO: Designing Application-specific Approximate Operators for FPGA-based Embedded Systems
abstract
Approximate arithmetic operators, such as adders and multipliers, are increasingly used to satisfy the energy and performance requirements of resource-constrained embedded systems. However, most of the available approximate operators have an application-agnostic design methodology, and the efficacy of these operators can only be evaluated by employing them in the applications. Furthermore, the various available libraries of approximate operators do not share any standard approximation-induction policy to design new operators according to an application’s accuracy and performance constraints. These limitations also hinder the utilization of machine learning models to explore and determine approximate operators according to an application’s requirements. In this work, we present a generic design methodology for implementing FPGA-based application-specific approximate arithmetic operators. Our proposed technique utilizes lookup tables and carry-chains of FPGAs to implement approximate operators according to the input configurations. For instance, for an \( \text{M}\times \text{N} \) accurate multiplier utilizing K lookup tables, our methodology utilizes K -bit configurations to design \( 2^K \) approximate multipliers. We then utilize various machine learning models to evaluate and select configurations satisfying application accuracy and performance constraints. We have evaluated our proposed methodology for three benchmark applications, i.e., biomedical signal processing, image processing, and ANNs. We report more non-dominated approximate multipliers with better hypervolume contribution than state-of-the-art designs for these benchmark applications with the proposed design methodology.
Salim Ullah, Siva Satyendra Sahoo, Nemath Ahmed, Debabrata Chaudhury, Akash Kumar 0001
ACM Trans. Embed. Comput. Syst.5
2022 Plasticine: A Cross-layer Approximation Methodology for Multi-kernel Applications through Minimally Biased, High-throughput, and Energy-efficient SIMD Soft Multiplier-divider
abstract
The rapid evolution of error-resilient programs intertwined with their quest for high throughput has motivated the use of Single Instruction, Multiple Data (SIMD) components in Field-Programmable Gate Arrays (FPGAs). Particularly, to exploit the error-resiliency of such applications, Cross-layer approximation paradigm has recently gained traction, the ultimate goal of which is to efficiently exploit approximation potentials across layers of abstraction. From circuit- to application-level, valuable studies have proposed various approximation techniques, albeit linked to four drawbacks: First, most of approximate multipliers and dividers operate only in SISD mode. Second, imprecise units are often substituted, merely in a single kernel of a multi-kernel application, with an end-to-end analysis in Quality of Results (QoR) and not in the gained performance. Third, state-of-the-art (SoA) strategies neglect the fact that each kernel contributes differently to the end-to-end QoR and performance metrics. Therefore, they lack in adopting a generic methodology for adjusting the approximation knobs to maximize performance gains for a user-defined quality constraint. Finally, multi-level techniques lack in being efficiently supported, from application-, to architecture-, to circuit-level, in a cohesive cross-layer hierarchy. In this article, we propose Plasticine , a cross-layer methodology for multi-kernel applications, which addresses the aforementioned challenges by efficiently utilizing the synergistic effects of a chain of techniques across layers of abstraction. To this end, we propose an application sensitivity analysis and a heuristic that tailor the precision at constituent kernels of the application by finding the most tolerable degree of approximations for each of consecutive kernels, while also satisfying the ultimate user-defined QoR. The chain of approximations is also effectively enabled in a cross-layer hierarchy, from application- to architecture- to circuit-level, through the plasticity of SIMD multiplier-dividers, each supporting dynamic precision variability along with hybrid functionality. The end-to-end evaluations of Plasticine on three multi-kernel applications employed in bio-signal processing, image processing, and moving object tracking for Unmanned Air Vehicles (UAV) demonstrate 41%–64%, 39%–62%, and 70%–86% improvements in area, latency, and Area-Delay-Product (ADP), respectively, over 32-bit fixed precision, with negligible loss in QoR. To springboard future research in reconfigurable and approximate computing communities, our implementations will be available and open-sourced at https://cfaed.tu-dresden.de/pd-downloads.
Zahra Ebrahimi, Dennis Klar, Mohammad Aasim Ekhtiyar, Akash Kumar 0001
ACM Trans. Design Autom. Electr. Syst.4
2021 Efficient Accuracy Recovery in Approximate Neural Networks by Systematic Error Modelling
abstract
Approximate Computing is a promising paradigm for mitigating the computational demands of Deep Neural Networks (DNNs), by leveraging DNN performance and area, throughput or power. The DNN accuracy, affected by such approximations, can be then effectively improved through retraining. In this paper, we present a novel methodology for modelling the approximation error introduced by approximate hardware in DNNs, which accelerates retraining and achieves negligible accuracy loss. To this end, we implement the behavioral simulation of several approximate multipliers and model the error generated by such approximations on pre-trained DNNs for image classification on CIFAR10 and ImageNet. Finally, we optimize the DNN parameters by applying our error model during DNN retraining, to recover the accuracy lost due to approximations. Experimental results demonstrate the efficiency of our proposed method for accelerated retraining (11 x faster for CIFAR10 and 8x faster for ImageNet) for full DNN approximation, which allows us to deploy approximate multipliers with energy savings of up to 36% for 8-bit precision DNNs with an accuracy loss lower than 1%.
Cecilia De la Parra, Andre Guntoro, Akash Kumar 0001
ASP-DAC3
2021 CLAppED: A Design Framework for Implementing Cross-Layer Approximation in FPGA-based Embedded Systems
abstract
With the rising variation and complexity of embedded work-loads, FPGA-based systems are being increasingly used for many applications. The reconfigurability and high parallelism offered by FPGAs are used to enhance the overall performance of these applications. However, the resource constraints of embedded platforms can limit the performance in multiple ways. In recent years, Approximate Computing has emerged as a viable tool for improving the performance by utilizing reduced precision data structures and resource-optimized high-performance arithmetic operators. However, most of the related state-of-the-art research has mainly focused on utilizing approximate computing principles individually on different layers of the computing stack. Nonetheless, approximations across different layers of computing stack can substantially enhance the system’s performance. To this end, we present a framework to enable the intelligent exploration and highly accurate identification of the feasible design points in the large design space enabled by cross-layer approximations. Our framework proposes a novel polynomial regression-based method to model approximate arithmetic operators. The proposed method enables machine learning models to better correlate approximate operators with their impact on an application’s output quality. We use a 2D convolution operator as a test case and present the results for FPGA- based approximate hardware accelerators.
Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001
DAC3
2021 Resource-Efficient Database Query Processing on FPGAs
abstract
FPGA technology has introduced new ways to accelerate database query processing, that often result in higher performance and energy efficiency. This is thanks to the unique architecture of FPGAs using reconfigurable resources to behave like an application-specific integrated circuit upon programming. The limited amount of these resources restricts the number and type of modules that an FPGA can simultaneously support. In this paper, we propose "morphing sort-merge": a set of run-time configurable FPGA modules that achieves resource efficiency by reusing the FPGA's resources to support different pipeline-breaking database operators, namely sort, aggregation, and equi-join. The proposed modules use dynamic optimization mechanisms that adapt the implementation to the distribution of data at run-time, thus resulting in higher performance. Our benchmarks show that morphing sort-merge reaches an average speedup of 5x compared to MonetDB.
Mehdi Moghaddamfar, Christian Färber, Wolfgang Lehner, Norman May, Akash Kumar 0001
DaMoN5
2021 Knowledge Distillation and Gradient Estimation for Active Error Compensation in Approximate Neural Networks
abstract
Approximate computing is a promising approach for optimizing computational resources of error-resilient applications such as Convolutional Neural Networks (CNNs). However, such approximations introduce an error that needs to be compensated by optimization methods, which typically include a retraining or fine-tuning stage. To efficiently recover from the introduced error, this fine-tuning process needs to be adapted to take CNN approximations into consideration. In this work, we present a novel methodology for fine-tuning approximate CNNs with ultralow bit-width quantization and large approximation error, which combines knowledge distillation and gradient estimation to recover the lost accuracy due to approximations. With our proposed methodology, we demonstrate energy savings of up to 38% in complex approximate CNNs with weights quantized to 4 bits and 8-bit activations, with less than 3% accuracy loss w.r.t. the full precision model.
Cecilia De la Parra, Xuyi Wu, Andre Guntoro, Akash Kumar 0001
DATE4
2021 Nano Security: From Nano-Electronics to Secure Systems
abstract
The field of computer hardware stands at the verge of a revolution driven by recent breakthroughs in emerging nanodevices. “Nano Security” is a new Priority Program recently approved by DFG, the German Research Council. This initial-stage project initiative at the crossroads of nano-electronics and hardware-oriented security includes 11 projects with a total of 23 Principal Investigators from 18 German institutions. It considers the interplay between security and nano-electronics, focusing on a dichotomy which emerging nano-devices (and their architectural implications) have on system security. The projects within the Priority Program consider both: potential security threats and vulnerabilities stemming from novel nano-electronics, and innovative approaches to establishing and improving system security based on nano-electronics. This paper provides an overview of the Priority Program's overall philosophy and discusses the scientific objectives of its individual projects.
Ilia Polian, Frank Altmann, Tolga Arul, Christian Boit, Ralf Brederlow, Lucas Davi, Rolf Drechsler, Nan Du 0004, Thomas Eisenbarth 0001, Tim Güneysu, Sascha Hermann, Matthias Hiller, Rainer Leupers, Farhad Merchant, Thomas Mussenbrock, Stefan Katzenbeisser 0001, Akash Kumar 0001, Wolfgang Kunz, Thomas Mikolajick, Vivek Pachauri, Jean-Pierre Seifert, Frank Sill, Jens Trommer
DATE17
2021 Vertical IP Protection of the Next-Generation Devices: Quo Vadis?
abstract
With the advent of 5G and IoT applications, there is a greater thrust in terms of hardware security due to imminent risks caused by high amount of intercommunication between various subsystems. Security gaps in integrated circuits, thus represent high risks for both-the manufacturers and the users of electronic systems. Particularly in the domain of Intellectual Property (IP) protection, there is an urgent need to devise security measures at all levels of abstraction so that we can be one step ahead of any kind of adversarial attacks. This work presents IP protection measures from multiple perspectives-from system-level down to device-level security measures, from discussing various attack methods such as reverse engineering and hardware Trojan insertions to proposing new-age protection measures such as multi-valued logic locking and secure information flow tracking. This special session will give a holistic overview at the current state-of-the-art measures and how well we are prepared for the next generation circuits and systems.
Shubham Rai, Siddharth Garg, Christian Pilato, Vladimir Herdt, Elmira Moussavi, Dominik Germek, Ramesh Karri, Rolf Drechsler, Farhad Merchant, Akash Kumar 0001
DATE10
2021 Perspectives on Emerging Computation-in-Memory Paradigms
abstract
The traditional Von-Neumann architecture is reaching its limits and finding it difficult to cope up with the ever-increasing demands of modern workloads like artificial intelligence. This demand has fueled the search of technologies that can mimic human brain to efficiently combine both memory and computation within a single device. In this work, we present the state-of-the-art research in the domain of computation-in-memory. In particular, we take a look at memristors and its widespread application in neuromorphic computation. We introduce ReRAMs in terms of their novel computing paradigms and present ReRAM-specific design flows. We address the various circuit opportunities and challenges related to reliability and fault tolerance associated with them. Another high-potential candidate to leverage memory and computation from a single device is Ferroelectric Field-effect Transistor (FeFET). Here we present a co-integration of such FeFETs with another emerging nanotechnology concept, called Reconfigurable Field Effect Transistor (RFET) and discuss the impact of the higher amount of states provided by this combination.
Shubham Rai, Anteneh Gebregiorgis, Debjyoti Bhattacharjee, Krishnendu Chakrabarty, Said Hamdioui, Anupam Chattopadhyay, Jens Trommer, Akash Kumar 0001
DATE9
2021 Logic Synthesis Meets Machine Learning: Trading Exactness for Generalization
abstract
Logic synthesis is a fundamental step in hardware design whose goal is to find structural representations of Boolean functions while minimizing delay and area. If the function is completely-specified, the implementation accurately represents the function. If the function is incompletely-specified, the implementation has to be true only on the care set. While most of the algorithms in logic synthesis rely on SAT and Boolean methods to exactly implement the care set, we investigate learning in logic synthesis, attempting to trade exactness for generalization. This work is directly related to machine learning where the care set is the training set and the implementation is expected to generalize on a validation set. We present learning incompletely-specified functions based on the results of a competition conducted at IWLS 2020. The goal of the competition was to implement 100 functions given by a set of care minterms for training, while testing the implementation using a set of validation minterms sampled from the same function. We make this benchmark suite available and offer a detailed comparative analysis of the different approaches to learning.
Shubham Rai, Walter Lau Neto, Yukio Miyasaka, Xinpei Zhang, Mingfei Yu, Qingyang Yi, Masahiro Fujita 0004, Guilherme B. Manske, Matheus F. Pontes, Leomar S. da Rosa Jr., Marilton S. de Aguiar, Paulo F. Butzen, Po-Chun Chien, Yu-Shan Huang, Hoa-Ren Wang, Jie-Hong Roland Jiang, Jiaqi Gu 0002, Zheng Zhao 0003, Zixuan Jiang, David Z. Pan, Brunno Abreu, Isac de Souza Campos, Augusto Andre Souza Berndt, Cristina Meinhardt, Jônata Tyska Carvalho, Mateus Grellert, Sergio Bampi, Aditya Lohana, Akash Kumar 0001, Wei Zeng 0015, Azadeh Davoodi, Rasit Onur Topaloglu, Jordan Dotzel, Yichi Zhang 0006, Hanyu Wang 0005, Zhiru Zhang, Valerio Tenace, Pierre-Emmanuel Gaillardon, Alan Mishchenko, Satrajit Chatterjee
DATE29
2021 Preserving Self-Duality During Logic Synthesis for Emerging Reconfigurable Nanotechnologies
abstract
Emerging reconfigurable nanotechnologies allow the implementation of self-dual functions with a fewer number of transistors as compared to traditional CMOS technologies. To achieve better area results for Reconfigurable Field-Effect Transistors (RFET)-based circuits, a large portion of a logic representation must be mapped to self-dual logic gates. This, in turn, depends upon how self-duality is preserved in the logic representation during logic optimization and technology mapping. In the present work, we develop Boolean size-optimization methods-a rewriting and a resubstitution algorithm using Xor-Majority Graphs (XMGs) as a logic representation aiming at better preserving self-duality during logic optimization. XMGs are more compact for both unate and binate logic functions as compared to conventional logic representations such as And-Inverter Graphs (AIGs) or Majority-Inverter Graphs (MIGs). We evaluate the proposed algorithm over crafted benchmarks (with various levels of self-duality) and cryptographic benchmarks. For cryptographic benchmarks with a high self-duality ratio, the XMG-based logic optimisation flow can achieve an area reduction of up to 17% when compared to AIG-based optimization flows implemented in the academic logic synthesis tool ABC.
Shubham Rai, Heinz Riener, Giovanni De Micheli, Akash Kumar 0001
DATE4
2021 Improving the Timing Behaviour of Mixed-Criticality Systems Using Chebyshev's Theorem
abstract
In Mixed-Criticality (MC) systems, there are often multiple Worst-Case Execution Times (WCETs) for the same task, corresponding to system operation mode. Determining the appropriate WCETs for lower criticality modes is non-trivial; while on the one hand, a low WCET for a mode can improve the processor utilization in that mode, on the other hand, using a larger WCET ensures that the mode switches are minimized, thereby maximizing the quality-of-service for all tasks, albeit at the cost of processor utilization. Although there are many studies to determine WCET in the highest criticality mode, no analytical solutions are proposed to determine WCETs in other lower criticality modes. In this regard, we propose a scheme to determine WCETs by Chebyshev theorem to make a trade-off between the number of scheduled tasks at design-time and the number of dropped low-criticality tasks at runtime as a result of frequent mode switches. Our experimental results show that our scheme improves the utilization of state-of-the-art MC systems by up to 85.29%, while maintaining 9.11% mode switching probability in the worst-case scenario.
Behnaz Ranjbar, Ali Hoseinghorban, Siva Satyendra Sahoo, Alireza Ejlali, Akash Kumar 0001
DATE5
2021 NMPO: Near-Memory Computing Profiling and Offloading
abstract
Real-world applications are now processing big-data sets, often bottlenecked by the data movement between the compute units and the main memory. Near-memory computing (NMC), a modern data-centric computational paradigm, can alleviate these bottlenecks, thereby improving the performance of applications. The lack of NMC system availability makes simulators the primary evaluation tool for performance estimation. However, simulators are usually time-consuming, and methods that can reduce this overhead would accelerate the early-stage design process of NMC systems. This work proposes Near-Memory computing Profiling and Offloading (NMPO), a high-level framework capable of predicting NMC offloading suitability employing an ensemble machine learning model. NMPO predicts NMC suitability with an accuracy of 85.6% and, compared to prior works, can reduce the prediction time by using hardware-dependent applications features by up to 3 order of magnitude.
Stefano Corda, Madhurya Kumaraswamy, Ahsan Javed Awan, Roel Jordans, Akash Kumar 0001, Henk Corporaal
DSD5
2021 AMAH-Flex: A Modular and Highly Flexible Tool for Generating Relocatable Systems on FPGAs
abstract
In this work, we present a solution to a common problem encountered when using FPGAs in dynamic, ever-changing environments. Even when using dynamic function exchange to accommodate changing workloads, partial bitstreams are typically not relocatable. So the runtime environment needs to store all reconfigurable partition/reconfigurable module combinations as separate bitstreams. We present a modular and highly flexible tool (AMAH-Flex) that converts any static and reconfigurable system into a 2 dimensional dynamically relocatable system. It also features a fully automated floorplanning phase, closing the automation gap between synthesis and bitstream relocation. It integrates with the Xilinx Vivado toolchain and supports both FPGA architectures, the 7-Series and the UltraScale+. In addition, AMAH-Flex can be ported to any Xilinx FPGA family, starting with the 7-Series. We demonstrate the functionality of our tool in several reconfiguration scenarios on four different FPGA families and show that AMAH-Flex saves up to 80% of partial bitstreams.
Najdet Charaf, Christoph Tietz, Michael Raitza, Akash Kumar 0001, Diana Göhringer
FPT4
2021 MemOReL: A Memory-oriented Optimization Approach to Reinforcement Learning on FPGA-based Embedded Systems
abstract
Reinforcement Learning (RL) represents the machine learning method that has come closest to showing human-like learning. While Deep RL is becoming increasingly popular for complex applications such as AI-based gaming, it has a high implementation cost in terms of both power and latency. Q-Learning, on the other hand, is a much simpler method that makes it more feasible for implementation on resource-constrained embedded systems for control and navigation. However, the optimal policy search in Q-Learning is a compute-intensive and inherently sequential process and a software-only implementation may not be able to satisfy the latency and throughput constraints of such applications. To this end, we propose a novel accelerator design with multiple design trade-offs for implementing Q-Learning on FPGA-based SoCs. Specifically, we analyze the various stages of the Epsilon-Greedy algorithm for RL and propose a novel microarchitecture that reduces the latency by optimizing the memory access during each iteration. Consequently, we present multiple designs that provide varying trade-offs between performance, power dissipation, and resource utilization of the accelerator. With the proposed approach, we report considerable improvement in throughput with lower resource utilization over state-of-the-art design implementations.
Siva Satyendra Sahoo, Akhil Raj Baranwal, Salim Ullah, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI4
2021 Exploring Physical Synthesis for Circuits based on Emerging Reconfigurable Nanotechnologies
abstract
Recently proposed ambipolar nanotechnologies allow the development of reconfigurable circuits with low area and power overheads as compared to the conventional CMOS technology. However, using a conventional physical synthesis flow for circuits that include gates based on reconfigurable FETs (RFETs) leads to sub-optimal results. This is due to the fact that the physical synthesis flow for circuits based on RFETs has to cater to the additional gate terminal per RFET transistors. In the present work, we explore three important verticals that lead to an optimized physical synthesis flow for RFET-based circuits with circuit-level reconfigurability: (1) designing optimized layouts of reconfigurable gates, (2) utilize special driver cells to drive the reconfigurable portions of a circuit, and (3) optimized placement of these reconfigurable parts in separate power domains. Experimental evaluations over EPFL benchmarks using our proposed approach show a reduction in chip area of up to 17.5% when compared to conventional flows.
Andreas Krinke, Shubham Rai, Akash Kumar 0001, Jens Lienig
ICCAD3
2021 RL-Guided Runtime-Constrained Heuristic Exploration for Logic Synthesis
abstract
Within logic synthesis, most optimization scripts are well-defined heuristics that generalize over a variety of Boolean circuits. These heuristic-based scripts comprise various optimization algorithms which are applied sequentially in a specific order over a logic graph representation of Boolean circuits (typically in the form of And Inverter Graphs (AIGs) or Majority Inverter Graphs (MIGs)). These heuristics, despite being well-defined generalizations, may not perform well over all kinds of circuits. In order to develop custom heuristics specific to a particular Boolean circuit that performs well, we propose a runtime-constrained reinforcement learning (RL) approach which is able to generate scripts to carry out logic synthesis flows. Within our approach, we incorporate a graph convolution network (GCN) in order to perform a holistic exploration of the search space. To carry out an extensive evaluation, we identify three different classes of environments consisting of different baseline optimization sequences. The experimental results reveal that our model outperforms the prevalent state-of-the-art work [24] and the best heuristic-based scripts of Berkeley-ABC [4]. Our evaluations show that our framework provides up to an average of 8.3 % further reduction in level over the EPFL Benchmark Suite [8] as compared to the Berkeley-ABC scripts. Further, we develop a framework for the EPFL mockturtle [20] logic synthesis libraries and generate custom scripts using our RL-based approach.
Yasasvi V. Peruvemba, Shubham Rai, Kapil Ahuja, Akash Kumar 0001
ICCAD4
2021 dCSR: A Memory-Efficient Sparse Matrix Representation for Parallel Neural Network Inference
abstract
Reducing the memory footprint of neural networks is a crucial prerequisite for deploying them in small and low-cost embedded devices. Network parameters can often be reduced significantly through pruning. We discuss how to best represent the indexing overhead of sparse networks for the coming generation of Single Instruction, Multiple Data (SIMD)-capable microcontrollers. From this, we develop Delta-Compressed Storage Row (dCSR), a storage format for sparse matrices that allows for both low overhead storage and fast inference on embedded systems with wide SIMD units. We demonstrate our method on an ARM Cortex-M55 MCU prototype with M-Profile Vector Extension (MVE). A comparison of memory consumption and throughput shows that our method achieves competitive compression ratios and increases throughput over dense methods by up to$2.9\times$for sparse matrix-vector multiplication (SpMV)-based kernels and$1.06\times$for sparse matrix-matrix multiplication (SpMM). This is accomplished through handling the generation of index information directly in the SIMD unit, leading to an increase in effective memory bandwidth.
Elias Trommer, Bernd Waschneck, Akash Kumar 0001
ICCAD3
2021 BioCare: An Energy-Efficient CGRA for Bio-Signal Processing at the Edge
abstract
Coarse Grained Reconfigurable Architectures (CGRAs) have proved to be viable platforms for health monitoring applications. Targeting energy-efficiency, state-of-the-art (SoA) CGRAs are augmented with approximation techniques, while still maintain acceptable accuracy at final Quality of Result (QoR). However, such CGRAs suffer from overheads of collecting separate Add/Mul/Div units. We propose BioCare as an area- and energy-efficient CGRA for health-monitoring edge devices, which exploits the synergistic effects of multiple approximations across HW/SW stack. BioCare offers different levels of energy-accuracy trade-off through the plasticity of its small PEs, each can support precision-adaptability with a Single Instruction, Multiple Data (SIMD) manner. BioCare demonstrates its superiority over SoAs, by achieving up to 32% and 67% area- and energy-savings, with 3.6 χ higher throughput. In addition to analysis on multiple kernels, evaluations on a multi-kernel ECG application shows that BioCare speed-ups the QRS detection latency by 61%, with 0% loss in accuracy. Our implementations will be available at https://cfaed.tu-dresden.de/pd-downloads.
Zahra Ebrahimi, Akash Kumar 0001
ISCAS2
2021 Exploiting Resiliency for Kernel-Wise CNN Approximation Enabled by Adaptive Hardware Design
abstract
Efficient low-power accelerators for Convolutional Neural Networks (CNNs) largely benefit from quantization and approximation, which are typically applied layer-wise for efficient hardware implementation. In this work, we present a novel strategy for efficient combination of these concepts at a deeper level, which is at each channel or kernel. We first apply layer-wise, low bit-width, linear quantization and truncation-based approximate multipliers to the CNN computation. Then, based on a state-of-the-art resiliency analysis, we are able to apply a kernel-wise approximation and quantization scheme with negligible accuracy losses, without further retraining. Our proposed strategy is implemented in a specialized framework for fast design space exploration. This optimization leads to a boost in estimated power savings of up to 34% in residual CNN architectures for image classification, compared to the base quantized architecture.
Cecilia De la Parra, Ahmed El-Yamany, Taha Soliman, Akash Kumar 0001, Norbert Wehn, Andre Guntoro
ISCAS4
2021 Metastability with Emerging Reconfigurable Transistors: Exploiting Ambipolarity for Throughput
abstract
In this work, we leverage ambipolar transistors in the context of metastability for random number generation. We propose designs of a Minority-based SR latch and a dual-edge triggered True Single Phase Clock D-Flip-Flop (TSPC DFF) to sample two random bits in a single clock cycle. We demonstrate how metastable circuits based on ambipolar transistors allow doubling the throughput as compared to a similar standard CMOS-based design. The proposed design is compact in terms of the number of transistors per block (60% less transistors), power consumption (saving 94.5% leakage power and 70.7% dynamic power) and path delay (77.3% reduction) with respect to its CMOS counterpart.
Abhiroop Bhattacharjee, Shubham Rai, Ansh Rupani, Michael Raitza, Akash Kumar 0001
VLSI-SoC5
2021 CLEO-CoDe: Exploiting Constrained Decoding for Cross-Layer Energy Optimization in Heterogeneous Embedded Systems
abstract
System-level design for low-power and energy efficiency in embedded systems using Heterogeneous Multi-Processor System-on-Chip (HMPSoC) is a challenging task due to the large design space. The related Design Space Exploration (DSE) suffers from scaling due to various degrees of freedom across multiple layers of the compute stack. Traditional multi-objective metaheuristic approaches work well for unconstrained system-level design, but do not scale well with the additional system and user constraints. Using a SATisfiablity problem (SAT) solver as a decoder for the meta-heuristics has been explored in related research as a better solution for a discrete constrained multi-objective optimization problem. This approach restricts the problem to the feasible space, hence improving the quality of the results. In this paper, we explore the ways in which constrained decoding such as the SAT decoding approach can be leveraged for cross-layer design space exploration. Low-power methodologies such as Dynamic Voltage and Frequncy Scaling (DVFS) and application-specific implementations are integrated, thus scaling the design space. Additionally, we demonstrate how user constraints on the system synthesis problem can be learned by our proposed approach to prune the meta-heuristic design space and improve the quality of the solutions. As the constraints on the problem increase and the design space scales, the constrained decoding approach outperforms a typical meta-heuristic approach.
Siva Satyendra Sahoo, Akash Kumar 0001
VLSI-SoC2
2021 Using Monte Carlo Tree Search for EDA - A Case-study with Designing Cross-layer Reliability for Heterogeneous Embedded Systems
abstract
Continued transistor scaling and increasing power density have led to considerable increase in fault-rates in silicon nanotechnology-based real-time systems. Cross-layer fault tolerance techniques present a more cost-efficient methodology for adapting to such increased fault rates by distributing fault-tolerance to different layers. To this end, we propose a methodology for integrating the design space exploration (DSE) for taskmapping on heterogeneous hardware-platforms with designing cross-layer reliability. Specifically, we model the DSE for task-mapping with cross-layer reliability as a tree search problem and use Monte Carlo Tree Search for task-mapping and scheduling applications with specific reliability requirements. The proposed methodology results in considerable improvements over a standalone approach to task-mapping and implementing cross-layer reliability.
Siva Satyendra Sahoo, Akash Kumar 0001
VLSI-SoC2
2021 Area-Optimized Accurate and Approximate Softcore Signed Multiplier Architectures
abstract
Multiplication is one of the most extensively used arithmetic operations in a wide range of applications. In order to provide resource-efficient and high-performance multipliers, previous works have proposed different designs of accurate and approximate multipliers-mainly for ASIC-based systems. However, the architectural differences between ASICs- and FPGA-based systems limit the effectiveness of these multipliers for FPGA-based systems. Moreover, most of these multiplier designs are valid only for unsigned numbers. To bridge this gap, we propose a novel implementation technique for designing resource-efficient and low-power accurate and approximate signed multipliers which are optimized for FPGA-based systems. Compared to Vivado's area-optimized multiplier IPs, the designs obtained using our proposed technique occupy 47 to 63 percent less area (Lookup Tables). To accelerate further research in this direction and reproduce the presented results, the RTL and behavioral models of our proposed methodology are available as an open-source library.11.Online. [Available]: https://cfaed.tu-dresden.de/pd-downloads.
Salim Ullah, Hendrik Schmidl, Siva Satyendra Sahoo, Semeen Rehman, Akash Kumar 0001
IEEE Trans. Computers5
2021 ReLAccS: A Multilevel Approach to Accelerator Design for Reinforcement Learning on FPGA-Based Systems
abstract
Reinforcement learning (RL), specifically Q-learning, with human-like learning abilities to learn from experience without any a priori data, is being increasingly used in embedded systems in the field of control and navigation. However, finding the optimal policy in this approach can be highly compute-intensive, and a software-only implementation may not satisfy the application's timing constraints. To this end, we propose optimization methods at multiple levels of accelerator design for RL. Specifically, at the architecture-level, we exploit the instruction-level parallelism and the spatial parallelism in FPGAs to improve the throughput over state-of-the-art designs by up to 34%. Further, we propose lookup table-level optimizations to reduce the resource utilization and power dissipation of the accelerator. Finally, we propose algorithm-level approximation that can be used for acceleration of Q-learning problems with more states and for reducing the peak power dissipation. We report up to 10× reduction in power dissipation with marginal degradation in quality of results.
Akhil Raj Baranwal, Salim Ullah, Siva Satyendra Sahoo, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 Power-Aware Runtime Scheduler for Mixed-Criticality Systems on Multicore Platform
abstract
In modern multicore mixed-criticality (MC) systems, a rise in peak power consumption due to parallel execution of tasks with maximum frequency, specially in the overload situation, may lead to thermal issues, which may affect the reliability and timeliness of MC systems. Therefore, managing peak power consumption has become imperative in multicore MC systems. In this regard, we propose an online peak power and thermal management heuristic for multicore MC systems. This heuristic reduces the peak power consumption of the system as much as possible during runtime by exploiting dynamic slack and per-cluster dynamic voltage and frequency scaling (DVFS). Specifically, our approach examines multiple tasks ahead to determine the most appropriate one for slack assignment, that has the most impact on the system peak power and temperature. However, changing the frequency and selecting a proper task for slack assignment and a proper core for task remapping at runtime can be time-consuming and may cause deadline violation which is not admissible for high-criticality tasks. Therefore, we analyze and then optimize our runtime scheduler and evaluate it for various platforms. The proposed approach is experimentally validated on the ODROID-XU3 (DVFS-enabled heterogeneous multicore platform) with various embedded real-time benchmarks. Results show that our heuristic achieves up to 5.25% reduction in system peak power and 20.33% reduction in maximum temperature compared to an existing method while meeting deadline constraints in different criticality modes.
Behnaz Ranjbar, Tuan D. A. Nguyen, Alireza Ejlali, Akash Kumar 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 CGRA-EAM - Rapid Energy and Area Estimation for Coarse-grained Reconfigurable Architectures
abstract
Reconfigurable architectures are quickly gaining in popularity due to their flexibility and ability to provide high energy efficiency. However, reconfigurable systems allow for a huge design space. Iterative design space exploration (DSE) is often required to achieve good Pareto points with respect to some combination of performance, area, and/or energy. DSE tools depend on information about hardware characteristics in these aspects. These characteristics can be obtained from hardware synthesis and net-list simulation, but this is very time-consuming. Therefore, architecture models are common. This work introduces CGRA-EAM (Coarse-Grained Reconfigurable Architecture - Energy & Area Model), a model for energy and area estimation framework for coarse-grained reconfigurable architectures. The model is evaluated for the Blocks CGRA. The results demonstrate that the mean absolute percentage error is 15.5% and 2.1% for energy and area, respectively, while the model achieves a speedup of close to three orders of magnitude compared to synthesis.
Mark Wijtvliet, Henk Corporaal, Akash Kumar 0001
ACM Trans. Reconfigurable Technol. Syst.3
2020 LeAp: Leading-one Detection-based Softcore Approximate Multipliers with Tunable Accuracy
abstract
Approximate multipliers are ubiquitously used in diverse applications by exploiting circuit simplification, mainly specialized for Application-Specific Integrated Circuit (ASIC) platforms. However, the intrinsic architectural specifications of Field-Programmable Gate Arrays (FPGAs) prohibited comparable resource gains when directly applying these techniques. LeAp is an area-, throughput-, and energy-efficient approximate multiplier for FPGAs which efficiently utilizes 6-input Look-up Tables (6-LUTs) and fast carry chains in its novel approximate log calculator to implement Mitchell's algorithm. Moreover, three novel error-refinement schemes with negligible area overhead and independent from multiplier-size, have boosted accuracy to>99%. Experimental results obtained from Vivado, Artificial Neural Network (ANN) and image processing applications indicate superiority of proposed multiplier over accurate and state-of-the-art approximate counterparts. In particular, LeAp outperforms the 32x32 accurate multiplier by achieving 69.7%, 14.7%, 42.1%, and 37.1% improvement in area, throughput, power, and energy, respectively. The library of RTL and behavioral implementations will be open-sourced at https://cfaed.tu-dresden.de/pd-downloads.
Zahra Ebrahimi, Salim Ullah, Akash Kumar 0001
ASP-DAC3
2020 CL(R)Early: An Early-stage DSE Methodology for Cross-Layer Reliability-aware Heterogeneous Embedded Systems
abstract
Cross-layer reliability (CLR) presents a cost-effective alternative to traditional single-layer design in resource-constrained embedded systems. CLR provides the scope for leveraging the inherent fault-masking of multiple layers and exploiting application-specific tolerances to degradation in some Quality of Service (QoS) metrics. However, it can also lead to an explosion in the design complexity. State-of-the art approaches to such joint optimization across multiple degrees of freedom can lead to degradation in the system-level Design Space Exploration (DSE) results. To this end, we propose a DSE methodology for enabling CLR-aware task-mapping in heterogeneous embedded systems. Specifically, we present novel approaches to both task and system-level analysis for performing an early-stage exploration of various design decisions. The proposed methodology results in considerable improvements over other state-of-the-art approaches and shows significant scaling with application size.
Siva Satyendra Sahoo, Bharadwaj Veeravalli, Akash Kumar 0001
DAC3
2020 ProxSim: GPU-based Simulation Framework for Cross-Layer Approximate DNN Optimization
abstract
Through cross-layer approximation of Deep Neural Networks (DNN) significant improvements in hardware resources utilization for DNN applications can be achieved. This comes at the cost of accuracy degradation, which can be compensated through different optimization methods. However, DNN optimization is highly time-consuming in existing simulation frameworks for cross-layer DNN approximation, as they are usually implemented for CPU usage only. Specially for large-scale image processing tasks, the need of a more efficient simulation framework is evident. In this paper we present ProxSim, a specialized, GPU-accelerated simulation framework for approximate hardware, based on Tensorflow, which supports approximate DNN inference and retraining. Additionally, we propose a novel hardware-aware regularization technique for approximate DNN optimization. By using ProxSim, we report up to 11× savings in execution time, compared to a multi-thread CPU-based framework, and an accuracy recovery of up to 30% for three case studies of image classification with MNIST, CIFAR-10 and ImageNet.
Cecilia De la Parra, Andre Guntoro, Akash Kumar 0001
DATE3
2020 DiSCERN: Distilling Standard-Cells for Emerging Reconfigurable Nanotechnologies
abstract
Recent attempts on circuits based on emerging reconfigurable nanotechnologies have primarily focused on using the traditional CMOS design flow involving similar-styled standard-cells. In the present work, we show that logic gates which implement self-dual functions can be efficiently implemented using reconfigurable nanotechnologies. We propose an algorithm which analyses the truth-tables of cuts in a mapped circuit to list all such potential reconfigurable logic gates for a particular circuit. Technology mapping with these new logic gates (or standard-cells) leads to a better mapping in terms of area and delay. Experiments employing our methodology over EPFL benchmarks, show average improvements of around 13%, 16% and 11.5% in terms of area, number of edges and delay respectively as compared to the conventional CMOS-centric standard-cell based mapping.
Shubham Rai, Michael Raitza, Siva Satyendra Sahoo, Akash Kumar 0001
DATE4
2020 L2L: A Highly Accurate Log_2_Lead Quantization of Pre-trained Neural Networks
abstract
Deep Neural Networks are one of the machine learning techniques which are increasingly used in a variety of applications. However, the significantly high memory and computation demands of deep neural networks often limit their deployment on embedded systems. Many recent works have considered this problem by proposing different types of data quantization schemes. However, most of these techniques either require post-quantization retraining of deep neural networks or bear a significant loss in output accuracy. In this paper, we propose a novel quantization technique for parameters of pre-trained deep neural networks. Our technique significantly maintains the accuracy of the parameters and does not require retraining of the networks. Compared to the single-precision floating-point numbers-based implementation, our proposed 8-bit quantization technique generates only ~1% and the ~0.4%, loss in top-1 and top-5 accuracies respectively for VGG16 network using ImageNet dataset.
Salim Ullah, Siddharth Gupta 0004, Kapil Ahuja, Aruna Tiwari, Akash Kumar 0001
DATE5
2020 Introducing FPGA-based Machine Learning on the Edge to Undergraduate Students
abstract
This innovative practice category work in progress paper describes a project in a final-year un-dergraduate course on implementing a neural accelerator on an FPGA for edge computing. In our university, an undergraduate course on Embedded Hardware System Design introduces students to advanced hardware design techniques with the goal of integrating the created hardware into a complete system. Students learn concepts such as high-level synthesis (HLS), logic synthesis and physical design, with an emphasis on FPGA-based designs. They also understand bus systems such as Advanced eXtensible Interface (AXI) to interconnect the various components. The concepts are put into practice through a project. A series of labs provide scaffolding to students through the course of implementing the project. These labs take students systematically through an introduction to hardware-software co-design, hardware design, creation and interfacing of custom co-processors and HLS. Important hardware design and optimization concepts, as well as managing the data interaction between hardware and software were reinforced through the project. The project also provided many students with the opportunity to be introduced to neural networks and machine learning (ML). Quantitative and qualitative results from a survey indicate that students gained a lot of knowledge and experience through the course of the project. The current form of the project streams in data from a local computer to which the FPGA is connected. Future work includes true and direct cloud connectivity and improved use cases for making it a true Internet of Things (IoT) project.
Rajesh C. Panicker, Akash Kumar 0001, Chacko John Deepu
FIE2
2020 Maximizing the Serviceability of Partially Reconfigurable FPGA Systems in Multi-tenant Environment
abstract
In cloud computing, software is transitioning from monolithic to microservices architecture to improve the maintainability, upgradability and the flexibility of the applications. They are able to request a service with different implementations of the same functionality, including hardware accelerator, depending on cost and performance. This model opens up a new opportunity to integrate reconfigurable hardware, specifically, FPGA, in the cloud to offer such services. There are many research works discussing solutions for this problem but they focus primarily on the high-level aspects of resource manager, hypervisor or hardware architecture. The low-level physical design choices of FPGA to maximize the accelerator allocation success rate (called serviceability) is largely untouched. In this paper, we propose a design space exploration algorithm to determine the best configuration of partially reconfigurable regions (PRRs) to host the accelerators. Besides, the algorithm is capable of estimating the actual resources occupied by the PRRs on the FPGA even before floorplanning. We systematically study the effects of having more PRRs on the system in various aspects, i.e., serviceability, waiting time and resource wastage. The experiments show that at a certain number of PRRs, upto 91% serviceability can be achieved for 12 concurrent users. It is a significant improvement from 52% without our approach. The average amount of time that each request has to wait to be served is also reduced by 6.3X. Furthermore, the cumulative unused FPGA resources is reduced almost by half.
Tuan D. A. Nguyen, Akash Kumar 0001
FPGA2
2020 SIMDive: Approximate SIMD Soft Multiplier-Divider for FPGAs with Tunable Accuracy
abstract
The ever-increasing quest for data-level parallelism and variable precision in ubiquitous multimedia and Deep Neural Network (DNN) applications has motivated the use of Single Instruction, Multiple Data (SIMD) architectures. To alleviate energy as their main resource constraint, approximate computing has re-emerged, albeit mainly specialized for their Application-Specific Integrated Circuit (ASIC) implementations. This paper, presents for the first time, an SIMD architecture based on novel multiplier and divider with tunable accuracy, targeted for Field-Programmable Gate Arrays (FPGAs). The proposed hybrid architecture implements Mitchell's algorithms and supports precision variability from 8 to 32 bits. Experimental results obtained from Vivado, multimedia and DNN applications indicate superiority of proposed architecture (both in SISD and SIMD) over accurate and state-of-the-art approximate counterparts. In particular, the proposed SISD divider outperforms the accurate Intellectual Property (IP) divider provided by Xilinx with 4x higher speed and 4.6x less energy and tolerating only 0.8% error. Moreover, the proposed SIMD multiplier-divider supersede accurate SIMD multiplier by achieving up to 26%, 45%, 36%, and 56% improvement in area, throughput, power, and energy, respectively.
Zahra Ebrahimi, Salim Ullah, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI3
2020 Full Approximation of Deep Neural Networks through Efficient Optimization
abstract
Approximate Computing is a promising paradigm for mitigating computational requirements of Deep Neural Networks (DNN), by taking advantage of their inherent error resilience. Specifically, the use of approximate multipliers in DNN inference can lead to significant improvements in power consumption of embedded DNN applications. This paper presents a methodology for efficient approximate multiplier selection and for full and uniform approximation of large DNNs, through retraining and minimization of the approximation error. We evaluate our methodology using 422 approximate multipliers from the EvoApprox library, with three different Residual architectures trained with Cifar10, and achieve energy savings of up to 18% surpassing the original floating-point accuracy, and of up to 58% with an accuracy loss of 0.73%.
Cecilia De la Parra, Andre Guntoro, Akash Kumar 0001
ISCAS3
2020 Improving approximate neural networks for perception tasks through specialized optimization
Cecilia De la Parra, Andre Guntoro, Akash Kumar 0001
Future Gener. Comput. Syst.3
2019 A Hybrid Agent-based Design Methodology for Dynamic Cross-layer Reliability in Heterogeneous Embedded Systems
abstract
Technology scaling and architectural innovations have led to increasing ubiquity of embedded systems across applications with widely varying and often constantly changing performance and reliability specifications. However, the increasing physical fault-rates in electronic systems have led to single-layer reliability approaches becoming infeasible for resource-constrained systems. Dynamic Cross-layer reliability (CLR) provides scope for efficient adaptation to such QoS variations and increasing unreliability. We propose a design methodology for enabling QoS-aware CLR-integrated runtime adaptation in heterogeneous MPSoC-based embedded systems. Specifically, we propose a combination of reconfiguration cost-aware optimization at design-time and an agent-based optimization at run-time. We report a reduction of up to 51% and 37% in average reconfiguration cost and average energy consumption respectively over state-of-the-art approaches.
Siva Satyendra Sahoo, Bharadwaj Veeravalli, Akash Kumar 0001
DAC3
2019 High-Throughput BitPacking Compression
abstract
To efficiently support analytical applications from a data management perspective, in-memory column store database systems are state-of-the art. In this kind of database system, lossless lightweight integer compression schemes are crucial to keep the memory storage as low as possible and to speedup query processing. In this specific compression domain, BitPacking is one of the most frequently applied compression scheme. However, (de) compression should not come with any additional cost during run time, but should be provided transparently without compromising the overall system performance. To achieve that, we focus on acceleration of BitPacking using Field Programmable Gate Arrays (FPGAs). Therefore, we outline several FPGA designs for BitPacking in this paper. As we are going to show in our evaluation, our specific designs provide the BitPacking compression scheme with high-throughput.
Nusrat Jahan Lisa, Tuan D. A. Nguyen, Dirk Habich, Akash Kumar 0001, Wolfgang Lehner
DSD4
2019 Online Peak Power and Maximum Temperature Management in Multi-core Mixed-Criticality Embedded Systems
abstract
In this work, we address peak power and maximum temperature in multi-core Mixed-Criticality (MC) systems. In these systems, a rise in peak power consumption may generate more heat beyond the cooling capacity. Additionally, the reliability and timeliness of MC systems may be affected due to excessive temperature. Therefore, managing peak power consumption has become imperative in multi-core MC systems. In this regard, we propose an online peak power management heuristic for multi-core MC systems. This heuristic reduces the peak power consumption of the system as much as possible during runtime by exploiting dynamic slack and Dynamic Voltage and Frequency Scaling (DVFS). Specifically, our approach examines multiple tasks ahead to determine the most appropriate one for slack assignment instead of just one task as in the literature. The selection is based on the impact of the tasks on peak power and temperature of the system. The DVFS is then applied to that task to reduce the system peak power and maximum temperature. Further, a re-mapping technique is proposed to further improve the results. Our experimental results show that our heuristic achieves up to 18.2% reduction in system peak power consumption and 8.1% reduction in maximum temperature compared to an existing method. The inherent energy consumption is also reduced by up to 50%.
Behnaz Ranjbar, Tuan D. A. Nguyen, Alireza Ejlali, Akash Kumar 0001
DSD4
2019 Exploiting Emerging Reconfigurable Technologies for Secure Devices
abstract
In the present work, we show how new and emerging reconfigurable technologies provide promising improvement over CMOS in the field of hardware security and encryption. We demonstrate how security features are a natural outcome of the circuits based on Silicon Nanowire reconfigurable transistors. This forms the basis of authentication key based security technique. Using the authentication key based system, we obtained the maximum possible key-length for MCNC benchmark circuits. Further, we formulated security as a tunable aspect for a circuit, by introducing don't care adjustment. A combination of the above two is used to establish security in terms of Shannon's entropy. We show that using the above concepts, Shannon's entropy increases for 99.1% benchmarks out of which maximum entropy is reached for 38.5% of all the benchmarks. We demonstrate these concepts using a case study for a 2-bit Ripple Carry Adder (RCA) based on SiNW RFETs and compare the design with its CMOS counterpart.
Ansh Rupani, Shubham Rai, Akash Kumar 0001
DSD3
2019 Design Methodology for Embedded Approximate Artificial Neural Networks
abstract
Artificial neural networks (ANNs) have demonstrated significant promise while implementing recognition and classification applications. The implementation of pre-trained ANNs on embedded systems requires representation of data and design parameters in low-precision fixed-point formats; which often requires retraining of the network. For such implementations, the multiply-accumulate operation is the main reason for resultant high resource and energy requirements. To address these challenges, we present Rox-ANN, a design methodology for implementing ANNs using processing elements (PEs) designed with low-precision fixed-point numbers and high performance and reduced-area approximate multipliers on FPGAs. The trained design parameters of the ANN are analyzed and clustered to optimize the total number of approximate multipliers required in the design. With our methodology, we achieve insignificant loss in application accuracy. We evaluated the design using a LeNet based implementation of the MNIST digit recognition application. The results show a 65.6%, 55.1% and 18.9% reduction in area, energy consumption and latency for a PE using 8-bit precision weights and activations and approximate arithmetic units, when compared to 16-bit full precision, accurate arithmetic PEs.
Adarsha Balaji, Salim Ullah, Anup Das 0001, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI4
2019 Shouji: a fast and efficient pre-alignment filter for sequence alignment
abstract
MOTIVATION: The ability to generate massive amounts of sequencing data continues to overwhelm the processing capability of existing algorithms and compute infrastructures. In this work, we explore the use of hardware/software co-design and hardware acceleration to significantly reduce the execution time of short sequence alignment, a crucial step in analyzing sequenced genomes. We introduce Shouji, a highly parallel and accurate pre-alignment filter that remarkably reduces the need for computationally-costly dynamic programming algorithms. The first key idea of our proposed pre-alignment filter is to provide high filtering accuracy by correctly detecting all common subsequences shared between two given sequences. The second key idea is to design a hardware accelerator that adopts modern field-programmable gate array (FPGA) architectures to further boost the performance of our algorithm. RESULTS: Shouji significantly improves the accuracy of pre-alignment filtering by up to two orders of magnitude compared to the state-of-the-art pre-alignment filters, GateKeeper and SHD. Our FPGA-based accelerator is up to three orders of magnitude faster than the equivalent CPU implementation of Shouji. Using a single FPGA chip, we benchmark the benefits of integrating Shouji with five state-of-the-art sequence aligners, designed for different computing platforms. The addition of Shouji as a pre-alignment step reduces the execution time of the five state-of-the-art sequence aligners by up to 18.8×. Shouji can be adapted for any bioinformatics pipeline that performs sequence alignment for verification. Unlike most existing methods that aim to accelerate sequence alignment, Shouji does not sacrifice any of the aligner capabilities, as it does not modify or replace the alignment step. AVAILABILITY AND IMPLEMENTATION: https://github.com/CMU-SAFARI/Shouji. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mohammed Alser, Hasan Hassan, Akash Kumar 0001, Onur Mutlu, Can Alkan
Bioinform.3
2019 Multi-objective design space exploration for system partitioning of FPGA-based Dynamic Partially Reconfigurable Systems
Siva Satyendra Sahoo, Tuan D. A. Nguyen, Bharadwaj Veeravalli, Akash Kumar 0001
Integr.4
2019 Architecture and Advanced Electronics Pathways Toward Highly Adaptive Energy- Efficient Computing
abstract
With the explosion of the number of compute nodes, the bottleneck of future computing systems lies in the network architecture connecting the nodes. Addressing the bottleneck requires replacing current backplane-based network topologies. We propose to revolutionize computing electronics by realizing embedded optical waveguides for onboard networking and wireless chip-to-chip links at 200-GHz carrier frequency connecting neighboring boards in a rack. The control of novel rate-adaptive optical and mm-wave transceivers needs tight interlinking with the system software for runtime resource management.
Gerhard P. Fettweis, Meik Dörpinghaus, Jerónimo Castrillón, Akash Kumar 0001, Christel Baier, Karlheinz Bock, Frank Ellinger, Andreas Fery, Frank H. P. Fitzek, Hermann Härtig, Kambiz Jamshidi, Thomas Kissinger, Wolfgang Lehner, Michael Mertig, Wolfgang E. Nagel, Giang T. Nguyen 0002, Dirk Plettemeier, Michael Schröter, Thorsten Strufe
Proc. IEEE4
2019 Designing Efficient Circuits Based on Runtime-Reconfigurable Field-Effect Transistors
abstract
An early evaluation in terms of circuit design is essential in order to assess the feasibility and practicability aspects for emerging nanotechnologies. Reconfigurable nanotechnologies, such as silicon or germanium nanowire-based reconfigurable field-effect transistors, hold great promise as suitable primitives for enabling multiple functionalities per computational unit. However, contemporary CMOS circuit designs when applied directly with this emerging nanotechnology often result in suboptimal designs. For example, 31% and 71% larger area was obtained for our two exemplary designs. Hence, new approaches delivering tailored circuit designs are needed to truly tap the exciting feature set of these reconfigurable nanotechnologies. To this effect, we propose six functionally enhanced logic gates based on a reconfigurable nanowire technology and employ these logic gates in efficient circuit designs. We carry out a detailed comparative study for a reconfigurable multifunctional circuit, which shows better normalized circuit delay (20.14%), area (32.40%), and activity as the power metric (40%) while exhibiting similar functionality as compared with the CMOS reference design. We further propose a novel design for a 1-bit arithmetic logic unit-based on silicon nanowire reconfigurable FETs with the area, normalized circuit delay, and activity gains of 30%, 34%, and 36%, respectively, as compared with the contemporary CMOS version.
Shubham Rai, Jens Trommer, Michael Raitza, Thomas Mikolajick, Walter M. Weber, Akash Kumar 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2018 Lifetime-aware design methodology for dynamic partially reconfigurable systems
abstract
Dynamic Partial Reconfiguration (DPR) in reconfigurable platforms can be used for the mitigation of aging-related permanent faults. We propose an application-specific system-level design methodology for determining the appropriate number of Partially Reconfigurable Regions and their compatibility with Partially Reconfigurable Modules for maximizing the system lifetime. Specifically, we propose a lifetime-aware scheduler that maximizes system MTTF. We use the scheduler along with an automated floorplanner for design space exploration at design-time to generate a heterogeneous PRR system. Our experiments show that the heterogeneous systems can offer up to 2x lifetime improvement over homogeneous ones.
Siva Satyendra Sahoo, Tuan D. A. Nguyen, Bharadwaj Veeravalli, Akash Kumar 0001
ASP-DAC4
2018 SMApproxlib: library of FPGA-based approximate multipliers
abstract
The main focus of the existing approximate arithmetic circuits has been on ASIC-based designs. However, due to the architectural differences between ASICs and FPGAs, comparable performance gains cannot be achieved for FPGA-based systems by using the approximations defined, particularly for ASIC-based systems. This paper exploits the structure of the 6-input lookup tables and associated carry chains of modern FPGAs to define a methodology for designing approximate multipliers optimized for FPGA-based systems. Using our presented methodology, we present SMApproxLib, an open source library of approximate multipliers with different bit-widths, output accuracies and performance gains. Being the first open source library of FPGA-based approximate multipliers, SMAp-proxLib can serve as a benchmark for designing and comparing future FPGA-based approximate arithmetic circuits.
Salim Ullah, Sanjeev Sripadraj Murthy, Akash Kumar 0001
DAC3
2018 Area-optimized low-latency approximate multipliers for FPGA-based hardware accelerators
abstract
The architectural differences between ASICs and FPGAs limit the effective performance gains achievable by the application of ASIC-based approximation principles for FPGA-based reconfigurable computing systems. This paper presents a novel approximate multiplier architecture customized towards the FPGA-based fabrics, an efficient design methodology, and an open-source library. Our designs provide higher area, latency and energy gains along with better output accuracy than those offered by the state-of-the-art ASIC-based approximate multipliers. Moreover, compared to the multiplier IP offered by the Xilinx Vivado, our proposed design achieves up to 30%, 53%, and 67% gains in terms of area, latency, and energy, respectively, while incurring an insignificant accuracy loss (on average, below 1% average relative error). Our library of approximate multipliers is open-source and available online at https://cfaed.tudresden.de/pd-downloads to fuel further research and development in this area, and thereby enabling a new research direction for the FPGA community.
Salim Ullah, Semeen Rehman, Bharath Srinivas Prabakaran, Florian Kriebel, Muhammad Abdullah Hanif, Muhammad Shafique 0001, Akash Kumar 0001
DAC7
2018 Column Scan Optimization by Increasing Intra-Instruction Parallelism
Nusrat Jahan Lisa, Annett Ungethüm, Dirk Habich, Tuan D. A. Nguyen, Akash Kumar 0001, Wolfgang Lehner
DATA5
2018 DeMAS: An efficient design methodology for building approximate adders for FPGA-based systems
abstract
The current state-of-the-art approximate adders are mostly ASIC-based, i.e., they focus solely on gate and/or transistor level approximations (e.g., through circuit simplification or truncation) to achieve area, latency, power and/or energy savings at the cost of accuracy loss. However, when these designs are synthesized for FPGA-based systems, they do not offer similar reductions in area, latency and power/energy due to the underlying architectural differences between ASICs and FPGAs. In this paper, we present a novel generic design methodology to synthesize and implement approximate adders for any FPGA-based system by considering the underlying resources and architectural differences. Using our methodology, we have designed, analyzed and presented eight different multi-bit adder architectures. Compared to the 16-bit accurate adder, our designs are successful in achieving area, latency and power-delay product gains of 50%, 38%, and 53%, respectively. We also compare our approximate adders to state-of-the-art approximate adders specialized for ASIC and FPGA fabrics and demonstrate the benefits of our approach. We will make the RTL and behavioral models of our and state-of-the-art designs open-source at https://sourceforge.net/projects/approxfpgas/ to further fuel the research and development in the FPGA community and to ensure reproducible research.
Bharath Srinivas Prabakaran, Semeen Rehman, Muhammad Abdullah Hanif, Salim Ullah, Ghazal Mazaheri, Akash Kumar 0001, Muhammad Shafique 0001
DATE6
2018 Technology mapping flow for emerging reconfigurable silicon nanowire transistors
abstract
Efficient circuit designs can make use of ambipolar nature of silicon nanowire (SiNW) over CMOS. Conventional circuit Design-Flow fails to use this inherent functional flexibility as CMOS based mapping considers a single logical output from logic gates. To address this, we propose an area-optimized technology mapping which uses this innate reconfigurability, offered by SiNW transistors for efficient circuit designs. To enable this objective, we use higher order functions (HOF) to encapsulate this extended functionality. Additionally, the electrical properties of SiNW allow us to take advantage of the available inverted forms of fan-ins for additional savings of area for XOR logic family. Experimental results using our technology mapping show that area of SiNW based logic design is less by an average of 18.38% as compared to CMOS flow for complete MCNC benchmarks suite. Further, we evaluate our flow for both reconfigurability-aware and static layout for SiNW based logic gates. The whole flow including the new SiNW based genlib and the modified ABC tool is made available under open source license to enable further research for any kind of emerging ambipolar transistors.
Shubham Rai, Michael Raitza, Akash Kumar 0001
DATE3
2018 A physical synthesis flow for early technology evaluation of silicon nanowire based reconfigurable FETs
abstract
Silicon Nanowire (SiNW) based reconfigurable field-effect transistors (RFETs) provide an additional gate terminal called the program gate which gives the freedom of programming p-type or n-type functionality for the same device at runtime. This enables the circuit designers to pack more functionality per computational unit. This saves processing costs as only one device type is required, and no doping and associated lithography steps are needed for this technology. In this paper, we present a complete design flow including both logic and physical synthesis for circuits based on SiNW RFETs. We propose layouts of logic gates, Liberty and LEF (Library Exchange Format) files to enable further research in the domain of these novel, functionally enhanced transistors. We show that in the first of its kind comparison, for these fully symmetrical reconfigurable transistors, the area after placement and routing for SiNW based circuits is 17% more than that of CMOS for MCNC benchmarks. Further, we discuss areas of improvement for obtaining better area results from the SiNW based RFETs from a fabrication and technology point of view. The future use of self-aligned techniques to structure two independent gates within a smaller pitch holds the promise of substantial area reduction.
Shubham Rai, Ansh Rupani, Dennis Walter, Michael Raitza, Andre Heinzig, Tim Baldauf, Jens Trommer, Christian Mayr 0001, Walter M. Weber, Akash Kumar 0001
DATE10
2018 Reloc - An Open-Source Vivado Workflow for Generating Relocatable End-User Configuration Tiles
abstract
As programmable logic is conquering data centers, FPGA resources become a ubiquitously available commodity available through the cloud. In these ecosystems, their employment must advance to fit the role of a dynamically managed, shared resource within multi-user environments. Programmable hardware vendors provide engineering workflows that only start to acknowledge the requirements of this thriving application context, which makes resource regularization, design migratability and isolated execution imperatives. Quite a few academic efforts have aimed at the key technique of design relocation on top of the vendor-provided workflows. Regularly, critical shortcomings remain and, in all cases, the engineering procedure is ridiculously complex. This paper identifies the significant gaps left by previous approaches and describes a thoroughly automated workflow for building a system infrastructure on FPGA that makes it a resource that is manageable by the OS. It highlights an end-to-end workflow to enable (a) the out-of-context synthesis of end-user designs against a thin module interface defined by appropriate synthesis constraints, which are (b) universally employable to all the user slots provisioned on an appropriately partitioned device without re-synthesis and (c) will operate there appropriately isolated merely communicating to the user host application through privately allocated communication channels.
Björn Gottschall, Thomas PreuBer, Akash Kumar 0001
FCCM3
2018 ParaDRo: A Parallel Deterministic Router Based on Spatial Partitioning and Scheduling
abstract
Routing of nets is one of the most time-consuming steps in the FPGA design flow. Existing works have described ways of accelerating the process through parallelization. However, only some of them are deterministic, and determinism is often achieved at the cost of speedup. In this paper, we propose ParaDRo, a parallel FPGA router based on spatial partitioning that achieves deterministic results while maintaining reasonable speedup. Existing spatial partitioning based routers do not scale well because the number of nets that can fully utilize all processors reduces as the number of processors increases. In addition, they route nets that are within a spatial partition sequentially. ParaDRo mitigates this problem by scheduling nets within a spatial partition to be routed in parallel if they do not have overlapping bounding boxes. Further parallelism is extracted by decomposing multi-sink nets into single-sink nets to minimize the amount of bounding box overlaps and increase the number of nets that can be routed in parallel. These improvements enable ParaDRo to achieve an average speedup of 5.4X with 8 threads with minimal impact on the quality of results.
Chin Hau Hoo, Akash Kumar 0001
FPGA2
2018 QoS-Aware Cross-Layer Reliability-Integrated FPGA-Based Dynamic Partially Reconfigurable System Partitioning
abstract
Dynamic Partial Reconfiguration (DPR) can be used for time-sharing of computing resources within Partially Reconfigurable Regions (PRRs) in FPGA-based systems. The heterogeneous partitioning in such systems allows the user to exploit the application-specific mapping of Partially Reconfigurable Modules (PRMs) to PRRs to implement more efficient designs. It offers increased opportunities in optimizing the reliability of the system across multiple layers - from the low-level physical one to the higher application layer. This method, called cross-layer reliability, can potentially exploit the application-specific tolerances to the quality of service (QoS) to tackle the increasing device fault-rates more cost-effectively by distributing the fault-mitigation to different layers. In this work, we propose a QoS-aware cross-layer reliability-integrated design methodology for FPGA-based DPR systems. Specifically, our methodology analyzes the requirements of the applications in terms of Functional Reliability, System Lifetime and Makespan to determine the best possible combinations of reliability-oriented design choices in different layers. We report up to an average of 24% and 30% performance improvements for single and multi-objective optimization-based system partitioning.
Siva Satyendra Sahoo, Tuan D. A. Nguyen, Bharadwaj Veeravalli, Akash Kumar 0001
FPT4
2018 Dataflow-Based Mapping of Spiking Neural Networks on Neuromorphic Hardware
abstract
Spiking Neural Networks (SNNs) are powerful computation engines for pattern recognition and image classification applications. Apart from application performance such as recognition and classification accuracy, system performance such as throughput becomes important when executing these applications on a hardware. We propose a systematic design-flow to map SNN-based applications on a crossbar-based neuromorphic hardware, guaranteeing application as well as system performance. Synchronous Dataflow Graphs (SDFGs) are used to model these applications with extended semantics to represent neural network topologies. Self-timed scheduling is then used to analyze throughput, incorporating hardware constraints such as synaptic memory, communication and I/O bandwidth of crossbars. Our design-flow integrates CARLsim, a GPU-accelerated application-level SNN simulator with SDF3, a tool for mapping SDFG on hardware. We conducted experiments with realistic and synthetic SNNs on representative neuromorphic hardware, demonstrating throughput-resource trade-offs for a given application performance. For throughput-constrained applications, we show average 20% reduction of hardware usage with 19% reduction in energy consumption. For throughput-scalable applications, we show an average 53% higher throughput compared to a state-of-the-art approach.
Anup Das 0001, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI2
2018 Protecting Communication in Many-Core Systems against Active Attackers
abstract
The advent of hardware Trojans is posing an increasing threat on complex integrated circuits. Network-on-Chip, the established communication infrastructure for many core systems-on-chip, are growing in complexity. Integration of third-party components, which are increasingly becoming valuable targets, exposes the surface for attacks through the injection of hardware Trojans. In this paper, we address active attacks on NoCs, and focus on the integrity of transmitted data. Basically, we use network coding for the transmission of data in order to increase efficiency and robustness.
Sadia Moriam, Elke Franz 0001, Paul Walther, Akash Kumar 0001, Thorsten Strufe, Gerhard P. Fettweis
ACM Great Lakes Symposium on VLSI4
2018 Emerging reconfigurable nanotechnologies: can they support future electronics?
abstract
Several emerging reconfigurable technologies have been explored in recent years offering device level runtime reconfigurability. These technologies offer the freedom to choose between p- and n-type functionality from a single transistor. In order to optimally utilize the feature-sets of these technologies, circuit designs and storage elements require novel design to complement the existing and future electronic requirements. An important aspect to sustain such endeavors is to supplement the existing design flow from the device level to the circuit level. This should be backed by a thorough evaluation so as to ascertain the feasibility of such explorations. Additionally, since these technologies offer runtime reconfigurability and often encapsulate more than one functions, hardware security features like polymorphic logic gates and on-chip key storage come naturally cheap with circuits based on these reconfigurable technologies. This paper presents innovative approaches devised for circuit designs harnessing the reconfigurable features of these nanotechnologies. New circuit design paradigms based on these nano devices will be discussed to brainstorm on exciting avenues for novel computing elements.
Shubham Rai, Srivatsa Rangachar Srinivasa, Patsy Cadareanu, Xunzhao Yin, Xiaobo Sharon Hu, Pierre-Emmanuel Gaillardon, Narayanan Vijaykrishnan, Akash Kumar 0001
ICCAD8
2018 A Self-Reconfiguring Cache Architecture to Improve Control Quality in Cyber-Physical Systems
abstract
Quality of control is a critical concern in Cyber-Physical Systems (CPS) which are comprised of multiple intercommunicating control applications. Due to complex timing behaviour of these systems, poor quality of control can lead to catastrophe. Recent studies showed that, conflict miss increment in the processor cache memory shared by concurrently running control applications can degrade control quality in CPS significantly. Increasing cache associativity can help to reduce conflict misses. However, the existing reconfigurable cache architectures that allow runtime modification of cache associativity are not capable to guaranty a newly chosen associativity's suitability for the forthcoming control quality requirement. Moreover, they have timing and energy related overheads. In this regard, this paper presents a novel, self-reconfiguring cache memory architecture "SeReMo". When conflict misses increase significantly, SeReMo reconfigures its associativity to better suit the current as well as future control quality demand. To trigger reconfiguration, a low overhead, non-strictly inclusive cache hierarchy-specific approach is used. Configurations with different associativity are generated using modules made of 4 cache lines and 7 special bits. Special replacement policy and indexing scheme are used to suit modular reconfiguration. SPEC CPU 2006 benchmark trace-driven simulation reveals that SeReMo reduces average number of conflict misses per line to 1/12951 of the state-of-the-art reconfigurable cache architecture at maximum (to 1/830 on average). As a result, execution time and energy consumption reduce by 48 hours at maximum (by 2/3 on average) and by 2907 Joules at maximum (86% on average) respectively.
Mohammad Shihabul Haque, Sriram Vasudevan, Alamuri Sriram Nihar, Arvind Easwaran, Akash Kumar 0001, Y. C. Tay
ISORC5
2017 Locality-Aware CTA Clustering for Modern GPUs
abstract
Cache is designed to exploit locality; however, the role of on-chip L1 data caches on modern GPUs is often awkward. The locality among global memory requests from different SMs (Streaming Multiprocessors) is predominantly harvested by the commonly-shared L2 with long access latency; while the in-core locality, which is crucial for performance delivery, is handled explicitly by user-controlled scratchpad memory. In this work, we disclose another type of data locality that has been long ignored but with performance boosting potential --- the inter-CTA locality. Exploiting such locality is rather challenging due to unclear hardware feasibility, unknown and inaccessible underlying CTA scheduler, and small in-core cache capacity. To address these issues, we first conduct a thorough empirical exploration on various modern GPUs and demonstrate that inter-CTA locality can be harvested, both spatially and temporally, on L1 or L1/Tex unified cache. Through further quantification process, we prove the significance and commonality of such locality among GPU applications, and discuss whether such reuse is exploitable. By leveraging these insights, we propose the concept of CTA-Clustering and its associated software-based techniques to reshape the default CTA scheduling in order to group the CTAs with potential reuse together on the same SM. Our techniques require no hardware modification and can be directly deployed on existing GPUs. In addition, we incorporate these techniques into an integrated framework for automatic inter-CTA locality optimization. We evaluate our techniques using a wide range of popular GPU applications on all modern generations of NVIDIA GPU architectures. The results show that our proposed techniques significantly improve cache performance through reducing L2 cache transactions by 55%, 65%, 29%, 28% on average for Fermi, Kepler, Maxwell and Pascal, respectively, leading to an average of 1.46x, 1.48x, 1.45x, 1.41x (up to 3.8x, 3.6x, 3.1x, 3.3x) performance speedups for applications with algorithm-related inter-CTA reuse.
Ang Li 0006, Shuaiwen Song, Weifeng Liu 0002, Xu Liu 0001, Akash Kumar 0001, Henk Corporaal
ASPLOS5
2017 Soft error-aware architectural exploration for designing reliability adaptive cache hierarchies in multi-cores
abstract
Mainstream multi-core processors employ large multilevel on-chip caches making them highly susceptible to soft errors. We demonstrate that designing a reliable cache hierarchy requires understanding the vulnerability interdependencies across different cache levels. This involves vulnerability analyses depending upon the parameters of different cache levels (partition size, line size, etc.) and the corresponding cache access patterns for different applications. This paper presents a novel soft error-aware cache architectural space exploration methodology and vulnerability analysis of multi-level caches considering their vulnerability interdependencies. Our technique significantly reduces exploration time while providing reliability-efficient cache configurations. We also show applicability/benefits for ECC-protected caches under multi-bit fault scenarios.
Arun Subramaniyan 0001, Semeen Rehman, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
DATE4
2017 Accounting for systematic errors in approximate computing
abstract
Approximate computing is gaining more and more attention as potential solution to the problem of increasing energy demand in computing. Several recent works focus on the application of deterministic approximate computing to arithmetic computations. Circuits for addition and multiplication are simplified, trading exactness for energy and/or speed. Recent approximation techniques for adders focus on modifications of individual full adders' truth tables or shortening carry chains. While the resulting error is usually characterized with statistical measures over the range of possible input/output combinations, the actual adder is a static nonlinear system regarding arithmetic operations and signal processing. The resulting unexpected effects present a challenge for adopting approximate computing as a widespread and standard application-level optimization technique. This paper focuses on the deterministic effects of approximate multi-bit adders, which are especially evident for certain input data in an otherwise well specified systems, showing the necessity to look beyond purely statistical measures. We show which fundamental principles are violated depending on the chosen approximation scheme, and how this choice affects practical applications. This can serve as a basis for designers to make informed decisions about the use of approximate adders at the application level.
Martin Bruestel, Akash Kumar 0001
DATE2
2017 Embracing approximate computing for energy-efficient motion estimation in high efficiency video coding
abstract
Approximate Computing is an emerging paradigm for developing highly energy-efficient computing systems. It leverages the inherent resilience of applications to trade output quality with energy efficiency. In this paper, we present a novel approximate architecture for energy-efficient motion estimation (ME) in high efficiency video coding (HEVC). We synthesized our designs for both ASIC and FPGA design flows. ModelSim gate-level simulations are used for functional and timing verification. We comprehensively analyze the impact of heterogeneous approximation modes on the power/energy-quality tradeoffs for various video sequences. To facilitate reproducible results for comparisons and further research and development, the RTL and behavioral models of approximate SAD architectures and constituting approximate modules are made available at https://sourceforge.net/projects/lpaclib/.
Walaa El-Harouni, Semeen Rehman, Bharath Srinivas Prabakaran, Akash Kumar 0001, Rehan Hafiz, Muhammad Shafique 0001
DATE4
2017 Exploiting transistor-level reconfiguration to optimize combinational circuits
abstract
Silicon nanowire reconfigurable field effect transistors (SiNW RFETs) abolish the physical separation of n-type and p-type transistors by taking up both roles in a configurable way within a doping-free technology. However, the potential of transistor-level reconfigurability has not been demonstrated in larger circuits, so far. In this paper, we present first steps to a new compact and efficient design of combinational circuits by employing transistor-level reconfiguration. We contribute new basic gates realized with silicon nanowires, such as 2/3-XOR and MUX gates. Exemplifying our approach with 4-bit, 8-bit and 16-bit conditional carry adders, we were able to reduce the number of transistors to almost one half. With our current case study we show that SiNW technology can reduce the required chip area by 16 despite larger size of the individual transistor, and improve circuit speed by 26%.
Michael Raitza, Akash Kumar 0001, Marcus Völp, Dennis Walter, Jens Trommer, Thomas Mikolajick, Walter M. Weber
DATE2
2017 ParaDiMe: A Distributed Memory FPGA Router Based on Speculative Parallelism and Path Encoding
abstract
The increase in speed and capacity of FPGAs is faster than the development of effective design tools to fully utilize it, and routing of nets remains as one of the most time-consuming stages of the FPGA design flow. While existing works have proposed methods of accelerating routing through parallelization, they are limited by the memory architecture of the system that they target. In this paper, we propose a distributed memory parallel FPGA router called ParaDiMe to address the limitations of existing works. ParaDiMe speculatively routes net in parallel and dynamically detects the need to reduce the number of active processes in order to achieve convergence. In addition, the synchronization overhead in ParaDiMe is significantly reduced through a careful design of the messaging protocol where paths to sinks are encoded in a space-efficient manner. Moreover, the frequency of synchronization is tuned to ensure convergence while minimizing the communication overhead. Compared to VTR, ParaDiMe achieves an average speedup of 19.8X with 32 processes while producing similar quality of results.
Chin Hau Hoo, Akash Kumar 0001
FCCM2
2017 Scrubbing Mechanism for Heterogeneous Applications in Reconfigurable Devices
abstract
Commercial off-the-shelf (COTS) reconfigurable devices have been recognized as one of the most suitable processing devices to be applied in nano-satellites, since they can satisfy and combine their most important requirements, namely processing performance, reconfigurability, and low cost. However, COTS reconfigurable devices, in particular Static-RAM Field Programmable Gate Arrays, can be affected by cosmic radiation, compromising the overall nano-satellite reliability. Scrubbing has been proposed as a mechanism to repair faults in configuration memory. However, the current scrubbing mechanisms are predominantly static, unable to adapt to heterogeneous applications and their runtime variations. In this article, a dynamically adaptive scrubbing mechanism is proposed. Through a window-based scrubbing scheduling, this mechanism adapts the scrubbing process to heterogeneous applications (composed of periodic/sporadic and streaming/DSP (Digital Signal Processing) tasks), as well as their reconfigurations and modifications at runtime. Conducted simulation experiments show the feasibility and the efficiency of the proposed solution in terms of system reliability metric and memory overhead.
Shyamsundar Venkataraman, Akash Kumar 0001
ACM Trans. Design Autom. Electr. Syst.3
2016 Critical points based register-concurrency autotuning for GPUs
Ang Li 0006, Shuaiwen Song, Akash Kumar 0001, Eddy Z. Zhang, Daniel G. Chavarría-Miranda, Henk Corporaal
DATE3
2016 Design and evaluation of reliability-oriented task re-mapping in MPSoCs using time-series analysis of intermittent faults
Siva Satyendra Sahoo, Akash Kumar 0001, Bharadwaj Veeravalli
DATE2
2016 A flexible inexact TMR technique for SRAM-based FPGAs
Shyamsundar Venkataraman, Akash Kumar 0001
DATE3
2016 PRFloor: An Automatic Floorplanner for Partially Reconfigurable FPGA Systems
abstract
Partial reconfiguration (PR) is gaining more attention from the research community because of its flexibility in dynamically changing some parts of the system at runtime. However, the current PR tools need the designer's involvement in manually specifying the shapes and locations for the PR regions (PRRs). It requires not only deep knowledge of the FPGA device, the system architecture, but also many trial-and-error attempts to find the best-possible floorplan. Therefore, many research works have been conducted to propose automatic floorplanners for PR systems. However, one of the most significant limitations of those works is that they only consider the PRRs and ignore all other static modules. In this paper, we propose a novel PR floorplanner called PRFloor. It takes into account all components in the system. The main ideas behind PRFloor are the unique recursive pseudo-bipartitioning heuristic using a new, simple, yet effective Nonlinear Integer Programming-based bipartitioner. The PRFloor performs very well in the experiments with various synthetic PR system setups with up to 130 modules, 24 PRRs and 85% of the FPGA resource. The average maximum clock frequency obtained for the actual PR systems implemented using PRFloor is even 3% higher than the similar systems without PR capability.
Tuan D. A. Nguyen, Akash Kumar 0001
FPGA2
2016 ParaFRo: A hybrid parallel FPGA router using fine grained synchronization and partitioning
abstract
Routing of nets is one of the most time-consuming steps in the FPGA design flow. While existing works have described ways of accelerating the process through parallelization, they are not scalable. In this paper, we propose ParaFRo, a two-phase hybrid parallel FPGA router using fine-grained synchronization and partitioning. The first phase of the router aims to exploit the maximum parallelism available by routing nets while minimizing load imbalance. Instead of resolving contention with expensive software transactional memory, synchronization among threads is realized using lightweight spin mutexes. In the case where the algorithm detects that convergence is not possible in phase one, it transitions into phase two where convergence is prioritized over maximum parallelism. To achieve convergence, each thread in phase two routes only congested nets that have been assigned to it by a partitioner. The partitioner aims to reduce the contention among threads at the cost of an unbalanced load. In addition, periodic rip up of the entire route tree is employed to break the algorithm out from a local minimum. When only congested nets are rerouted, ParaFRo with 8 threads achieves an average speedup of 26.2× relative to VTR. In contrast, existing works managed to obtain an average speedup of up to 9.42× with 8 threads. Besides, ParaFRo is able to maintain the high speedups while producing similar quality of result as VTR in terms of critical path delay. Finally, the quality of result is relatively independent of the number of the threads.
Chin Hau Hoo, Yajun Ha, Akash Kumar 0001
FPL3
2016 XNoC: A non-intrusive TDM circuit-switched Network-on-Chip
abstract
Network-on-Chip (NoC) is known as a scalable and high performance interconnect in Systems-on-Chip (SoCs) with multiple processing elements (PEs). Recently, the design paradigm of SoCs has shifted from static to dynamic runtime reconfigurable system. In these systems, the PEs can be loaded/unloaded on demand. Therefore, the NoC should be able to adapt as quickly as possible to the changes to maintain the performance of the systems. In this work, we present a non-intrusive runtime reconfigurable time-division-multiplexed circuit-switched NoC, XNoC, which offers the following benefits (1) it switches between different routes within a predictable latency that is strictly determined by the length of the route and the number of time slots; (2) the configuration process can be masked effectively by overlapping with communication and (3) the multi-cast service is supported with aggregate feedback from sink nodes. We propose an XSwitch which requires 3.5X less resource than the conventional switch with similar features. The overall resource cost of XNoC is also smaller than the most known NoC and the clock timing is up to 50% better. We also propose a novel distributed control plane to accelerate the reconfiguration process and to improve the scalability of NoC. The achieved reconfiguration speedup compared to the centralized control unit is up to 7.6X in certain conditions. On average, it takes only 74 clock cycles to activate a 12-hop connection.
Tuan D. A. Nguyen, Akash Kumar 0001
FPL2
2016 Architectural-space exploration of approximate multipliers
abstract
This paper presents an architectural-space exploration methodology for designing approximate multipliers. Unlike state-of-the-art, our methodology generates various design points by adapting three key parameters: (1) different types of elementary approximate multiply modules, (2) different types of elementary adder modules for summing the partial products, and (3) selection of bits for approximation in a wide-bit multiplier design. Generation and exploration of such a design space enables a wide-range of multipliers with varying approximation levels, each exhibiting distinct area, power, and output quality, and thereby facilitates approximate computing at higher abstraction levels. We synthesized our designs using Synopsys Design Compiler with a TSMC 45nm technology library and verified using ModelSim gate-level simulations. Power and quality evaluations for various designs are performed using PrimeTime and behavioral models, respectively. The selected designs are then deployed in a JPEG application. For reproducibility and to facilitate further research and development at higher abstraction layers, we have released the RTL and behavioral models of these approximate multipliers and adders as an open-source library at https://sourceforge.net/projects/lpaclib/.
Semeen Rehman, Walaa El-Harouni, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
ICCAD4
2016 SFU-Driven Transparent Approximation Acceleration on GPUs
abstract
Approximate computing, the technique that sacrifices certain amount of accuracy in exchange for substantial performance boost or power reduction, is one of the most promising solutions to enable power control and performance scaling towards exascale. Although most existing approximation designs target the emerging data-intensive applications that are comparatively more error-tolerable, there is still high demand for the acceleration of traditional scientific applications (e.g., weather and nuclear simulation), which often comprise intensive transcendental function calls and are very sensitive to accuracy loss. To address this challenge, we focus on a very important but long ignored approximation unit on today's commercial GPUs --- the special-function unit (SFU), and clarify its unique role in performance acceleration of accuracy-sensitive applications in the context of approximate computing. To better understand its features, we conduct a thorough empirical analysis on three generations of NVIDIA GPU architectures to evaluate all the single-precision and double-precision numeric transcendental functions that can be accelerated by SFUs, in terms of their performance, accuracy and power consumption. Based on the insights from the evaluation, we propose a transparent, tractable and portable design framework for SFU-driven approximate acceleration on GPUs. Our design is software-based and requires no hardware or application modifications. Experimental results on three NVIDIA GPU platforms demonstrate that our proposed framework can provide fine-grained tuning for performance and accuracy trade-offs, thus facilitating applications to achieve the maximum performance under certain accuracy constraints.
Ang Li 0006, Shuaiwen Song, Mark Wijtvliet, Akash Kumar 0001, Henk Corporaal
ICS4
2016 X: A Comprehensive Analytic Model for Parallel Machines
abstract
To continuously comply with Moore's Law, modern parallel machines become increasingly complex. Effectively tuning application performance for these machines therefore becomes a daunting task. Moreover, identifying performance bottlenecks at application and architecture level, as well as evaluating various optimization strategies, are becoming extremely difficult when the entanglement of numerous correlated factors is being presented. To tackle these challenges, we present a visual analytical model named "X". It is intuitive and sufficiently flexible to track all the typical features of a parallel machine. Different from the conventional analytic models that focus on the temporal state of a representative core or thread, our proposed X-model concentrates on the spatial state of the parallel machines -- the distribution of concurrent threads among different subsystems of these machines, while predicting the overall throughput based on such state. One major highlight of our model is its tractability as it only requires a small number of essential parameters from the application and architecture. Meanwhile, it is able to effectively help users investigate the combined-effects of different types of parallelism: the instruction-level-parallelism (ILP), the thread-level-parallelism (TLP), the memory-level-parallelism (MLP) and the data-level-parallelism (DLP). Through the X-model, developers and architects can quickly draw an intuitive figure called X-graph to identify performance bottlenecks and play "what-if " scenarios to evaluate the effectiveness of the proposed optimization techniques by investigating their individual and combined effects.
Ang Li 0006, Shuaiwen Song, Eric Brugel, Akash Kumar 0001, Daniel G. Chavarría-Miranda, Henk Corporaal
IPDPS4
2016 Machine Learning Approach to Generate Pareto Front for List-scheduling Algorithms
abstract
List Scheduling is one of the most widely used techniques for scheduling due to its simplicity and efficiency. In traditional list-based schedulers, a cost/priority function is used to compute the priority of tasks/jobs and put them in an ordered list. The cost function has been becoming more and more complex to cover increasing number of constraints in the system design. However, most of the existing list-based schedulers implement a static priority function that usually provides only one schedule for each task graph input. Therefore, they may not be able to satisfy the desire of system designers, who want to examine the trade-off between a number of design requirements (performance, power, energy, reliability ...). To address this problem, we propose a framework to utilize the Genetic Algorithm (GA) for exploring the design space and obtaining Pareto-optimal design points. Furthermore, multiple regression techniques are used to build predictive models for the Pareto fronts to limit the execution time of GA. The models are built using training task graph datasets and applied on incoming task graphs. The Pareto fronts for incoming task graphs are produced in time 2 orders of magnitude faster than the traditional GA, with only 4% degradation in the quality.
Pham Nam Khanh, Akash Kumar 0001, Khin Mi Mi Aung
SCOPES2
2016 Resource and Throughput Aware Execution Trace Analysis for Efficient Run-Time Mapping on MPSoCs
abstract
There have been several efforts on run-time mapping of applications on multiprocessor-systems-on-chip. These traditional efforts perform either on-the-fly processing or use design-time analyzed results. However, on-the-fly processing often leads to low-quality mappings, and design-time analysis becomes computationally costly for large-size problems and require huge storage for large number of applications. In this paper, we present a novel run-time mapping approach, where identification of an efficient mapping for a use-case is done by the online execution trace analysis of the active applications. The trace analysis facilitates for fast identification of the mapping while optimizing for the system resource usage and throughput of the active applications, leading to reduced energy consumption as well. By rapidly identifying the efficient mapping at run-time, the proposed approach overcomes the mappings' exploration time bottleneck for large-size problems and their storage overhead problem when compared to the traditional approaches. Our experiments show that on average the exploration time to identify the mapping is reduced $14 {\times }$ when compared to state-of-the-art approaches and storage overhead is reduced by 92%. Additionally, energy and resource savings are achieved along with identification of high-quality mapping.
Amit Kumar Singh 0002, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 Reliability and Energy-Aware Mapping and Scheduling of Multimedia Applications on Multiprocessor Systems
abstract
Lifetime reliability is an emerging concern in multiprocessor systems as escalating power density and hence temperature variation continues to accelerate wear-out leading to a growing prominence of device defects. In this paper, we propose a system-level approach that involves performance-aware mapping of multimedia applications on a multiprocessor system to jointly minimize energy consumption and temperature related wear-out. Fundamental to this approach is a simplified temperature model that incorporates not only the transient and the steady-state behavior (temporal effect), but also the temperature dependency on the surrounding cores (spatial effect). This model is validated against the temperature obtained using theHotSpottool with transient and steady-state simulations, and is shown to be accurate within 5.5°C, leading to an MTTF estimation accuracy of an average 21 percent with respect to the state-of-the-art approaches. The proposed temperature model is integrated in a gradient-based fast heuristic that controls the voltage and frequency of the cores to limit the average and peak temperature leading to a longer lifetime, simultaneously minimizing the energy consumption. Lifetime computation considers task remapping, which is a common feature available in modern multiprocessor systems. A linear programming approach is then proposed to distribute the cores of a multiprocessor system among concurrent applications to maximize the lifetime. Experiments conducted with a set of synthetic and real-life applications represented as synchronous data flow graphs demonstrate that the proposed approach minimizes energy consumption by an average 24 percent with 47 percent increase in lifetime. For concurrent applications, the proposed lifetime-aware core distribution results in an average 10 percent improvement in lifetime as compared to performance-based core distribution.
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli
IEEE Trans. Parallel Distributed Syst.2
2016 Analysis and Mapping for Thermal and Energy Efficiency of 3-D Video Processing on 3-D Multicore Processors
abstract
Three-dimensional video processing has high computation requirements and multicore processors realized in 3-D integrated circuits (ICs) provide promising high performance computing platforms. However, the conventional approaches to accelerate the computations involved in 3-D video processing do not exploit the high performance potential of 3-D ICs. In this paper, we propose an application-driven methodology that performs efficient mapping of 3-D video applications' components on 3-D multicores to achieve high performance (throughput). The methodology involves an extensive application analysis to exploit the spatial and temporal correlation available in 3-D neighborhood. Afterward, it leverages the correlation and thermal properties of different 3-D views to perform an efficient mapping of 3-D video processing on cores available at different layers of 3-D IC. The goal is to optimize energy consumption and peak temperature while meeting the throughput requirement. Experiments show 76% reduction in communication energy along with reduction in peak temperature when compared with approaches exploiting architecture characteristics only.
Amit Kumar Singh 0002, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
IEEE Trans. Very Large Scale Integr. Syst.3
2015 Accelerating non-volatile/hybrid processor cache design space exploration for application specific embedded systems
abstract
In this article, we propose a technique to accelerate nonvolatile or hybrid of volatile and nonvolatile processor cache design space exploration for application specific embedded systems. Utilizing a novel cache behavior modeling equation and a new accurate cache miss prediction mechanism, our proposed technique can accelerate NVM or hybrid FIFO processor cache design space exploration for SPEC CPU 2000 applications up to 249 times compared to the conventional approach.
Mohammad Shihabul Haque, Ang Li 0006, Akash Kumar 0001, Qingsong Wei
ASP-DAC3
2015 Dynamically adaptive scrubbing mechanism for improved reliability in reconfigurable embedded systems
abstract
Commercial off-the-shelf (COTS) reconfigurable devices have been recognized as one of the most suitable processing devices to be applied in satellites, since they can satisfy and combine their most important requirements, namely processing performance, reconfigurability and low cost. However, COTS reconfigurable devices, in particular Static-RAM Field Programmable Gate Arrays (FPGAs), can be affected by cosmic radiation, compromising the overall satellite reliability. Scrubbing has been proposed as a mechanism to repair faults in configuration memory. However, the current scrubbing mechanisms are predominantly static and unable to adapt to run-time variations in applications. In this paper, a dynamically adaptive scrubbing mechanism is proposed. Through a window-based scrubbing scheduling, this mechanism adapts the scrubbing process to the reconfigurations and modifications on the FPGA user-design at runtime. Conducted simulation experiments show the feasibility and the efficiency of the proposed solution in terms of system reliability and memory overhead.
Shyamsundar Venkataraman, Akash Kumar 0001
DAC3
2015 Workload uncertainty characterization and adaptive frequency scaling for energy minimization of embedded systems
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi
DATE2
2015 (AS)2: accelerator synthesis using algorithmic skeletons for rapid design space exploration
Shakith Fernando, Mark Wijtvliet, Cedric Nugteren, Akash Kumar 0001, Henk Corporaal
DATE4
2015 Exploiting loop-array dependencies to accelerate the design space exploration with high level synthesis
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001, Khin Mi Mi Aung
DATE3
2015 Dynamic reconfigurable puncturing for secure wireless communication
Jude Angelo Ambrose, Akash Kumar 0001, Sri Parameswaran
DATE3
2015 Designing inexact systems efficiently using elimination heuristics
Shyamsundar Venkataraman, Akash Kumar 0001, Jeremy Schlachter, Christian C. Enz
DATE2
2015 ParaLaR: A parallel FPGA router based on Lagrangian relaxation
abstract
Routing of nets is one of the most time consuming steps in the FPGA design flow. While existing works have described ways of accelerating the process through parallelization, they are not scalable. In this paper, we propose a scalable way of parallelizing the routing algorithm through Lagrangian relaxation. The FPGA routing problem is formulated as a linear programming problem, and the channel width constraints, which limit the amount of parallelism, are relaxed by incorporating them into the objective function. The result of the relaxation yields independent sub-problems that we solve using minimum Steiner tree algorithms. Our approach outperforms the state-of-the-art FPGA parallel router by producing an average self-relative speedup of 7.05X with 8 threads, reduces the total wire length by 22.4% on average and has similar channel width requirements as VPR, albeit at the cost of 7.5% longer critical path. Another advantage of our algorithm is that the number of threads and the order in which the nets are routed has totally no impact on the quality of result.
Chin Hau Hoo, Akash Kumar 0001, Yajun Ha
FPL2
2015 An automated technique to generate relocatable partial bitstreams for Xilinx FPGAs
abstract
Partial reconfiguration is a technique used to increase the flexibility of an FPGA-based system by reprogramming parts of the system dynamically without interrupting the operation of the other modules. Despite the runtime benefits offered by partially reconfigurable (PR) systems, creating and storing partial bitstreams (PBs) are becoming major concerns for system architects when the numbers of reconfigurable partitions (RPs) and PR modules (PRMs) increase. It takes significant amount of time to generate the PBs for PR systems with large number of RPs and PRMs. More importantly, when the mapping relationship between PRMs and RPs is many-to-many, several almost-identical PBs of one PRM must be stored separately which leads to inefficient utilization of the memory storage. Therefore, bitstream relocation is drawing interests from the research community as a viable solution. Yet almost none of the works are able to demonstrate a coherent method to not only create relocatable PBs for complex and large PRMs in variable-size RPs but also how to do that automatically to free the designer from the tedious and error prone manual processes. In this paper, we propose a new technique to fill that gap. The method is successfully developed for Xilinx Virtex 7 devices using Vivado design tool flow.
Roel Oomen, Tuan D. A. Nguyen, Akash Kumar 0001, Henk Corporaal
FPL3
2015 Transit: A Visual Analytical Model for Multithreaded Machines
abstract
With the extraordinary growth of cores and threads in today's multithreaded machines, analyzing and tuning the performance of such platforms becomes a challenging task. In this paper, we propose an intuitive and visualizable model for analyzing the performance of contemporary highly concurrent multithreaded machines. Based on flow balancing between service demand and service supply of the memory system, the model draws an intuitive figure to characterize machine state, identify bottlenecks and determine optimization directions. The tractability of the model is highlighted as it only requires two parameters from the workload. Our model achieves 90% and 83% prediction accuracy for computation throughput on Fermi and Kepler GPUs over the 16 applications from Rodinia benchmark.
Ang Li 0006, Y. C. Tay, Akash Kumar 0001, Henk Corporaal
HPDC3
2015 Fine-Grained Synchronizations and Dataflow Programming on GPUs
abstract
The last decade has witnessed the blooming emergence of many-core platforms, especially the graphic processing units (GPUs). With the exponential growth of cores in GPUs, utilizing them efficiently becomes a challenge. The data-parallel programming model assumes a single instruction stream for multiple concurrent threads (SIMT); therefore little support is offered to enforce thread ordering and fine-grained synchronizations. This becomes an obstacle when migrating algorithms which exploit fine-grained parallelism, to GPUs, such as the dataflow algorithms.
Ang Li 0006, Gert-Jan van den Braak, Henk Corporaal, Akash Kumar 0001
ICS4
2015 Generic scrubbing-based architecture for custom error correction algorithms
abstract
Scrubbing has been considered as an efficient mechanism to repair faults in the FPGA's configuration memory, when they are placed in harsh environments. By using this elementary mechanism, several academic solutions/algorithms based on error correction codes (ECCs) have been proposed. However, most of these proposed solutions are only theoretical and do not properly deal with the implementation concerns. With this paper we propose a generic scrubbing-based hardware architecture and design flow for implementing custom error correction algorithms based on ECCs. A conducted case study implementing and evaluating three different algorithms shows the feasibility and the efficiency of the proposed architecture and design flow.
Shyamsundar Venkataraman, Akash Kumar 0001
RSP3
2015 Adaptive and transparent cache bypassing for GPUs
abstract
In the last decade, GPUs have emerged to be widely adopted for general-purpose applications. To capture on-chip locality for these applications, modern GPUs have integrated multilevel cache hierarchy, in an attempt to reduce the amount and latency of the massive and sometimes irregular memory accesses. However, inferior performance is frequently attained due to serious congestion in the caches results from the huge amount of concurrent threads. In this paper, we propose a novel compile-time framework for adaptive and transparent cache bypassing on GPUs. It uses a simple yet effective approach to control the bypass degree to match the size of applications' runtime footprints. We validate the design on seven GPU platforms that cover all existing GPU generations using 16 applications from widely used GPU benchmarks. Experiments show that our design can significantly mitigate the negative impact due to small cache sizes and improve the overall performance. We analyze the performance across different platforms and applications. We also propose some optimization guidelines on how to efficiently use the GPU caches.
Ang Li 0006, Gert-Jan van den Braak, Akash Kumar 0001, Henk Corporaal
SC3
2015 Execution Trace-Driven Energy-Reliability Optimization for Multimedia MPSoCs
abstract
Multiprocessor systems-on-chip (MPSoCs) are becoming a popular design choice in current and future technology nodes to accommodate the heterogeneous computing demand of a multitude of applications enabled on these platform. Streaming multimedia and other communication-centric applications constitute a significant fraction of the application space of these devices. The mapping of an application on an MPSoC is an NP-hard problem. This has attracted researchers to solve this problem both as stand-alone (best-effort) and in conjunction with other optimization objectives, such as energy and reliability. Most existing studies on energy-reliability joint optimization are static—that is, design time based. These techniques fail to capture runtime variability such as resource unavailability and dynamism associated with application behaviors, which are typical of multimedia applications. The few studies that consider dynamic mapping of applications do not consider throughput degradation, which directly impacts user satisfaction. This article proposes a runtime technique to analyze the execution trace of an application modeled as Synchronous Data Flow Graphs (SDFGs) to determine its mapping on a multiprocessor system with heterogeneous processing units for different fault scenarios. Further, communication energy is minimized for each of these mappings while satisfying the throughput constraint. Experiments conducted with synthetic and real SDFGs demonstrate that the proposed technique achieves significant improvement with respect to the state-of-the-art approaches in terms of throughput and storage overhead with less than 20% energy overhead.
Anup Das 0001, Amit Kumar Singh 0002, Akash Kumar 0001
ACM Trans. Reconfigurable Technol. Syst.3
2015 Autonomous Soft-Error Tolerance of FPGA Configuration Bits
abstract
Field-programmable gate arrays (FPGAs) are increasingly susceptible to radiation-induced single event upsets (SEUs). These upsets are predominant in a space environment; however, with increasing use of static RAM (SRAM) in modern FPGAs, these SEUs are gaining prominence even in a terrestrial environment. SEUs can flip SRAM bits of FPGA, potentially altering the functionality of the implemented design. This has motivated FPGA designers to investigate techniques to protect the FPGA configuration bits against such inadvertent bit flips (soft error). Traditionally, triple modular redundancy (TMR) is used to protect the FPGA bit flips. Increasing design complexity and limited battery life motivate for alternative approaches for soft-error tolerance. In this article, we propose a technique to improve autonomous fault-masking capabilities of a design by maximizing the number of zeros or ones in lookup tables (LUTs). The technique analyzes critical configuration bits and utilizes spare resources (XOR gates and carry chains) of FPGAs to selectively manipulate the logic implemented in LUTs using two operations: LUT restructuring and LUT decomposition. We implemented the proposed approach for Xilinx Virtex-6 FPGAs and validated the same with a wide set of designs from the MCNC, IWLS 2005, and ITC99 benchmark suites. Results demonstrate that the proposed logic restructuring maximizes logic 0 (or 1) of LUTs by an average of 20%, achieving 80% fault masking with no area overhead. The fault rate of the entire design is reduced by 60% on average as compared to the existing techniques. Furthermore, the logic decomposition algorithm provides incremental fault-tolerance capabilities and achieves an additional 5% fault masking with an average 7% increase in slice usage. The complete methodology is implemented into a tool for Xilinx FPGA and is made available online for the benefit of the research community. The algorithms are lightweight, and the whole design flow (including Xilinx Place and Route) was completed in 75 minutes for the largest benchmark in the set.
Anup Das 0001, Shyamsundar Venkataraman, Akash Kumar 0001
ACM Trans. Reconfigurable Technol. Syst.3
2014 Reinforcement Learning-Based Inter- and Intra-Application Thermal Optimization for Lifetime Improvement of Multicore Systems
abstract
The thermal profile of multicore systems vary both within an application's execution (intra) and also when the system switches from one application to another (inter). In this paper, we propose an adaptive thermal management approach to improve the lifetime reliability of multicore systems by considering both inter- and intra-application thermal variations. Fundamental to this approach is a reinforcement learning algorithm, which learns the relationship between the mapping of threads to cores, the frequency of a core and its temperature (sampled from on-board thermal sensors). Action is provided by overriding the operating system's mapping decisions using affinity masks and dynamically changing CPU frequency using in-kernel governors. Lifetime improvement is achieved by controlling not only the peak and average temperatures but also thermal cycling, which is an emerging wear-out concern in modern systems. The proposed approach is validated experimentally using an Intel quad-core platform executing a diverse set of multimedia benchmarks. Results demonstrate that the proposed approach minimizes average temperature, peak temperature and thermal cycling, improving the mean-time-to-failure (MTTF) by an average of 2x for intra-application and 3x for inter-application scenarios when compared to existing thermal management techniques. Furthermore, the dynamic and static energy consumption are also reduced by an average 10% and 11% respectively.
Anup Das 0001, Rishad A. Shafik, Geoff V. Merrett, Bashir M. Al-Hashimi, Akash Kumar 0001, Bharadwaj Veeravalli
DAC5
2014 Temperature aware energy-reliability trade-offs for mapping of throughput-constrained applications on multimedia MPSoCs
abstract
This paper proposes a design-time (offline) analysis technique to determine application task mapping and scheduling on a multiprocessor system and the voltage and frequency levels of all cores (offline DVFS) that minimize application computation and communication energy, simultaneously minimizing processor aging. The proposed technique incorporates (1) the effect of the voltage and frequency on the temperature of a core; (2) the effect of neighboring cores' voltage and frequency on the temperature (spatial effect); (3) pipelined execution and cyclic dependencies among tasks; and (4) the communication energy component which often constitutes a significant fraction of the total energy for multimedia applications. The temperature model proposed here can be easily integrated in the design space exploration for multiprocessor systems. Experiments conducted with MPEG-4 decoder on a real system demonstrate that the temperature using the proposed model is within 5% of the actual temperature clearly demonstrating its accuracy. Further, the overall optimization technique achieves 40% savings in energy consumption with 6% increase in system lifetime.
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli
DATE2
2014 Combined DVFS and mapping exploration for lifetime and soft-error susceptibility improvement in MPSoCs
abstract
Energy and reliability optimization are two of the most critical objectives for the synthesis of multiprocessor systems-on-chip (MPSoCs). Task mapping has shown significant promise as a low cost solution in achieving these objectives as standalone or in tandem as well. This paper proposes a multi-objective design space exploration to determine the mapping of tasks of an application on a multiprocessor system and voltage/frequency level of each tasks (exploiting the DVFS capabilities of modern processors) such that the reliability of the platform is improved while fulfilling the energy budget and the performance constraint set by system designers. In this respect, the reliability of a given MPSoC platform incorporates not only the impact of voltage and frequency on the aging of the processors (wear-out effect) but also on the susceptibility to soft-errors - a joint consideration missing in all existing works in this domain. Further, the proposed exploration also incorporates soft-error tolerance by selective replication of tasks, making the proposed approach an interesting blend of reactive and proactive fault-tolerance. The combined objective of minimizing core aging together with the susceptibility to transient faults under a given performance/energy budget is solved by using a multi-objective genetic algorithm exploiting tasks' mapping, DVFS and selective replication as tuning knobs. Experiments conducted with reallife and synthetic application graphs clearly demonstrate the advantage of the proposed approach.
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli, Cristiana Bolchini, Antonio Miele
DATE2
2014 Accelerating Volume Image Registration through Correlation Ratio Based Methods on GPUs
abstract
Volume image registration is a basic component of medical image processing which traditionally requires long computation time. In this paper, we propose five Correlation Ratio based schemes that explore the design space for Graphics Processing Unit (GPU) acceleration. Through comparisons among these five schemes, we present the trade-off between benefits and overheads of introducing shadow histograms on various storage (shared memory, global memory) by different level execution units (thread, warp, thread block). Compared to Mutual Information based methods, these Correlation Ratio based methods require less resources for shadow histograms, a faster storage therefore could be exploited to achieve better performance which is shown in our experiments. Particularly, the fifth scheme completely avoids updating conflicts of histogram calculation, leading to a substantial performance improvement (over 18x speedup) over the native FLIRT version. It reduces the registration time from over 100s to less than 6s for two typical 256x256x160 3D images.
Ang Li 0006, Akash Kumar 0001
DSD2
2014 Design Space Exploration to Accelerate Nelder-Mead Algorithm Using FPGA
abstract
Nelder-Mead algorithm (NMA) is the best-known algorithm for multidimensional optimization without involving derivative computations. Due to the simplicity in implementation and the fast convergent property of NMA, it is widely used in the fields of statistics, engineering, physics and medical sciences. In practice, when objective function is complicated, the optimization procedure requires a lot of computation efforts, leading to a time-consuming process. This work introduces a NMA solver engine fully implemented on FPGA hardware and performs design space exploration to provide various solutions suitable for FPGA device.
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001, Khin Mi Mi Aung
FCCM3
2014 PR-HMPSoC: A versatile partially reconfigurable heterogeneous Multiprocessor System-on-Chip for dynamic FPGA-based embedded systems
abstract
FPGA-based heterogeneous Multiprocessor Systems-on-Chip (HMPSoCs) are becoming quite popular for high performance embedded systems because of their powerful computational ability and relatively flexible architecture to adapt to unexpected system requirement changes. However, with the insatiable demands of supporting an extensive range of applications beyond the limited resources of FPGA chip and shorter time-to-market, many research works on partially reconfigurable (PR) FPGA architectures have been conducted to fulfill the needs. Those have yet to fully provide a versatile framework to exploit the flexibility of PR such as hardware/software task migration and bitstream relocation; more importantly, the on-chip debug features to access all processors currently loaded in the system are compromised because of the lack of native-support from vendor tools. In this paper, a novel PR-HMPSoC architecture for dynamic FPGA-based embedded system is proposed to provide solutions for all of the above issues. The results from the experimental system consisting of one static Microblaze and three PR Microblaze/hardware accelerators connected by a Network-on-Chip show that the architecture is very promising with just 8% reduction in operating frequency.
Tuan D. A. Nguyen, Akash Kumar 0001
FPL2
2014 Criticality-aware scrubbing mechanism for SRAM-based FPGAs
abstract
Scrubbing has been considered as an effective mechanism to provide fault-tolerance in Static-RAM (SRAM)-based Field Programmable Gate Arrays (FPGAs). However, the current scrubbing techniques execute without considering the criticality and timing of the user tasks implemented in the FPGA. They often do not execute the scrubbing process in the right instant, which minimizes the probability of each task being executed without transient faults. Moreover, these current solutions are not adapted to the tasks' fault-tolerance requirements, since they may not properly protect the most critical tasks in the system. However, if they do it, they waste resources with the less critical tasks. In this paper, a new scrubbing mechanism is proposed. This new approach adapts the scrubbing mechanism to the tasks' execution, by a proper scheduling and according to their criticality. A proposed heuristic finds a feasible scrubbing schedule for each hardware task. Firstly, the minimum scrubbing periods are computed according to the criticality of each implemented hardware task. Secondly, a proper scrubbing schedule following the EDL (Earliest Deadline as Late as possible) algorithm is found, maximizing the reliability of the system. The experimental results show up to 79% improvements on the system reliability, achieved without wasting scrubbing resources.
Shyamsundar Venkataraman, Anup Das 0001, Akash Kumar 0001
FPL4
2014 A bit-interleaved embedded hamming scheme to correct single-bit and multi-bit upsets for SRAM-based FPGAs
abstract
Single Event Upsets (SEUs) inadvertently change the configuration bits of Static-RAM (SRAM)-based Field Programmable Gate Arrays (FPGAs), leading to erroneous output until the error has been corrected. Scrubbing using an Error Correction Code (ECC) such as hamming is a popular method to correct such faults. However, current works either require a large external memory to store the ECCs or can at most correct only one error in a frame. This paper proposes a novel bit-interleaved embedded hamming scheme along with scrubbing, to correct single (SBUs) and multi-bit upsets (MBUs) in SRAM-based FPGAs. This scheme does not require an external memory to store the ECCs, as they are embedded within the configuration memory itself. Experiments conducted on various benchmarks show that the proposed scheme can handle multiple errors per frame very well, with an embedding efficiency of over 99.3%.
Shyamsundar Venkataraman, Anup Das 0001, Akash Kumar 0001
FPL4
2014 Multi-directional error correction schemes for SRAM-based FPGAs
abstract
Readback scrubbing is considered as an effective mechanism to correct errors in Static-RAM (SRAM)-based Field Programmable Gate Arrays (FPGAs). However, current solutions have a low error correction percentage per unit area overhead. This paper proposes two new error detection/correction mechanisms that combine frame readback scrubbing with error correction codes (ECCs) that are applied in multiple directions, to achieve a high error correction percentage per unit area overhead. Experiments conducted show that the proposed schemes have an excellent error correction percentage (over 99%), especially for multi-bit upsets, while using up to 59.37% lesser area overhead compared with other state-of-the-art.
Shyamsundar Venkataraman, Sidharth Maheshwari, Akash Kumar 0001
FPL4
2014 Leakage and performance aware resource management for 2D dynamically reconfigurable FPGA architectures
abstract
The variety of applications for field programmable gate arrays (FPGAs) is continuously growing, thus it is important to address power consumption issues during the operation. As technological node shrinks, leakage power becomes increasingly critical in overall power consumption of FPGA. The technique of configuration pre-fetching (loads configurations as soon as possible) adopted to achieve high performance is one of the major reasons of leakage waste since regions containing reconfiguration information cannot be powered down in between the time gap of reconfiguration and execution. In this work, we present a heuristic approach to minimize the leakage power consumption for two-dimensional reconfigurable FPGA architectures. The heuristic scheduler is based on list scheduling and exploits dynamic priority for sorting the tasks into schedule order and a cost function for cell allocation. Farthest placement scheme is adopted for anti-fragmentation purpose. The cost function provides control to compromise between leakage dissipation and schedule length.
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001
FPL4
2014 FPGA-based high throughput XTS-AES encryption/decryption for storage area network
abstract
The key issue to improve the performance for secure large-scale Storage Area Network (SAN) applications lies in the speed of its encryption/decryption module. Software-based encryption/decryption cannot meet throughput requirements. To solve this problem, we propose a FPGA-based XTS-AES encryption/decryption to suit the needs for secure SAN applications with high throughput requirements. Besides throughput, area optimization is also considered in this proposed design. First, we reuse the same AES encryption to produce the tweak value and unify the operations of AES encryption/decryption in XTS-AES encryption/decryption. Second, we transfer the computations of AES encryption/decryption from GF(28) to GF(24)2, which enables us move the map and the inverse map functions outside the AES round. Third, we propose to support the SubBytes and the inverse SubBytes by the same hardware component. Finally, pipelined registers have been inserted into the proposed unrolled architecture for XTS-AES encryption/decryption. The experiments show that the proposed design achieves 36.2 Gbits/s throughput using 6784 slices on XC6VLX240T FPGA.
Yi Estelle Wang, Akash Kumar 0001, Yajun Ha
FPT2
2014 A multi-stage leakage aware resource management technique for reconfigurable architectures
abstract
Shrinking size of transistors has enabled us to integrate more and more logic elements into FPGA chips leading to higher computing power. However, it also brings serious concern to the leakage power dissipation of the FPGA devices. One of the major reasons for leakage power dissipation in FPGA is the utilization of prefetching technique to minimize the reconfiguration overhead (delay) in Partially Reconfigurable (PR) FPGAs. This technique creates delays between the reconfiguration and execution parts of a task, which may lead up to 44% leakage power of FPGA since the SRAM-cells containing reconfiguration information cannot be powered down. In this work, a resource management approach containing scheduling, placement and post-placement stages has been proposed to address the aforementioned issue. In scheduling stage, a leakage-aware cost function is derived to cope with the leakage power. The placement stage uses a cost function that allows designers to decide a trade-off between performance and leakage-saving. The post-placement stage employs a heuristic approach and shows further improvements. Experiments show that our approach can achieve large leakage savings for both synthetic and real life applications with acceptable extended deadline. Furthermore, different variants of the proposed approach can reduce leakage power by 40-65% when compared to a performance-driven approach and by 15-43% when compared to state-of-the-art works.
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001
ACM Great Lakes Symposium on VLSI3
2014 Design and robust scheduling of nano-satellite swarm for synthetic aperture radar applications
abstract
This paper presents the design and robust scheduling of nano-satellite (nanosat) swarm for synthetic aperture radar (SAR) applications. Based on power budget and bandwidth limit of nanosats, the nanosats' form factor and swarm size are chosen to ensure requirements on ground resolution and signal-to-noise ratio. An energy-efficient and robust scheduling considering stochastic failures is proposed using scenario optimization with convexification. The effectiveness of our proposed scheduling approach is verified with mathematical rigor as well as extensive simulation results on a realistic SAR application using strip and spot modes.
Chee Khiang Pang, Akash Kumar 0001, Cher-Hiang Goh, Cao Vinh Le
ICARCV2
2014 Lightweight Bare-Metal Stateful Firewall
abstract
A firewall is a crucial security element in modern computer networks. This work investigates and demonstrates the implementation of a lightweight TCP/IP firewall in a bare-metal environment, on a commercial embedded ARM device. Compared to an implementation having an operating system (OS), using bare-metal design enables reduction of exposure to potential vulnerabilities in OS code, and provides a more dependable system. The implemented firewall provides both static and stateful filtering capabilities, and is configurable in a user-friendly way. As the architecture of the commercial hardware used was not available under closed source licensing, it was discovered through analysis at both hardware and software levels. Some challenges were encountered, and tools were developed to address these. The prototype is validated through functional testing in a controlled environment successfully.
Yihuan Xing, Ford-Long Wong, Akash Kumar 0001
PRDC3
2014 A multi-stage thermal management strategy for 3D multicores
abstract
3D integration technology has the potential to enhance IC performance, improve functionality and lessen wiring of ICs. However, it poses several challenges, where the key challenge is heat generation from internal active layers due to power dissipation. To mitigate this challenge, thermal aware design has become a necessity. Towards thermal aware design, this paper proposes a two stage design technique. In the first stage, a temperature-power thermal model is created to calculate power dissipated by an IC at an input temperature. The proposed model calculates power dissipated by 2D and 3D ICs with an average error of 0.37% and 25% respectively. Power calculation helps in process variation, validation of power models and minimization of temperature gradients. In the second stage, thermal aware mapping is performed for the ICs. For thermal aware mapping, three mapping algorithms are proposed to account for different resource (processor) availability scenarios. Each algorithm utilizes temperature-power thermal model (from the first design stage) to map applications to processing elements in a 3D IC. The proposed two stage design technique performs faster temperature to power calculations than existing techniques. It provides a simplified approach to mapping compared to existing techniques by utilizing power dissipated by processing elements to map applications.
Dipika Suresh, Amit Kumar Singh 0002, Akash Kumar 0001
RSP3
2014 Communication and migration energy aware task mapping for reliable multiprocessor systems
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli
Future Gener. Comput. Syst.2
2014 Energy-aware task mapping and scheduling for reliable embedded computing systems
abstract
Task mapping and scheduling are critical in minimizing energy consumption while satisfying the performance requirement of applications enabled on heterogeneous multiprocessor systems. An area of growing concern for modern multiprocessor systems is the increase in the failure probability of one or more component processors. This is especially critical for applications where performance degradation (e.g., throughput) directly impacts the quality of service requirement. This article proposes a design-time (offline) multi-criterion optimization technique for application mapping on embedded multiprocessor systems to minimize energy consumption for all processor fault-scenarios. A scheduling technique is then proposed based on self-timed execution to minimize the schedule storage and construction overhead at runtime. Experiments conducted with synthetic and real applications from streaming and nonstreaming domains on heterogeneous MPSoCs demonstrate that the proposed technique minimizes energy consumption by 22% and design space exploration time by 100x, while satisfying the throughput requirement for all processor fault-scenarios. For scalable throughput applications, the proposed technique achieves 30% better throughput per unit energy, compared to the existing techniques. Additionally, the self-timed execution-based scheduling technique minimizes schedule construction time by 95% and storage overhead by 92%.
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli
ACM Trans. Embed. Comput. Syst.2
2013 TRISHUL: A single-pass optimal two-level inclusive data cache hierarchy selection process for real-time MPSoCs
abstract
Hitherto discovered approaches analyze the execution time of a real-time application on all the possible cache hierarchy setups to find the application specific optimal two-level inclusive data cache hierarchy to reduce cost, space and energy consumption while satisfying the time deadline in real-time Multi-Processor Systems on Chip (MPSoC). These brute-force like approaches can take years to complete. Alternatively, application's memory access trace driven crude estimation methods can find a cache hierarchy quickly by compromising the accuracy of results. In this article, for the first time, we propose a fast and accurate application's trace driven approach to find the optimal real-time application specific two-level inclusive data cache hierarchy. Our proposed approach “TRISHUL” predicts the optimal cache hierarchy performance first and then utilizes that information to find the optimal cache hierarchy quickly. TRISHUL can suggest a cache hierarchy, which has up to 128 times smaller size, up to 7 times faster compared to the suggestion of the state-of-the-art crude trace driven two-level inclusive cache hierarchy selection approach for the application traces analyzed.
Mohammad Shihabul Haque, Akash Kumar 0001, Yajun Ha, Shaobo Luo
ASP-DAC2
2013 Aging-aware hardware-software task partitioning for reliable reconfigurable multiprocessor systems
abstract
Homogeneous multiprocessor systems with reconfigurable area (also known as Reconfigurable Multiprocessor Systems) are emerging as a popular design choice in current and future technology nodes to meet the heterogeneous computing demand of a multitude of applications enabled on these platforms. Application specific mapping decisions on such a platform involve partitioning a given application into software tasks (executed on one or more of the general purpose processors, GPPs) and the hardware tasks (realized as dedicated hardware on the reconfigurable area) to optimize and/or satisfy design constraints such as reliability, performance and design cost. Improving the reliability considering transient faults by increasing the number of checkpoints negatively impacts the reliability considering permanent faults. This trade-off is ignored in all prior studies on task mapping and scheduling. This paper proposes an optimization technique to decide the optimal number of checkpoints for the software tasks which minimizes aging of the GPPs while maximizing the transient fault-tolerance of the overall platform (GPPs and the reconfigurable area) and satisfying design cost and performance. Experiments conducted with synthetic and real-life application task graphs (cyclic and acyclic) demonstrate that the proposed technique minimizes aging and improves the platform lifetime by an average 60% as compared to the existing transient fault-aware techniques. Further, a gradient-based heuristic is proposed to minimize the design space exploration time by upto 500× with less than 5% deviation from optimal solution.
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli
CASES2
2013 Energy optimization by exploiting execution slacks in streaming applications on multiprocessor systems
abstract
Dynamic voltage and frequency scaling (DVFS) offers great potential for optimizing the energy efficiency of Multiprocessor Systems-on-Chip (MPSoCs). The conventional approaches for processor voltage and frequency adjustment are not suitable for streaming multimedia applications due to the cyclic nature of dependencies in the executing tasks which can potentially violate the throughput constraints. In this paper, we propose a methodology that applies DVFS for such cyclic dependent tasks. The methodology involves an off-line analysis that assumes worst-case execution times of tasks to identify the executions that can be slowed down and an on-line analysis to utilize the slacks arising from tasks that finish their execution before the worst-case execution times. Thus, the methodology minimizes energy consumption during both off-line and on-line analysis while satisfying the throughput constraints. Experiments based on models of real-life streaming multimedia applications show that the proposed methodology reduces the overall energy consumption by 43% when compared to existing approaches.
Amit Kumar Singh 0002, Anup Das 0001, Akash Kumar 0001
DAC3
2013 Mapping on multi/many-core systems: survey of current and emerging trends
abstract
The reliance on multi/many-core systems to satisfy the high performance requirement of complex embedded software applications is increasing. This necessitates the need to realize efficient mapping methodologies for such complex computing platforms. This paper provides an extensive survey and categorization of state-of-the-art mapping methodologies and highlights the emerging trends for multi/many-core systems. The methodologies aim at optimizing system's resource usage, performance, power consumption, temperature distribution and reliability for varying application models. The methodologies perform design-time and run-time optimization for static and dynamic workload scenarios, respectively. These optimizations are necessary to fulfill the end-user demands. Comparison of the methodologies based on their optimization aim has been provided. The trend followed by the methodologies and open research challenges have also been discussed.
Amit Kumar Singh 0002, Muhammad Shafique 0001, Akash Kumar 0001, Jörg Henkel
DAC3
2013 Reliability-driven task mapping for lifetime extension of networks-on-chip based multiprocessor systems
abstract
Shrinking transistor geometries, aggressive voltage scaling and higher operating frequencies have negatively impacted the lifetime reliability of embedded multi-core systems. In this paper, a convex optimization-based task-mapping technique is proposed to extend the lifetime of a multiprocessor systems-on-chip (MPSoCs). The proposed technique generates mappings for every application enabled on the platform with variable number of cores. Based on these results, a novel 3D-optimization technique is developed to distribute the cores of an MPSoC among multiple applications enabled simultaneously. Additionally, reliability of the underlying network-on-chip links is also addressed by incorporating aging of links in the objective function. Our formulations are developed for directed acyclic graphs (DAGs) and synchronous dataflow graphs (SDFGs), making our approach applicable for streaming as well as non-streaming applications. Experiments conducted with synthetic and real-life application graphs demonstrate that the proposed approach extends the lifetime of an MPSoC by more than 30% when applications are enabled individually as well as in tandem.
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli
DATE2
2013 Communication and migration energy aware design space exploration for multicore systems with intermittent faults
abstract
Shrinking transistor geometries, aggressive voltage scaling and higher operating frequencies have negatively impacted the dependability of embedded multicore systems. Most existing research works on fault-tolerance have focused on transient and permanent faults of cores. Intermittent faults are a separate class of defects resulting from on-chip temperature, pressure and voltage variations and lasting for a few cycles to several seconds or more. Operations of cores impacted by intermittent faults are suspended during these cycles but come back alive when conditions become favorable. This paper proposes a technique to model the availability of multiprocessor systems-on-chip (MPSoCs) with intermittent and reparable device defects. This model is based on Markov chain with stochastic fault distribution and can be applied even for permanent faults. Based on this model, a design space pruning technique is proposed to select a set of task mappings (with variable resource usage), which minimizes the task communication energy while satisfying the MPSoC availability constraint. Moreover, task migration overhead is also minimized, which is an important consideration for frequently occurring intermittent and temperature related faults, where prolonged system downtime during task re-mapping is not desired. Experiments conducted with real-life and synthetic application task graphs demonstrate that the proposed technique minimizes communication energy by 30% and reduces migration overhead by 50% as compared to the existing approaches.
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli
DATE2
2013 Incorporating Energy and Throughput Awareness in Design Space Exploration and Run-Time Mapping for Heterogeneous MPSoCs
abstract
The advancement in process technology has enabled integration of different types of processing cores into a single chip towards creating heterogeneous Multiprocessor Systems-on-Chip (MPSoCs). While providing high level of computation power to support complex applications, these modern systems also introduce novel challenges for system designers, like managing a huge number of mappings (application tasks to processing cores allocations) that increases exponentially with the number of cores and their types. This paper presents a mapping approach that computes multiple energy-throughput trade-off points (mappings) at design-time and uses one of these points at run-time based on desired throughput and current resource availability while optimizing for the overall energy consumption. While significantly reducing the complexity of the design space exploration (DSE) to compute mappings at design-time, the proposed strategy still evaluates mappings for all the resource combinations of the platform, providing efficient mapping solutions for all the scenarios of system architecture at run-time. Moreover, the proposed approach performs energy-aware mapping at run-time while utilizing the DSE results. Experimental results show that proposed strategy achieves better energy-throughput trade-off points, covers all the resource combinations and reduces energy consumption up to 24.93% at design-time and additionally 17.8% at run-time when compared to state-of-the-art techniques.
Pham Nam Khanh, Amit Kumar Singh 0002, Akash Kumar 0001, Khin Mi Mi Aung
DSD3
2013 RAPIDITAS: RAPId Design-Space-Exploration Incorporating Trace-Based Analysis and Simulation
abstract
Simulation-based Design Space Exploration (DSE) to evaluate all possible mappings for a given application and Multiprocessor-System-on-Chip (MPSoC) platform is computationally costly for large problems. Even using efficient exploration methodologies to evaluate the mappings cannot overcome the evaluation time bottleneck. This paper presents a novel DSE methodology that analyzes the execution trace to prune the vast design space. Simulations are employed only on the pruned design points (mappings), hence reducing the number of simulations. The methodology performs iterative exploration and provides premier mappings requiring different number of processors, which can be used at run-time subject to desired performance and available platform processors. We evaluate our methodology by using models of real-life multimedia applications and demonstrate that the DSE time is reduced by 72% while generating high quality mappings.
Amit Kumar Singh 0002, Anup Das 0001, Akash Kumar 0001
DSD3
2013 High Speed Video Processing Using Fine-Grained Processing on FPGA Platform
abstract
This summary paper1proposes an FPGA-based array processor which performs Laplacian filtering on a 40 by 40 pixel grayscale video. The architecture comprises of bit-serial pixel processors interconnected to give a two-dimensional mesh array. This architecture features the novel use of partial reconfiguration which transfers data to and fro the array. Each processor occupies a configurable logic block and achieves a target frame rate of 10000 frames per second, at an operating frequency of 0.31 MHz on the Virtex-6 ML605 Evaluation Kit. The detailed correspondence between the contents of slice lookup tables and the Virtex-6 bitstream format is also documented.
Zhi Ping Ang, Akash Kumar 0001, Yajun Ha
FCCM2
2013 Improving autonomous soft-error tolerance of FPGA through LUT configuration bit manipulation
abstract
Soft-errors in LUT configuration bits of FPGAs can alter the functionality of an implemented design, rendering it useless, unless re-programmed. This paper proposes a technique to improve autonomous fault-masking capabilities of a design by maximizing the number of zeros or ones in LUTs. The technique utilizes spare resources (XOR gates and carry chain) of FPGA devices to selectively manipulate LUT contents using two operations - LUT restructuring and LUT decomposition. Experiments conducted with a wide set of benchmarks from MCNC, IWLS 2005 and ITC99 benchmark suite on Xilinx Virtex 6 FPGA board demonstrate that the proposed methodology maximizes logic 0/1 of LUTs by an average 20% achieving 80% fault-masking with no area overhead. The fault-rate of the entire design is reduced by 60% on average as compared to the existing techniques. Further, an additional 5% fault-masking can be achieved with a 7% increase in slice usage.
Anup Das 0001, Shyamsundar Venkataraman, Akash Kumar 0001
FPL3
2013 MAMPSX: A demonstration of rapid, predictable HMPSOC synthesis
abstract
Heterogeneous Multiprocessor systems-on-chip (HMPSoC) are becoming popular as a means of meeting energy efficiency requirements of modern embedded systems. However, as these HMPSoCs run multimedia applications as well, they also need to meet realtime requirements. Designing HMPSoCs with predictable timing behavior is a key challenge, as the current design methods for these platforms are semi-automated, non-predictable, or support limited heterogeneity. In this demonstration, we present a design framework to rapidly generate and implement predictable HMPSoC designs. It takes the application specifications and the architecture model as input and generates the entire HMPSoC, for FPGA prototyping, that meets the throughput constraints of the application. We also present results of a case study that computes the performance-power tradeoffs of an industrial vision application. A tool-chain targeting the Xilinx Zynq FPGA is also presented.
Shakith Fernando, Mark Wijtvliet, Firew Siyoum, Yifan He 0002, Sander Stuijk, Akash Kumar 0001, Henk Corporaal
FPL6
2013 A directional coarse-grained power gated FPGA switch box and power gating aware routing algorithm
abstract
Leakage power has become an important component of the total power consumption in FPGAs as process technology shrinks. In addition, a significant amount of leakage power in FPGAs is consumed by the routing resources. Therefore, leakage power reduction in FPGAs should begin with the routing resources. In this paper, we propose a novel directional coarse-grained power gating architecture for switch boxes. In addition, the existing VPR routing algorithm has been adapted with a new cost function to support the new power gating architecture. Results have shown that the new cost function yields an average improvement of 22% as compared to the existing VPR cost function in terms of the number of power gating regions that can be turned off.
Chin Hau Hoo, Yajun Ha, Akash Kumar 0001
FPL3
2013 Real-time and low power embedded ℓ1-optimization solver design
abstract
Basis pursuit denoising (BPDN) is an optimization method used in cutting edge computer vision and compressive sensing research. Although hosting a BPDN solver on an embedded platform is desirable because analysis can be performed in real-time, existing solvers are generally unsuitable for embedded implementation due to either poor run-time performance or high memory usage. To address the aforementioned issues, this paper proposes an embedded-friendly solver which demonstrates superior run-time performance, high recovery accuracy and competitive memory usage compared to existing solvers. For a problem with 5000 variables and 500 constraints, the solver occupies a small memory footprint of 29 kB and takes 0.14 seconds to complete on the Xilinx Zynq Z-7020 system-on-chip. The same problem takes 0.19 seconds on the Intel Core i7-2620M, which runs at 4 times the clock frequency and 114 times the power budget of the Z-7020. Without sacrificing runtime performance, the solver has been highly optimized for power constrained embedded applications. By far this is the first embedded solver capable of handling large scale problems with several thousand variables.
Zhi Ping Ang, Akash Kumar 0001
FPT2
2013 MAMPSx: A design framework for rapid synthesis of predictable heterogeneous MPSoCs
abstract
Heterogeneous Multiprocessor System-on-Chips (HMPSoC) are becoming popular as a means of meeting energy efficiency requirements of modern embedded systems. However, as these HMPSoCs run multimedia applications as well, they also need to meet real-time requirements. Designing these predictable HMPSoCs is a key challenge, as the current design methods for these platforms are either semi-automated, non-predictable, or have limited heterogeneity. In this paper, we propose a design framework to generate and program HMPSoC designs in a rapid and predictable manner. It takes the application specifications and the architecture model as input and generates the entire HMPSoC, for FPGA prototyping, that meets the throughput constraints. The experimental results show that our framework can provide a conservative bound on the worst-case throughput of the FPGA implementation. We also present results of a case study that computes the area-power trade-offs of an industrial vision application. The entire design space exploration of all configurations was completed in 8 hours. A tool-chain targeting the Xilinx Zynq FPGA is also presented.
Shakith Fernando, Firew Siyoum, Yifan He 0002, Akash Kumar 0001, Henk Corporaal
RSP4
2013 CADSE: communication aware design space exploration for efficient run-time MPSoC management
Amit Kumar Singh 0002, Akash Kumar 0001, Jigang Wu, Thambipillai Srikanthan
Frontiers Comput. Sci.2
2012 Minimizing Power Consumption of Spatial Division Based Networks-on-Chip Using Multi-path and Frequency Reduction
abstract
With an increasing number of processing elements being integrated on a single die, networks-on-chip (NoCs) are emerging as a significant contributor to overall chip power consumption. While some solutions have been proposed to reduce this power consumption, none of them can be applied to spatial division multiplexing (SDM)-based NoCs. In this paper, we introduce a method to minimize the power consumption of an SDM-based NoC by frequency minimization, while still satisfying the bandwidth requirements. The problem is integrated with the connection-routing problem which is modeled as a mixed-integer quadratic constrained problem (MIQCP). However, solving this MIQCP formulation directly using existing solvers is infeasible for large use-cases. We propose a two-step approach by first computing the minimum feasible frequency for the entire network taking bandwidth of all connections into consideration. This first step reduces the frequency-minimization-routing MIQCP problem into a routing-only mixed-integer linear programming (MILP) problem. In the second step, this MILP problem is solved using a standard ILP solver. Two other techniques are proposed to solve the routing and frequency minimization problem. Experiments are performed with synthetic examples and a case-study with JPEG decoder to evaluate the performance and results of the three methods. MILP-based approach achieves up to 55% power reduction as compared to the other methods albeit at the cost of higher execution time.
Sheng Hao Wang, Anup Das 0001, Akash Kumar 0001, Henk Corporaal
DSD3
2012 Acceleration of distance-to-default with hardware-software co-design
abstract
The role of Credit Rating Agencies has come under intense scrutiny in the recent past due to their failure to accurately rate the issuers of debt obligations and instruments. With growing uncertainty in the markets and the need for accurate results, credit rating algorithms are getting more and more complex by the day. Distance-to-default (DTD), or the leverage indicator, is one of the key indicators in credit research that determines the probabilities-of-default of firms. The greater the amount of historic data available for a given firm, the higher is the accuracy of the DTD results. However, this directly translates to a higher processing time and increases the costs of computation. The DTD computation features a linear workflow that is suited for implementation on a Field Programmable Gate Array (FPGA) which could lead to a more efficient and low-cost solution. Application of embedded platforms in implementing such algorithms have the potential to reduce the power consumption through parallelism and utilise an optimised solution offered by reconfigurable logic and customized hardware. In addition to the hardware solution, the right balance of software implementation can give the performance of such complex and intensive processes, an added boost. In this paper, we explore that very prospect of a hardware-software co-design and suitably implement a prototype of the DTD algorithm. The software in our design is partly run on a 2.9GHz Intel processor and the FPGA soft-core processor (Microblaze) which is implemented on a Xilinx Virtex-6 ML605 FPGA, accelerated by hardware coprocessors. This resulted in a 16.6× and 317.17× speedup in the computation of the implied asset value and the log-likelihood function respectively as compared to a pure software implementation on a 2.9GHz Intel processor.
Izaan Allugundu, Pranay Puranik, Yat Piu Lo, Akash Kumar 0001
FPL4
2012 An area-efficient partially reconfigurable crossbar switch with low reconfiguration delay
abstract
With the increasing number of processors in Multiprocessor System-on-Chips (MPSoCs), Network-on-Chips (NoCs) are replacing conventional buses as the interprocessor communication architecture. Since different use cases might be running on MPSoCs, there is a need for dynamically reconfigurable NoC. However, most dynamically reconfigurable NoCs have a large area overhead due to the additional reconfiguration logic. While recently some dynamically reconfigurable NoCs have been proposed based on partial reconfiguration (PR), they have high reconfiguration delay and require off-line bitstream generation for all possible scenarios. The problem lies with the design of the crossbar switch, which is the fundamental component of a NoC. In this paper, a novel partially reconfigurable crossbar switch design with low area requirement, low reconfiguration delay and runtime bitstream generation is presented. The crossbar switch is built from lookup tables (LUTs), and reconfiguration is done by modifying the LUTs' content through PR. Reconfiguration delay is minimized by constraining the placement of the LUTs into the least number of configurable logic block columns and identifying the configuration frames that are responsible for LUTs' content. The novel crossbar switch design achieves an area saving of up to 84% and reconfiguration delay minimization of up to 78%. It can be used to realize any network topology, and Clos, Benes and single stage crossbar topologies are evaluated in the paper.
Chin Hau Hoo, Akash Kumar 0001
FPL2
2012 Development of an FPGA-based real-time P300 speller
abstract
A Brain Computer Interface (BCI) is a system that allows direct communication between a computer and the human brain. Though the main application for BCIs is in rehabilitation of disabled patients, they are increasingly being used in other application scenarios as well. Most of the current BCI systems are based on personal computers. However, there is an increased interest in implementing BCIs for portable platforms as well, such as mobile phones and Field Programmable Gate Arrays (FPGAs) owing to low cost, power and portability. This paper proposes a low-cost FPGA based BCI speller application. The proposed system combines a stimulation panel, data acquisition and FPGA based real-time signal processing. The BCI system demonstrated here is a speller, which allows the user to use his/her brain signals to communicate directly with the application and spell out words by merely looking at the screen. The system achieves an accuracy of 65.37% when utilizing 2 rounds of data per character and an accuracy of 100% when utilizing 20 rounds of data per character.
Kanav Khurana, Rajesh Chandrasekhara Panicker, Akash Kumar 0001
FPL4
2012 Energy-Aware Communication and Remapping of Tasks for Reliable Multimedia Multiprocessor Systems
abstract
Shrinking transistor geometries, aggressive voltage scaling and higher operating frequencies have negatively impacted the dependability of embedded multiprocessor systems-on-chip (MPSoCs). Fault-tolerance and energy efficiency are the two most desired features of modern-day MPSoCs. For most of the multimedia applications, task communication energy constitutes more than 40% of the overall application energy. In this paper, an integer linear programming (ILP) based approach is proposed to reduce the communication energy and fault-tolerant migration overhead of throughput-constrained multimedia applications modeled using synchronous data flow graphs (SDFGs). The ILP is solved at compile-time for all fault-scenarios to generate task-core mappings satisfying an application throughput requirement. These mappings are stored in a table which is looked up at run-time as and when faults occur. Experiments conducted with real and synthetic applications demonstrate that the proposed technique reduces communication energy by an average 40% and migration overhead by 33% as compared to the existing fault-tolerant techniques.
Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli
ICPADS2
2012 Fault-aware task re-mapping for throughput constrained multimedia applications on NoC-based MPSoCs
abstract
Shrinking transistor geometry and aggressive voltage scaling are leading to growing concerns on the reliability of multiprocessor systems. Majority of streaming multimedia applications are characterized by fixed throughput requirements; violation of which directly impacts user experience. None of the prior research considers joint treatment of throughput and task-migration overhead, both of which are essential for fault-tolerance of throughput-constrained multimedia multiprocessor systems. In this paper, we propose to remap tasks from faulty processors with the objective of minimizing the migration overhead while satisfying throughput constraints. The proposed technique is based on extensive design-time analysis of different fault scenarios to determine optimal mappings from the throughput-migration overhead Pareto space. These mappings are stored in a table and are looked-up at run-time to migrate tasks as and when faults occur. Applications are modeled using Synchronous Data Flow graphs (SDFG) to consider cyclic dependencies of tasks, typically found in multimedia systems. Experiments performed with synthetic and real application graphs demonstrate that the migration overhead can be reduced by 26% on average while still meeting throughput constraints. Moreover, by selecting an appropriate initial processor-task mapping, migration overhead can be further reduced by 15% on average.
Anup Das 0001, Akash Kumar 0001
RSP2
2012 A design flow for partially reconfigurable heterogeneous multi-processor platforms
abstract
Modern multiprocessor systems-on-chip (MPSoCs) are expected to handle multi-application use cases. As the number and complexity of these applications scale, resource allocation to meet the application throughput requirement is becoming quite a challenge. In this paper, a complete design flow is proposed for partially reconfigurable heterogeneous MPSoC platforms. The proposed flow determines the minimum resources required to map and guarantee the throughput of applications in all use-cases. Further, a suitable mapping for each application is chosen so that energy consumption is minimized. Experiments conducted with a set of synthetic benchmarks and real-life applications clearly demonstrate the advantage of our approach over homogeneous or fully reconfigurable designs. The proposed design flow achieves more than 50% energy savings when the number of configurations is not optimized. With configuration-optimization, our flow results in 75% reduction in the number of configurations with 5% reduction in energy.
Jiashu Li, Anup Das 0001, Akash Kumar 0001
RSP3
2012 Accelerating throughput-aware runtime mapping for heterogeneous MPSoCs
abstract
Modern embedded systems need to support multiple time-constrained multimedia applications that often employ multiprocessor-systems-on-chip (MPSoCs). Such systems need to be optimized for resource usage and energy consumption. It is well understood that a design-time approach cannot provide timing guarantees for all the applications due to its inability to cater for dynamism in applications. However, a runtime approach consumes large computation requirements at runtime and hence may not lend well to constrained-aware mapping. In this article, we present a hybrid approach for efficient mapping of applications in such systems. For each application to be supported in the system, the approach performs extensive design-space exploration (DSE) at design time to derive multiple design points representing throughput and energy consumption at different resource combinations. One of these points is selected at runtime efficiently, depending upon the desired throughput while optimizing for energy consumption and resource usage. While most of the existing DSE strategies consider a fixed multiprocessor platform architecture, our DSE considers a generic architecture, making DSE results applicable to any target platform. All the compute-intensive analysis is performed during DSE, which leaves for minimum computation at runtime. The approach is capable of handling dynamism in applications by considering their runtime aspects and providing timing guarantees. The presented approach is used to carry out a DSE case study for models of real-life multimedia applications: H.263 decoder, H.263 encoder, MPEG-4 decoder, JPEG decoder, sample rate converter, and MP3 decoder. At runtime, the design points are used to map the applications on a heterogeneous MPSoC. Experimental results reveal that the proposed approach provides faster DSE, better design points, and efficient runtime mapping when compared to other approaches. In particular, we show that DSE is faster by 83% and runtime mapping is accelerated by 93% for some cases. Further, we study the scalability of the approach by considering applications with large numbers of tasks.
Amit Kumar Singh 0002, Akash Kumar 0001, Thambipillai Srikanthan
ACM Trans. Design Autom. Electr. Syst.2
2011 A hybrid strategy for mapping multiple throughput-constrained applications on MPSoCs
abstract
Modern embedded systems are based on Multiprocessor-Systems-on-Chip (MPSoCs) to meet the strict timing deadlines of multiple applications. MPSoC resources must be utilized efficiently by mapping the applications in throughput-aware manner in order to meet throughput constraints for each of them. A design-time methodology is applicable only to predefined set of applications with static behavior, which is incapable of handling dynamism in applications. On the other hand, a run-time approach can cater to the dynamism but cannot provide timing guarantees for all the applications due to large computation requirements at run-time. This paper presents a hybrid flow which performs compute intensive analysis at design-time to derive multiple resource-throughput trade-off points and selects one of these at run-time subject to available resources and desired throughput. Experimental results show that the design-time analysis is faster by 39%, provides better trade-off points and the run-time mapping is speeded up by 93% when compared to state-of-the-art techniques.
Amit Kumar Singh 0002, Akash Kumar 0001, Thambipillai Srikanthan
CASES2
2010 Mapping real-life applications on run-time reconfigurable NoC-based MPSoC on FPGA
abstract
Multiprocessor systems-on-chip (MPSoC) are required to fulfill the performance demand of modern real-life embedded applications. These MPSoCs are employing Network-on-Chip (NoC) for reasons of efficiency and scalability. Additionally, these systems need to support run-time reconfiguration of their components to cater to dynamically changing demands of the system. Designing and programming such systems for real-life applications prove to be a major challenge. This paper demonstrates the designing of reconfigurable NoC-based MPSoC and programming it for real-life applications. The NoC is reconfigured at run-time to support different combinations of multiple applications at different times. The platform is verified with a case study executing the parallelized C-codes of a simple producer-consumer and JPEG decoder applications on a NoC-based MPSoC on a Xilinx FPGA. Based on our investigations to map the applications on a 3 × 3 platform, we show that the NoC reconfiguration overhead is kept at a minimum and the platform utilizes 85% of the total available slices of Virtex-5 FPGA. Moreover, we show that the proposed approach is highly scalable when targeting for large number of applications.
Amit Kumar Singh 0002, Akash Kumar 0001, Thambipillai Srikanthan, Yajun Ha
FPT2
2010 An area-efficient dynamically reconfigurable Spatial Division Multiplexing network-on-chip with static throughput guarantee
abstract
With an increasing trend to implement Network-on-Chip (NoC)-based Multi-Processor Systems-on-Chips (MPSoCs), NoCs need to have guaranteed services and be dynamically reconfigurable. Many current NoCs consume too much area and cannot support dynamic reconfiguration. In this paper, we present an area-efficient Spatial Division Multiplexing (SDM)-based NoC. We replaced area consuming 32-bit to M-bit serializers with 32-bit to 1-bit serializers in the network interface and incur almost no loss in performance. We also restrict flexibility in the router to achieve further area reduction. A separate area-efficient control network, with an overhead of 3.9% of the total area of the NoC, is developed to support dynamic reconfiguration.
Zhiyao Joseph Yang, Akash Kumar 0001, Yajun Ha
FPT2
2010 CA-MPSoC: An automated design flow for predictable multi-processor architectures for multiple applications
Ahsan Shabbir, Akash Kumar 0001, Sander Stuijk, Bart Mesman, Henk Corporaal
J. Syst. Archit.2
2010 Communication-aware heuristics for run-time task mapping on NoC-based MPSoC platforms
Amit Kumar Singh 0002, Thambipillai Srikanthan, Akash Kumar 0001, Jigang Wu
J. Syst. Archit.3
2010 Iterative Probabilistic Performance Prediction for Multi-Application Multiprocessor Systems
abstract
Modern embedded devices are increasingly becoming multiprocessor with the need to support a large number of applications to satisfy the demands of users. Due to a huge number of possible combinations of these multiple applications, it becomes a challenge to predict their performance. This becomes even more important when applications may be dynamically started and stopped in the system. Since modern embedded systems allow users to download and add applications at run-time, a complete design-time analysis is not always possible. This paper presents a new technique to accurately predict the performance of multiple applications mapped on a multiprocessor platform. Iterative probabilistic analysis is used to estimate the time spent by tasks during their contention phase, and thereby predicting the performance of applications. The approach is scalable with the number of applications and processors in the system. As compared to earlier techniques, this approach is much faster and scalable, while still improving the accuracy. The analysis takes 300 ¿s on a 500 MHz processor for ten applications. Since multimedia applications are increasingly becoming more dynamic, results of a case-study with applications with varying execution times are also presented. In addition, results of a case-study with real applications executing on a field-programmable gate array multiprocessor platform are shown.
Akash Kumar 0001, Bart Mesman, Henk Corporaal, Yajun Ha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2008 Vectorization of Reed Solomon Decoding and Mapping on the EVP
abstract
Reed Solomon (RS) codes are used in a variety of (wireless) communication systems. Although commonly implemented in dedicated hardware, this paper explores the mapping of high-throughput RS decoding on vector DSPs. The four modules of such a decoder, viz. syndrome computation, key equation solver, Chien search, and Forney pose different vectorization challenges. Their vectorizations are explained in detail, including optimizations specific for embedded vector processor (EVP). For RS (255,239), this solution is benchmarked vs published implementations, and scalability up to vector size 64 is explored. The best and the worst case throughput of our implementation is 8 times and 2 times higher respectively than other architectures.
Akash Kumar 0001, Kees van Berkel 0001
DATE1
2008 Analyzing composability of applications on MPSoC platforms
Akash Kumar 0001, Bart Mesman, Bart D. Theelen, Henk Corporaal, Yajun Ha
J. Syst. Archit.1
2008 Multiprocessor systems synthesis for multiple use-cases of multiple applications on FPGA
abstract
Future applications for embedded systems demand chip multiprocessor designs to meet real-time deadlines. The large number of applications in these systems generates an exponential number of use-cases. The key design automation challenges are designing systems for these use-cases and fast exploration of software and hardware implementation alternatives with accurate performance evaluation of these use-cases. These challenges cannot be overcome by current design methodologies which are semiautomated, time consuming, and error prone. In this article, we present a design methodology to generate multiprocessor systems in a systematic and fully automated way for multiple use-cases . Techniques are presented to merge multiple use-cases into one hardware design to minimize cost and design time, making it well suited for fast design-space exploration (DSE) in MPSoC systems. Heuristics to partition use-cases are also presented such that each partition can fit in an FPGA, and all use-cases can be catered for. The proposed methodology is implemented into a tool for Xilinx FPGAs for evaluation. The tool is also made available online for the benefit of the research community and is used to carry out a DSE case study with multiple use-cases of real-life applications: H263 and JPEG decoders. The generation of the entire design takes about 100 ms, and the whole DSE was completed in 45 minutes, including FPGA mapping and synthesis. The heuristics used for use-case partitioning reduce the design-exploration time elevenfold in a case study with mobile-phone applications.
Akash Kumar 0001, Shakith Fernando, Yajun Ha, Bart Mesman, Henk Corporaal
ACM Trans. Design Autom. Electr. Syst.1
2007 A Probabilistic Approach to Model Resource Contention for Performance Estimation of Multi-featured Media Devices
abstract
The number of features that are supported in modern multimedia devices is increasing faster than ever. Estimating the performance of such applications when they are running on shared resources is becoming increasingly complex. Simulation of all possible use-cases is very time-consuming and often undesirable. In this paper, a new technique is proposed based on probabilistically estimating the performance of concurrently executing applications that share resources. Two different methods of employing this approach are presented and compared with state-of-the-art technique, and with achieved performance found through extensive simulations. The results are within 15% of simulation result (considered as reference case) and up to ten times better than a worst-case estimation approach. The approach scales very well with increasing number of applications, and can also be applied at run-time for admission control.
Akash Kumar 0001, Bart Mesman, Henk Corporaal, Bart D. Theelen, Yajun Ha
DAC1
2007 Interactive presentation: An FPGA design flow for reconfigurable network-based multi-processor systems on chip
abstract
Multi-processor systems on chip (MPSoC) platforms are becoming increasingly more heterogeneous and are shifting towards a more communication-centric methodology. Networks on chip (NoC) have emerged as the design paradigm for scalable on-chip communication architectures. As the system complexity grows, the problem emerges as how to design and instantiate such a NoC-based MPSoC platform in a systematic and automated way. This paper presents an integrated flow to automatically generate a highly configurable NoC-based MPSoC for FPGA instantiation. The system specification is done on a high level of abstraction, relieving the designer of error-prone and time consuming work. The flow uses the state-of-the-art /Ethereal NoC, and silicon hive processing cores, both configurable at design- and run-time. The authors use this flow to generate a range of sample designs whose functionality has been verified on a Celoxica RC300E development board. The board, equipped with a Xilinx Virtex II 6000, also offers a huge number of peripherals, and shows how the insertion is automated in the design for easy debugging and prototyping
Akash Kumar 0001, Andreas Hansson 0001, Jos Huisken, Henk Corporaal
DATE1
2007 Multi-processor System-level Synthesis for Multiple Applications on Platform FPGA
abstract
Multiprocessor systems-on-chip (MPSoC) are being developed in increasing numbers to support the high number of applications running on modern embedded systems. Designing and programming such systems prove to be a major challenge. Most of the current design methodologies rely on creating the design by hand, and are therefore error-prone and time-consuming. This also limits the number of design points that can be explored. While some efforts have been made to automate the flow and raise the abstraction level, these are still limited to single-application designs. In this paper, we present a design methodology to generate and program MPSoC designs in a systematic and automated way for multiple applications. The architecture is automatically inferred from the application specifications, and customized for it. The flow is ideal for fast design space exploration (DSE) in MPSoC systems. We present results of a case study to compute the buffer-throughput trade-offs in real-life applications, H263 and JPEG decoders. The generation of the entire project takes about 100ms, and the whole DSE was completed in 45 minutes, including the FPGA mapping and synthesis.
Akash Kumar 0001, Shakith Fernando, Yajun Ha, Bart Mesman, Henk Corporaal
FPL1
2006 Global Analysis of Resource Arbitration for MPSoC
abstract
Modern day applications require use of multi-processor systems for reasons of scalability and power efficiency. As more and more applications are integrated on a single device, mapping and analyzing them on a multi-processor system becomes a multi-dimensional problem. Each possible set of applications that can be active simultaneously leads to a different use-case (also referred to as scenario) that the system has to be verified and tested for. Analyzing the feasibility and resource utilization of all possible use-cases is very demanding and often infeasible. In this paper, we highlight the issue of composability, i.e. being able to analyze applications in isolation while still reason about their overall behavior. We observe that arbitration plays an important role in this analysis. We compare two simple, yet commonly used arbitration mechanisms, and highlight the properties that are important for such analysis. We conclude that none of this arbitration mechanism is ideal for such an analysis and propose some variations to make them more suited for the analysis
Akash Kumar 0001, Bart Mesman, Henk Corporaal, Jef L. van Meerbergen, Yajun Ha
DSD1
2005 Efficient techniques for improved QoS performance in WDM optical burst switched networks
Gurusamy Mohan, Akash Kumar 0001, M. Ashish
Comput. Commun.2