Amit Acharyya

dblp:54/8180 · DBLP profile ↗
← Back
51ranked-venue papers
2as first author
30since 2021 · last 2025
0000-0002-5636-0676ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 2 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 A Novel Deep-Learning Method for Obstructive Sleep Apnea Detection from Single Channel Photoplethysmography
abstract
Early detection of obstructive sleep apnea (OSA) is extremely necessary to control its’ rising prevalence worldwide. Conventional diagnostic method like polysomnography (PSG) is uncomfortable, intrusive, and costly, thus, limiting its’ easy accessibility among people. To address this, we propose a novel deep learning method for detecting OSA using photoplethysmography (PPG) signal, a non-invasive method that is commonly available in wearable devices. We introduce a novel methodology using Multivariate Long Short-Term Memory-Fully Convolutional Network (MLSTM-FCN) model that effectively captures both temporal dependencies and local features in PPG signals for OSA detection. A new windowing technique was introduced to ensure apneic events are centered within each window to enhance the model’s ability to detect delayed physiological responses to apnea. The model was trained and evaluated on the Multi-Ethnic Study of Atherosclerosis (MESA) dataset, achieving an improvement of 11.3% in accuracy over the state-of-the-art method. The method obtained an accuracy of 93.44%, precision of 0.94, recall of 0.91, and an F1-score of 0.93. These results demonstrate the potential of our method in accurately identifying OSA events. This method offers a unobtrusive, comfortable, and cost-effective alternative to traditional diagnostic tools, making it suitable for long-term, home-based monitoring.
Prateek Agrawal, Rashmi Kumari, Pabitra Das, Surita Sarkar, Amit Acharyya
ISCAS5
2025 Energy-Efficient Reconfigurable Skyrmion-Based Counter for Nanoscale Applications
abstract
Counters are essential building blocks in digital systems, widely used for tasks such as tracking events, generating timing signals, and performing arithmetic operations like counting, frequency division, and sequencing. As the demand for low-power, high-performance computing grows, particularly in applications like IoT and edge devices, energy-efficient counter designs become increasingly crucial. Skyrmions have recently emerged as promising candidates for future logic device design due to their low energy, non-volatility, and stability. Their integration into information processing systems, such as racetrack memory and logic circuits, highlights their potential to overcome the limitations of conventional CMOS technology. This work introduces a reconfigurable skyrmion-based 4-bit counter for high-speed nanoscale computing applications. This architecture offers low energy, reliable performance, and enhanced scalability, making it an attractive solution for next-generation digital circuits and emerging nanoscale applications such as machine learning. The simulation results demonstrate the counter’s ability to reduce energy consumption to approximately 1.27 aJ per transition, which is nearly 99% lower than traditional CMOS-based counters.
C. Kishore, Santhosh Sivasubramani, Sarwath Sara, Arabinda Haldar, Chandrasekhar Murapaka, Rishad A. Shafik, Amit Acharyya
ISCAS7
2025 A Novel Methodology for Obstructive Sleep Apnea Detection from ECG using Deep Learning Approach
abstract
Obstructive sleep apnea (OSA) has become a serious health concern with increasing morbidity worldwide. Even though polysomnography is widely used by physicians for diagnosing OSA, the process is costly, time-consuming, and uncomfortable for patients. This increases the demand for developing unobtrusive, cost-effective, and reliable solutions for detecting OSA and reducing patient discomfort. Several machine learning-based and some deep learning based methods using extracted ECG features for OSA detection from ECG are found to be less reliable due to the manual feature extraction process, very few studies(included in the comparison table of below section) have used only deep learning methods for OSA detection from ECG signals. In this study, we proposed a novel deep learning method that leverages convolutional neural networks (CNN) and long short-term memory (LSTM) networks to learn spatial and temporal features for detecting OSA from ECG data. Our model was trained and evaluated on the publicly available MESA and Apnea-ECG datasets to assess its robustness. In a prudent data windowing step, we center the apnea data within each data window, enabling the model to better learn apnea patterns and resulting in achieving an accuracy of 95.4%, 92.2%, Specificity of 95.1%, 92.8%, Sensitivity of 94.8%, 91.6% and F1 score of 94.1%, 92.1% respectively on MESA and Apnea-ECG dataset. Our model yields an increase of 0.57% and 9.31% in specificity, 1.87% and 50.47% in sensitivity, 4.89% in F1-score, and 3.47% and 17.77% in accuracy against the best result we found using the Apnea-ECG and MESA dataset respectively. The results show our proposed model outperforms state-of-the-art methods on the MESA dataset and achieves equally good results on the Apnea-ECG dataset. Implementation of our model on Jetson orin AGX board gains comparable results with the above-stated accuracy showing the possible practical use of our methodology. These findings highlight the model’s stability, robustness, and high accuracy in detecting OSA.
Rashmi Kumari, Prateek Agrawal, Surita Sarkar, Pabitra Das, Amit Acharyya
ISCAS5
2025 Digital Twin Assisted Performance Aware Power Management for FPGA MPSoC using High Speed Reinforcement Learning
abstract
Performance-aware power manager development poses severe challenge in FPGA MPSoC. Conventional power managers are unaware of the underlying application’s performance requirements. The aspect of application performance awareness can be effectively handled by algorithms based on Reinforcement Learning (RL). The problem though is adaptability to new applications. To adjust to new application scenarios, RL-based algorithms need to be retrained but this causes deadline misses. This motivated us to propose a power manager based on RL assisted by a digital twin framework. The proposed architecture offloads the RL training process to the digital twin stage which emulates reward computation and state estimation. Performance-aware RL training aids in choosing the best frequency and core combination which uses less energy than the current approaches. The proposed methodology is implemented on the FPGA MPSoC ZCU104 with Mi Benchmarks and NATS Benchmarks. It reduces the FPGA energy consumption by 31.395% compared to the FPGA Built-in power manager. Retraining episodes on Deployed FPGA MPSoC are completely eliminated and the convergence time of RL for new workloads decreased to 86 seconds which is significantly less compared to state-of-the-art RL-based methods.
Kartik Laad, Ratnala Vinay, Parveen Nisha, Vidhumouli Hunsigida, Appa Rao Nali, Amit Acharyya
ISCAS6
2025 Flexible PVDF-ZnO Composite Sensor with High Sensitivity and Piezoelectric Response for Enhanced Acoustic Emission Detection
abstract
Flexible sensors are increasingly essential in applications requiring adaptability to complex surfaces and resilience under mechanical stress, yet traditional rigid sensors often lack these capabilities, limiting their effectiveness. This study addresses these challenges by enhancing the piezoelectric properties of PVDF by incorporating ZnO nanorods, creating a flexible composite sensor with superior sensitivity. A homogenous PVDF-ZnO composite was fabricated via an optimized casting method, achieving up to 97% beta phase for high piezoelectric response. FESEM, XRD, and FTIR characterization confirmed uniform ZnO dispersion and phase transition The sensor was later evaluated for its effectiveness in acoustic emission detection for nondestructive testing applications. A pencil lead break test demonstrated a substantial voltage output (23 mV), surpassing a conventional PZT sensor. These findings highlight the potential of PVDF-ZnO composites for high-sensitivity acoustic emission sensing applications.
Akshaya Muraleedharan, Swati Ghosh Acharyya, Amit Acharyya
ISCAS3
2025 REVBiT 2.0: REVerse Engineering of BiTstream for LUT Extraction, Boolean Logic, and Pin Combination Identification
abstract
Field-Programmable Gate Arrays (FPGAs) are extensively utilized in various fields due to their inherent flexibility and ability to be reconfigured. The functionality of digital designs within FPGAs is stored as configuration frames within the bitstream. Previous studies introduced tools like BIL, RapidSmith, Debit, DAT and BitFREE, which reverse-engineer the bitstream to reveal the Boolean logic of Look-Up Tables (LUTs) using the Xilinx ISE design suite. This tool produces both the bitstream and a textual representation of the placed design in the form of the Xilinx Design Language (XDL) file, enabling deeper insights into the FPGA’s internal structure. However, with the introduction of the AMD Xilinx Vivado design suite, support for XDL and text-based hardware adjustments discontinued, making it more challenging to reverse-engineer modern bitstreams. Our prior study, REVBiT, utilized the AMD Xilinx Vivado Design Suite, which extracted the LUTs and identified Boolean logic but failed to identify the pin combination in its present form. The pin combination of LUTs is also essential information for determining the correct configuration of LUTs in the bitstream. To address the limitation of state-of-the-art methods, we introduce REVBiT 2.0, a methodology for extracting LUTs and their Boolean logic, along with the pins connection of the LUT. Our proposed deep-learning models have been trained and tested with 92,82,950 data samples for the pin combination and Boolean logic identification. The experiment for bitstream extraction is carried out on a real FPGA board with the help of a logic analyzer. Our proposed methodology has been experimentally validated on AMD Xilinx 7-Series, Ultrascale, and Ultrascale+ FPGA device families for 2, 3, and 4-input LUTs using the AMD Xilinx Vivado design suite and relying on bitstream without any additional information. We achieved ≈ 100% accuracy for the LUT extraction, more than 87.50% prediction accuracy for pin combination and 92.43% for Boolean logic identification from the bitstream.
Anmol Singh Narwariya, Aniruddha Paradkar, Pabitra Das, Amit Acharyya
ISCAS4
2025 Digital Twin-Based Architecture for Run-Time Power Modeling for Sensorless Edge Devices
abstract
Run-time power and thermal management software are key to optimizing any embedded system to save energy. Power feedback is required to make an informed decision which can be taken either from power sensors or power models for power management. In most of the edge devices, there is a lack of dedicated power sensors as it is costly, and deploying them at scale across multiple devices can be difficult. Power models are the most used solution to estimate power and are trained on seen workloads. However, adaptability is a major concern if an unseen workload comes, resulting in the model’s accuracy drops increasing mean square error. This paper proposes a digital twin-based power model capable of handling unknown workloads without hampering the model’s accuracy removing the need for retraining on the edge. The proposed method has been proved on two power model techniques (i.e. Linear regression and Random forest) on Nvidia’s Jetson Nano platform which makes this generalized. There is a significant improvement in power estimation accuracy for new or unseen workloads. Specifically, when utilizing the linear regression model and the random forest model, the mean squared error (MSE) is reduced to 87% and 94% compared to the state-of-the-art method for the unseen workload respectively. Furthermore, the proposed architecture achieves an R2value of 0.99 and a mean average percentage error (MAPE) of 0.12% compared to state-of-the-art that has R2value of 0.92 and MAPE of 2.9% for unseen workloads and saving 20,061 2-input NAND gate equivalent area.
Parveen Nisha, Ratnala Vinay, Kartik Laad, Amit Acharyya
ISCAS4
2025 Simultaneous Series-Parallel High-Speed and Accurate Active Cell Balancing for Improved Battery Life
abstract
Active Cell Balancing based on DC-DC converter has become prominent due to its improved balancing accuracy, high energy conversion efficiency, and modularized approaches to balancing the charge in cells. It is widely employed in emerging applications such as electric vehicles, UAVs, and renewable grid storage to enhance battery life. The state-of-the-art active cell balancing techniques employ separate cell balancing for series and parallel connected cells. However, during separate balancing of cells in series and parallel configuration, the cell balancing system suffers from high system latency, low balancing speed, and higher power consumption due to a large number of balancing components. This article introduces a simultaneous series-parallel Flyback converter-based PWM duty cycle controlled Active Cell Balancing methodology with Constant Current-Constant Voltage (CC-CV) charging/discharging. The battery pack model was simulated with capacity and State-of-Charge (SOC) imbalances. The results show an improvement of 32.5% in balancing time and 81.47% reduction in power consumption without compromising the balancing accuracy when compared to balancing series and parallel cells separately in state-of-the-art active cell balancing methodologies. This helps to prolong battery lifetime and improve safety.
Arghadeep Sarkar, Rashi Dutt, Amit Acharyya
ISCAS3
2025 XVPE-Net: A Novel Methodology for Interpretable Vital Parameter and Cuffless Blood Pressure Estimation from PPG Signal
abstract
The interpretability of machine learning model is crucial in healthcare as it fosters reliability and supports clinicians in making informed decisions based on predictions. However, recent approaches for vital parameter estimation often act as black boxes, lacking clarity in their predictive reasoning. Therefore, accurate estimation of vital parameters is crucial not only for timely diagnosis and effective patient monitoring but also for ensuring that clinicians can trust and understand the model’s predictions. In this study, we introduce XVPE-Net, an explainable framework aimed at improving the state-of-the-art VPE-Net model for estimating vital parameters: heart rate (HR), respiratory rate (RR), systolic blood pressure (SBP), and diastolic blood pressure (DBP). Our model enhances interpretability by visually highlighting key parts of the PPG signal that significantly influence VPE-Net’s predictions of vital parameters. We test our model using 1,000 PPG segments from 38 subjects in the MIMIC-III dataset, which is available in the Physionet repository. The results show strong predictive capabilities, along with better transparency and reliability. This study helps us understand how the model makes decisions and supports the use of explainable AI in health monitoring, contributing to the development of more reliable models for medical applications.
Rahul Verma, Pabitra Das, Surita Sarkar, Prateek Agrawal, Rashmi Kumari, Amit Acharyya
ISCAS6
2025 Noninvasive Methodology for the Age Estimation of ICs Using Gaussian Process Regression
abstract
Age prediction for integrated circuits (ICs) is essential in establishing prevention and mitigation steps to avoid unexpected circuit failures in the field. Any electronic system would get benefit from an accurate age calculation. Additionally, it would assist in reducing the amount of electronic waste and the effort toward green computing. In this article, we propose a methodology to estimate the age of ICs using the Gaussian process regression (GPR). The output frequency of the ring oscillator (RO) is influenced by various factors, including the trackable path, voltage, temperature, and ageing. These dependencies are leveraged in the GPR model training. We demonstrate the RO’s frequency degradation by employing the Synopsys HSPICE tool with 32 nm predictive technology model (PTM) and the Synopsys technology library. We used temperature variation from 0 °C to 100 °C and voltage variation from 0.80 to 1.05 V for the data acquisition. Our methodology predicts age precisely; the minimum prediction accuracy with a month deviation on linear sampling rate is 85.36% for 13-Stage RO and 87.09% for 21-Stage RO, with a range of improvement in prediction accuracy compared to state-of-the-art (SOTA) is 9.74% to 16.99%. Similarly, on the logarithmic sampling rate, the prediction accuracy for 13-Stage RO and 21-Stage RO are 98.62% and 98.56%, respectively. The proposed methodology performs more accurately in terms of prediction accuracy and age prediction deviation from the SOTA methodology.
Anmol Singh Narwariya, Pabitra Das, S. Saqib Khursheed, Amit Acharyya
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 GRIPT: Graph Attention-Assisted Inductive Methodology for Fast and Accurate Average Power Estimation From RTL Simulation Skipping Gate-Level Simulation
abstract
This paper proposes GRIPT, a novel graph attention-based inductive methodology that enables a fast and accurate average power estimation of synthesized ASIC Design from RTL simulation, bypassing gate-level simulation. The proposed GRIPT methodology with features-aided attention mechanism-based inductive Graph Neural Network (GNN) model propagates the input wires’ toggle rates acquired from the RTL SAIF file through the circuit. We examine the proposed GRIPT methodology’s versatility by testing circuits across technology nodes and foundries like TSMC 65nm, 40nm, 90nm, 130nm and GF 40nm while training exclusively on the TSMC 65nm technology node. We introduce unseen and untrained logic cells to test the transferability of the proposed GRIPT model in toggle rate prediction across technology nodes. We evaluate the scalability of the proposed GRIPT by testing circuits of considerable size, like circuit trigonometry from OpenCores, and circuits arbiter, multiplier, div, log2, sin, square and sqrt from the EPFL benchmark suite, without training them. The proposed GRIPT surpasses the state-of-the-art GRANNITE model in predicting the unseen and untrained logic cell’s toggle rates to justify the proposed GRIPT’s transferability for designs across technology nodes to achieve an average improvement of 6.74%, 5.53%, 8.93%, 6.88% and 7.22%, respectively. To demonstrate versatility and scalability, the proposed GRIPT methodology outperforms the commercial RTL power estimation tool and GRANNITE in estimating the average power of circuits spanning TSMC 65nm, 40nm, 90nm, 130nm, and GF 40nm technology nodes, with a mean improvement of 23.42%, 16.59%, 16.35%, 27.16%, and 28.87%; 2.01%, 0.97%, 0.98%, 1.6% and 2.72%, respectively. The proposed GRIPT is 12.57X and 1.1X faster, with an average inference throughput (number of cycles inferred per second) of 1384.6 Hz compared to the commercial gate-level power estimation tool and GRANNITE, respectively.
Pabitra Das, Sai Pranav K. R, Amit Acharyya
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 Inductive GNN-Based Methodology for Accurate and Fast Average Power Estimation of Synthesized ASIC Designs From RTL Simulation Bypassing Gate-Level Simulation
abstract
This paper proposes an inductive Graph Neural Network (GNN) based methodology for accurate and fast average power estimation of logic-synthesized and RTL-simulated ASIC Design, eliminating gate-level simulation. With the novel variation of inductive GNN architecture, the proposed model propagates the input wires’ toggle rates acquired from RTL simulation through the synthesized design. We only train the proposed model on circuits synthesized from TSMC 65nm technology node but test on the circuits synthesized across TSMC 65nm, 40nm, 90nm, 130nm and GF 40nm technology nodes. We test the inductivity of the proposed model to predict the output wires’ toggle rates of unseen and untrained logic cells of the designs. We compute the proposed methodology’s average power inference throughput (number of cycles inferred per second) for speed comparison. The proposed model does better than state-of-the-art architecture GRANNITE to predict the unseen and untrained logic cell’s toggle rates of the designs across TSMC 65nm, 40nm, 90nm, 130nm, and GF 40nm technology nodes, showing an average improvement of 6.67%, 8.49%, 9.26%, 7.47% and 6.36%, respectively. The proposed methodology is more accurate than the commercial RTL average power estimation tool and GRANNITE in estimating the average power of circuits across TSMC 65nm, 40nm, 90nm, 130nm, and GF 40nm technology nodes by achieving a mean improvement of 24.94%, 16.77%, 17.65%, 29.72%, and 32.84%; 2.75%, 1.1%, 1.81%, 2.59% and 4.49%; respectively. The proposed methodology is 11.07X faster, with an average inference throughput of 1.218kHz, than the commercial gate-level average power estimation tool.
Pabitra Das, Sai Pranav K. R, Amit Acharyya
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Leveraging IO Pad Protection Diodes for Recycled IC Detection and Age Estimation Using Polynomial Regression
abstract
The presence of counterfeit recycled ICs (CRICs) in the global semiconductor supply chain is a major concern in the present-day world. These CRICs are less reliable and have become a serious threat to the ICs employed in safety–critical systems. Accurate age prediction for integrated circuits (ICs) is crucial for implementing preventative and mitigation strategies to avoid unexpected failures in the field. By precisely estimating the age of an IC, electronic systems can benefit from improved reliability and performance, as maintenance and replacements can be scheduled proactively, and reducing the risk of sudden breakdowns. Furthermore, accurate age prediction plays a vital role in extending the lifespan of electronic devices, which in turn helps to minimize electronic waste. This not only reduces the environmental impact but also supports the broader goal of green computing by promoting more sustainable and resource-efficient technology practices. In this article, we introduce a method for detecting a CRIC and estimating its age by utilizing the existing input-output (IO) pad structures targeting sensorless chips. The proposed methodology estimates age by measuring the voltage drop across the protection diodes present in the IO pad structure and applying this voltage drop to the proposed polynomial regression model. This methodology requires no additional sensory circuit, resulting in no area overhead. As there is no requirement for a special on-chip sensor, the proposed methodology can be used to detect the age of an IC in production. Our proposed polynomial regression model achieves a mean squared error (MSE) of 1.77 h, with a minimum improvement of 99.7% over the state-of-the-art methodologies.
Anmol Singh Narwariya, Srisubha Kalanadhabhatta, Amit Acharyya
IEEE Trans. Very Large Scale Integr. Syst.3
2024 ANN-based Accurate and Fast Post-Route QoR Data Prediction Methodology from Pre-Clock Tree Synthesis by Skipping CTS and Routing
abstract
In physical design, back-end flow, clock tree synthesis, and routing steps are vital in design optimization. Early and accurate QoR parameter prediction is paramount for design optimization and timely design delivery. In this paper, we propose an accurate and fast post-route QoR data prediction methodology by skipping CTS and Routing steps from the physical design back-end flow. This proposed methodology focuses on database generation and prediction of 26 QoR report parameters of post-Route QoR report. In this methodology, we are considering nine benchmark circuits from ISCAS89, IWLS 2005, and ISPD 2013. We use the random split method to split 50% of the data for training and the remaining 50% for testing purposes. We are validating our proposed method for two different technology nodes, i.e., TSMC 40nm and TSMC 65nm, which ensure robustness and reusability of the proposed methodology. Experimental results show that the average mean square error for all the parameters for both technologies is 10-2, while most of the parameter MSE is in the range of 10-4to 10-6for both technology nodes. It is also shown that the accuracy significantly improved (95% in TSMC 40nm and 98% in TSMC 65nm) and is comparable to the other work, along with the advantage of skipping CTS and Route both steps. Skipping of CTS and Routing optimizations and use of the artificial neural network (ANN) makes the proposed method a fast method as ANN is a less weighted network.
Pabitra Das, Amit Acharyya
ISCAS3
2024 Bayesian Inference Accelerator for Spiking Neural Networks
abstract
Bayesian neural networks offer better estimates of model uncertainty compared to frequentist networks. However, inference involving Bayesian models requires multiple instantiations or sampling of the network parameters, requiring significant computational resources. Compared to traditional deep learning networks, spiking neural networks (SNNs) have the potential to reduce computational area and power, thanks to their event-driven and spike-based computational framework. Most works in literature either address frequentist SNN models or non-spiking Bayesian neural networks. In this work, we demonstrate an optimization framework for developing and implementing efficient Bayesian SNNs in hardware by additionally restricting network weights to be binary-valued to further decrease power and area consumption. We demonstrate accuracies comparable to Bayesian binary networks with full-precision Bernoulli parameters, while requiring up to 25× less spikes than equivalent binary SNN implementations. We show the feasibility of the design by mapping it onto Zynq-7000, a lightweight SoC, and achieve a 6.5× improvement in GOPS/DSP while utilizing up to 30 times less power compared to the state-of-the-art.
Prabodh Katti, Anagha Nimbekar, Amit Acharyya, Bashir M. Al-Hashimi, Bipin Rajendran
ISCAS4
2024 P2E-LGAN: PPG to ECG Reconstruction Methodology using LSTM based Generative Adversarial Network
abstract
Cardiovascular diseases (CVDs) are the major cause of global morbidity and mortality. CVDs can be preliminarily diagnosed by analyzing a patient’s electrocardiogram (ECG), which requires long-term continuous ECG monitoring to suitably detect the onset of the disease. However, ECG data acquisition involves multiple lead attachments and requires regular intervention by an expert, thereby making the process cumbersome and inappropriate for continuous health monitoring due to limited portability and discomfort caused to the patients. Nowadays, automated ECG measurement techniques are gaining popularity in wearable health monitoring applications to seamlessly identify cardiac abnormalities even in a home environment. On the other hand, photoplethysmography (PPG) signals can be acquired from the wrist or fingertip of a patient by using a lead-less patch-less set-up that can be easily integrated with smart wearable devices. Therefore, to address the aforementioned demerits associated with ECG devices, a few researchers have fostered the idea of reconstructing ECG from photoplethysmogram (PPG) signals to generate simple yet effective CVD monitoring methodologies. Hence, in this paper, we propose P2E-LGAN, a hybrid generative adversarial network (GAN) based framework for generating ECG from PPG. The proposed network is evaluated on a benchmark database combined with ECG and PPG data. The inclusion of LSTM in the GAN network reduces the root mean square (RMSE), mean absolute error of heart rate (MAE(HR)) and percentage root mean square difference (PRD) by 35.7%, 37.2% and 9.8% respectively. Individual graphical analysis and performance evaluation of different metrics with state-of-the-art methods demonstrate the effectiveness of the proposed framework for reconstructing ECG from PPG.
Rashmi Kumari, Surita Sarkar, Debeshi Dutta, Pabitra Das, Amit Acharyya
ISCAS5
2024 REVBiT: REVerse Engineering of BiTstream for LUT Extraction & Logic Identification
abstract
Field-Programmable Gate Arrays (FPGAs) are widely used in various applications due to their flexibility and reconfigurability, and they store the functionality of digital design in the form of configuration frames within the bitstream. In the earlier studies, state-of-art methodologies, such as BIL and RapidSmith reverse engineer the bitstream to identify the boolean logic of LUTs using the Xilinx ISE tool, which provides bitstream and textual information of placed design in the form of a Xilinx Design Language (XDL) file. However, the more recent tool, Xilinx Vivado, does not include XDL support or text-based hardware adjustments. To resolve the above problem, here we introduce a methodology called REVBiT for LUT extraction and boolean logic identification that offers the potential to verify functionality against a trusted reference or rectify corrupted bitstream data by correcting it. Also, our propose methodology verified on AMD Xilinx 7-Series, Ultrascale and Ultrascale+ device families FPGAs using the Xilinx Vivado tool and does not rely on additional information besides the bitstream. We achieved 100% accuracy for the LUT extraction and 93.86%, 96.26%, and 95.16% accuracy for the boolean function identification for 7-Series, Ultrascale and Ultrascale+ device families, respectively.
Anmol Singh Narwariya, Chetan Talele, Pabitra Das, Amit Acharyya
ISCAS4
2023 GRASPE: Accurate Post-Synthesis Power Estimation from RTL using Graph Representation Learning
abstract
In this paper, we propose GRASPE, a graph representation learning-based methodology to accurately estimate post-synthesis average power consumption from the RTL to expedite the time to market in the ASIC design. Our proposed methodology uses novel graph neural network architecture (GNN) to work on unoptimized and unmapped post-translated netlist files. The GRASPE learns to propagate the average toggle rates with embedded feature values as vectors on each logic cell during training and then predicts the average toggle rates of a new design during testing. We attain a mean improvement of 19.84% and 4.42% in average toggle rates prediction, 14.12% and 2.67% in average power estimation over the commercial RTL power estimation tool and Graph Convolutional Network as GNN, respectively and 17.96X faster than the commercial gate-level power estimation tool. Subsequently, we evaluate GRASPE with the state-of-the-art GRANNITE for inference latency and average power estimation and demonstrate an average improvement of 3.985X and 1.28%, respectively.
Pabitra Das, Anant Terkar, Amit Acharyya
ISCAS4
2023 Selective Binarization based Architecture Design Methodology for Resource-constrained Computation of Deep Neural Networks
abstract
In this paper, we introduced a novel selective binarization based architecture design methodology for the compute-intensive deep neural networks (DNNs) implementation on memory-constrained platforms with insignificant compromise in accuracy. To demonstrate the advantages of our proposed architecture design methodology, we performed a detailed layer-wise performance analysis of a DNN with the help of metrics like memory savings and accuracies. Subsequently, we validated the proposed design methodology by implementing it on the resource-constrained AMD-Xilinx Kintex-7 field-programmable gate array (FPGA) chip for the considered DNN. A thorough analysis of our architecture design methodology results shows a significant memory savings of about 93% or compression of 15.7× with less than a 1% marginal reduction in accuracy.
Ramesh Reddy Chandrapu, Dubacharla Gyaneshwar, Sumohana S. Channappayya, Amit Acharyya
ISCAS4
2023 DeepAttack: A Deep Learning Based Oracle-less Attack on Logic Locking
abstract
Logic locking is one of the most promising design-for-trust technique for protecting intellectual property from reverse engineering, IP piracy, and modification throughout the electronic supply chain. However, oracle-less deobfuscation attacks that do not require an activated chip have been successful in obtaining the secret key of locked designs. This requires a detailed determination of the extent of vulnerability available in obfuscated circuitry. In this paper, we propose the oracle-less DeepAttack: an attack on logic locking that is capable of extracting the activation key of the locked netlist using a deep learning model. Based on the ISCAS-85 and EPFL benchmarks evaluation, DeepAttack achieves an average key prediction accuracy of 93.39%, outperforming the oracle-less state-of-the-art attacks SAIL, SnapShot, and OMLA by 21.28, 10.73, and 3.84 percentage points, respectively.
Anand Raj, Nikhitha Avula, Pabitra Das, Dominik Germek, Farhad Merchant, Amit Acharyya
ISCAS6
2023 Battery States Co-estimation Methodology Using Dual Square Root Unscented Kalman Filter
abstract
Real-time and accurate estimation of battery internal states is immensely critical for emerging applications such as Electric Vehicles (EV), smart grids, and space applications. Model-based state estimation methodology provides highly robust and accurate battery state estimation. However, separate estimation of states, such as State-of-Charge (SOC), State-of-Health (SOH), and State-of-Power (SOP), leads to erroneous estimation since the states are highly interdependent. A co-estimation methodology for SOC, SOH, and SOP using a highly accurate and stable formulation of the Kalman filter, i.e., the Dual Square Root Unscented Kalman filter (D-SRUKF) is proposed in this paper. The proposed battery states co-estimation methodology has been validated using experimental battery test data. The results show that SOC estimation error is 0.404 %, with an improvement of 77.60% compared to separate state estimation using the D-SRUKF estimator and 58.02% compared to state-of-the-art EKF-RLS co-estimation methodology. SOH and SOP are also co-estimated within the same filter, leading to accurate estimation without adding to the computational complexity of the system. The accuracy of SOH estimation is improved by 16.98% compared to the EKF-RLS co-estimation.
Souris Sahu, Rashi Dutt, Amit Acharyya
ISCAS3
2023 Graphene-based area efficient power planning architecture design methodology for nanomagnetic logic implementation
Santhosh Sivasubramani, Sanghamitra Debroy, Swati Ghosh Acharyya, Amit Acharyya
J. Supercomput.4
2022 Dual Square Root Unscented Kalman Filter based Single Channel Blind Source Separation Methodology
abstract
Single channel Blind Source Separation (SCBSS) is a challenging problem for several real-world practical applications. The existing SCBSS methodologies depend upon the properties of the sources present in the mixture and hence do not remain truly blind. Also, the solutions are found to be suboptimal and limited in application. In this paper, we present an SCBSS methodology using a state-parameter estimation approach to eliminate the constraints on the source signals such as statistical independence and frequency disjoint spectra. A Dual Square Root Unscented Kalman Filter (D-SRUKF) estimator has been proposed, which demonstrates higher numerical accuracy and improved stability compared to the widely used Dual Extended Kalman Filter (D-EKF). Simulations have been performed for separating mixed signals with overlapping spectra such as speech and biomedical signals. The proposed methodology demonstrates higher Signal-to-Interference Ratio (SIR) and Signal-to-Distortion Ratio (SDR) when the current methodologies even fail to separate the sources. The results also show that the proposed DSRUKF SCBSS is 15% more accurate than the state-of-the-art D-EKF SCBSS and has higher stability owing to the square root formulation of D-SRUKF source estimator.
Rashi Dutt, Amit Acharyya, Israr Sheikh
ISCAS2
2022 Artificial Neural Network Based Post-CTS QoR Report Prediction
abstract
In this paper, we propose two models to predict 26 parameters data of post clock tree synthesis (post-CTS) quality of results (QoR) report without running the CTS optimization step. In model 1, we considered 9 benchmark circuits (6 from ISCAS89 and 3 from open cores). We randomly split 50% of the total data into the training sample and the other 50% in the testing sample. In model 2, we use 6 benchmark circuit data for training purposes and use 3 benchmark circuit data for testing purposes which are unseen to the model. We utilize a regression neural network for predictions. To ensure robustness and reusability of the proposal, we validate our proposed models for two different technology nodes i.e. TSMC 65nm and TSMC 90nm. Experimental results show that the average mean square error for all the parameters for both the technologies is of the order of 10-3while most of the parameter MSE is in the range of 10-5to 10-7for both the technology nodes. These data ensure robustness and re-usability of the proposal with a high level of accuracy.
Pabitra Das, Amit Acharyya
ISCAS3
2022 An FPGA Based Energy-Efficient Read Mapper With Parallel Filtering and In-Situ Verification
abstract
In the assembly pipeline of Whole Genome Sequencing (WGS), read mapping is a widely used method to re-assemble the genome. It employs approximate string matching and dynamic programming-based algorithms on a large volume of data and associated structures, making it a computationally intensive process. Currently, the state-of-the-art data centers for genome sequencing incur substantial setup and energy costs for maintaining hardware, data storage and cooling systems. To enable low-cost genomics, we propose an energy-efficient architectural methodology for read mapping using a single system-on-chip (SoC) platform. The proposed methodology is based on the q-gram lemma and designed using a novel architecture for filtering and verification. The filtering algorithm is designed using a parallel sorted q-gram lemma based method for the first time, and it is complemented by an in-situ verification routine using parallel Myers bit-vector algorithm. We have implemented our design on the Zynq Ultrascale+ XCZU9EG MPSoC platform. It is then extensively validated using real genomic data to demonstrate up to 7.8× energy reduction and up to 13.3× less resource utilization when compared with the state-of-the-art software and hardware approaches.
Venkateshwarlu Y. Gudur, Sidharth Maheshwari, Amit Acharyya, Rishad A. Shafik
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 PLEDGER: Embedded Whole Genome Read Mapping using Algorithm-HW Co-design and Memory-aware Implementation
abstract
With over 6000 known genetic disorders, genomics is a key driver to transform the current generation of healthcare from reactive to personalized, predictive, preventive and participatory (P4) form. High throughput sequencing technologies produce large volumes of genomic data, making genome reassembly and analysis computationally expensive in terms of performance and energy. In this paper, we propose an algorithm-hardware co-design driven acceleration approach for enabling translational genomics. Core to our approach is a Pyopencl based tooL for gEnomic workloaDs tarGeting Embedded platforms (PLEDGER). PLEDGER is a scalable, portable and energy-efficient solution to genomics targeting low-cost embedded platforms. It is a read mapping tool to reassemble genome, which is a crucial prerequisite to genomics. Using bit-vectors and variable level optimisations, we propose a low-memory footprint, dynamic programming based filtration and verification kernel capable of accelerated parallel heterogeneous executions. We demonstrate, for the first time, mapping of real reads to whole human genome on a memory-restricted embedded platform using novel memory-aware preprocessed data structures. We compare the performance and accuracy of PLEDGER with state-of-the-art RazerS3, Hobbes3, CORAL and REPUTE on two systems: 1) Intel i7-8750H CPU + Nvidia GTX 1050 Ti, 2) Odroid N2 with 6 cores: 4xCortex-A73 + 2xCortex-A53 and Mali GPU. PLEDGER demonstrates persistent energy and accuracy advantages compared to state-of-the-art read mappers producing up to 11× speedups and 5.9× energy savings compared to state-of-the-art hardware resources.
Sidharth Maheshwari, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Venkateshwarlu Y. Gudur, Amit Acharyya
DATE6
2021 Single Channel Blind Source Separation Using Dual Extended Kalman Filter
abstract
Single channel Blind Source Separation (SCBSS) is an important source separation technique gaining prominence in many emerging applications. It is a special case of the well-defined Blind Source Separation (BSS) where only a single mixed signal is recorded to estimate the unknown sources. In this paper, we propose a simultaneous state-parameter estimation methodology for SCBSS using Dual Extended Kalman Filter (D-EKF). The proposed methodology eliminates the inherent frequency disjoint and statistical independence limitations of the state-of-the-art SCBSS approaches such as single channel Independent Component Analysis (SCICA). A frame- based Kalman processing technique has been proposed to ensure faster convergence of the proposed methodology. Simulation results have been presented for mixed sources with overlapping spectra and compared with SCICA and other BSS algorithms. The results demonstrate the superior performance of the proposed methodology with improved Signal-to-Interference Ratio (SIR) and Signal-to-Distortion Ratio (SDR) for real-world practical applications.
Rashi Dutt, Sayon Mondal, Amit Acharyya
ISCAS3
2021 IC Age Estimation Methodology Using IO Pad Protection Diodes for Prevention of Recycled ICs
abstract
Recycled ICs have become a major threat to the ICs used in safety critical systems. In the current state-of-the-art techniques, recycled ICs are detected by measuring the frequency, current, path delay or power-up values to estimate the HCI, BTI and EM effects on the transistors with age. Some of the state- of-the-art techniques require additional on-chip sensors to detect and estimate the age of an IC while others use existing logic like SRAM and Flip-flops to detect the recycled ICs. In this paper, we provide a methodology to detect a recycled IC and also to estimate its age by using the existing IO pad structures. For the first time, age is estimated by measuring voltage drop across the protection diodes present in IO pad structure. With this methodology, no additional sensors have to be added and hence there is no area overhead. With this proposed methodology, ICs that are used for a minimum period of a day can be effectively detected by using the concept of extended Kalman filtering technique for the first time in this domain. By stressing the part for five days, our proposed methodology can estimate the age of the IC aged between 1 month to 5 years with 95% percent of accuracy.
Srisubha Kalanadhabhatta, Rashi Dutt, S. Saqib Khursheed, Amit Acharyya
ISCAS4
2021 Control Strategy for Efficient Utilisation of Regenerative Power through Optimal Load Distribution in Hybrid Energy Storage System
abstract
With an increased focus on green transportation, hybrid energy storage systems (HESS) for Electric Vehicles (EV) are gaining importance in recent years. In this paper, we propose a control strategy for optimal distribution of power demand between battery and supercapacitor (SC) in a HESS, as well as efficient utilization of the regenerative braking power. A DC/DC boost converter based power splitting strategy has been proposed. Two control blocks are proposed for charging and discharging of energy storage, which reduces the rate of discharge current, thereby improving the health and life of the battery. Simulation results are presented for the proposed circuit and compared with the state-of-the-art topology. The results show that the proposed control strategy improves State-of-Charge (SOC) of the battery from 84% to 94.4% in comparison with only battery held systems. The depth of discharge (DOD) also improved by 10.4%, which helps to enhance the range of the HESS based EV from a single charge.
Souris Sahu, Rashi Dutt, Amit Acharyya
ISCAS3
2021 CORAL: Verification-Aware OpenCL Based Read Mapper for Heterogeneous Systems
abstract
Genomics has the potential to transform medicine from reactive to a personalized, predictive, preventive, and participatory (P4) form. Being a Big Data application with continuously increasing rate of data production, the computational costs of genomics have become a daunting challenge. Most modern computing systems are heterogeneous consisting of various combinations of computing resources, such as CPUs, GPUs, and FPGAs. They require platform-specific software and languages to program making their simultaneous operation challenging. Existing read mappers and analysis tools in the whole genome sequencing (WGS) pipeline do not scale for such heterogeneity. Additionally, the computational cost of mapping reads is high due to expensive dynamic programming based verification, where optimized implementations are already available. Thus, improvement in filtration techniques is needed to reduce verification overhead. To address the aforementioned limitations with regards to the mapping element of the WGS pipeline, we propose a Cross-platfOrm Read mApper using opencL (CORAL). CORAL is capable of executing on heterogeneous devices/platforms, simultaneously. It can reduce computational time by suitably distributing the workload without any additional programming effort. We showcase this on a quadcore Intel CPU along with two Nvidia GTX 590 GPUs, distributing the workload judiciously to achieve up to 2× speedup compared to when, only, the CPUs are used. To reduce the verification overhead, CORAL dynamically adapts k-mer length during filtration. We demonstrate competitive timings in comparison with other mappers using real and simulated reads. CORAL is available at: https://github.com/nclaes/CORAL.
Sidharth Maheshwari, Venkateshwarlu Y. Gudur, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Amit Acharyya
IEEE ACM Trans. Comput. Biol. Bioinform.6
2020 REPUTE: An OpenCL based Read Mapping Tool for Embedded Genomics
abstract
Genomics is transforming medicine from reactive to personalized, predictive, preventive and participatory (P4). The massive amount of data produced by genomics is a major challenge as it requires extensive computational capabilities, consuming large amounts of energy. A crucial prerequisite for computational genomics is genome assembly but the existing mapping tools used are predominantly software based, optimized for homogeneous high-performance systems. In this paper, we propose an OpenCL based REad maPper for heterogeneoUs sysTEms (REPUTE), which can use diverse and parallel compute and storage devices effectively. Core to this tool are dynamic programming based filtration and verification kernel to map the reads on multiple devices, concurrently. We show hardware/ software co-design and implementations of REPUTE across different platforms, and compare it with state-of-the-art mappers. We demonstrate the performance of mappers on two systems: 1) Intel CPU + 2×Nvidia GPUs; 2) HiKey970 embedded SoC with ARM Cortex-A73/A53 cores. The results show that REPUTE outperforms other read mappers in most cases producing up to 13× speedup with better or comparable accuracy. We also demonstrate that the embedded implementation can achieve up to 27× energy savings, enabling low-cost genomics.
Sidharth Maheshwari, Rishad A. Shafik, Ian Wilson 0006, Alexandre Yakovlev, Amit Acharyya
DATE5
2020 Real-Time and Accurate State-of-Charge Estimation Methodology using Dual Square Root Unscented Kalman Filter
abstract
Real-time and accurate estimation of battery states has gained immense importance in recent years due to emerging applications of Battery Energy Storage Systems (BESS) in smart grid and Electric Vehicles. The behaviour of BESS modelled as a 2-RC Circuit and State-of-Charge (SOC) and RC parameter estimation using Unscented Kalman Filter (UKF) has emerged as an optimal model for online Battery Management Systems (BMS). However, the stability of UKF degrades due to error covariance matrix becoming ill-conditioned. This paper presents a dual Square Root Unscented Kalman Filter (SRUKF) based SOC and parameter estimation algorithms for BMS. The proposed SRUKF methodology improves the stability of the system as the square root form of the error covariance matrix always remains positive semi-definite. The methodology has been designed and implemented in MATLAB/Simulink and compared with dual EKF and state-of-the-art dual UKF algorithms. The results show that dual SRUKF is 74% more accurate than the state-of-the-art and remains stable once it converges to true SOC value.
Rashi Dutt, Murali Chodisetti, Amit Acharyya
ISCAS3
2020 Accelerated Filtering and in situ Verification for Energy-Optimized Genome Read Mapping
abstract
Whole genome sequencing (WGS) includes sequencing and assembly pipelines to extract biological genomes for new advances in healthcare, agriculture and environmental research. It produces small random sections of the genome, called reads, and then re-assembled by mapping those reads to a reference genome. This process called read mapping produces a large volume of data, which are disparately processed by compute- and memory-intensive filtering and verification algorithms. As such, the problem of energy-frugal read mapping has remained an open challenge. In this paper, we propose an accelerated read mapping methodology with combined filtering and verification, implemented on an FPGA platform. Core to our methodology is an algorithm based on q-gram lemma for filtration with Myers bit-vector for verification in tandem. Through in situ verification, the proposed implementation optimizes resource utilization between filtration and verification and introduces parallel pipelines in computation and storage processes. Our experimental analysis shows that this methodology gives up to 8.7× energy efficiency when implemented on the Zynq Ultrascale+ FPGA platform, compared with the state-of-the-art software and hardware approaches.
Venkateshwarlu Y. Gudur, Sidharth Maheshwari, Rishad A. Shafik, Amit Acharyya
ISCAS4
2020 CardioNet: Deep Learning Framework for Prediction of CVD Risk Factors
abstract
The recent progressions in semiconductor and computing technology have empowered the PPG utilization in medical diagnosis. This paper presents a reconfigurable deep learning framework `CardioNet' for early diagnosis of cardiovascular risk factors or most common diseases (such as diabetes, hypertension, cerebrovascular, cerebra-infraction) using the PPG data. The proposed model has a light-weight architecture, designed by exploiting the deep learning framework of convolutional neural network, exhibiting inherent capability of feature extraction, thereby, eliminating the cost effective steps of feature selection and extraction. The performance demonstration of the proposed model is done on a healthy dataset comprising 657 data segments of 219 subjects holding records of common CVD risk factors (diabetes, hypertension, cerebrovascular, cerebra-infraction). The obtained results of an overall accuracy of 97% for diagnosis of CVD risk factors, show the efficiency of the proposed model for real-time usability. The clinical significance of this work to provide an accurate and non-invasive method for early diagnosis and monitoring of cardio-risk factors.
Madhuri Panwar, Arvind Gautam, Rashi Dutt, Amit Acharyya
ISCAS4
2020 Hardware-Software Codesign Based Accelerated and Reconfigurable Methodology for String Matching in Computational Bioinformatics Applications
abstract
Research for new technologies and methods in computational bioinformatics has resulted in many folds biological data generation. To cope with the ever increasing growth of biological data, there is a need for accelerated solutions in various domains of computational bioinformatics. In these domains, string matching is a most versatile operation performed at various stages of the computational pipeline. For search patterns that are updated with time, there is a need for accelerated and reconfigurable string matching to perform faster searching in the ever-growing biological databases. In this paper, we have proposed an accelerated and real-time reconfigurable methodology for string matching using hardware-software codesign. Using state of the art field programmable gate arrays, we have proposed a complete system-on-chip solution for applications that require accelerated as well as real-time reconfigurable string matching. The proposed methodology is the first of its kind novel approach for high-speed string matching that also supports quick reconfiguration by patterns changing with time. It is verified at the string matching stage of protein identification. Experimental results show that the architectures designed using our proposed methodology are 4X faster than state-of-the-art software implementation running on a workstation and 1.5X-4X faster than hardware accelerators available in the literature.
Venkateshwarlu Y. Gudur, Amit Acharyya
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 PUF-Based Secure Chaotic Random Number Generator Design Methodology
abstract
Pseudorandom number generators (PRNGs) play a pivotal role in generating key sequences of cryptographic protocols. Among different schemes, a simple chaotic PRNG (CPRNG) exhibits the property of being extremely sensitive to the initial seed and, hence, unpredictable. However, CPRNG is vulnerable if the initial seed is compromised. In this brief, we propose a novel physical unclonable function-based CPRNG (PUF-CPRNG), where the initial seed is secured by generating it from PUF. Furthermore, the proposed PUF-CPRNG includes dynamic refreshing logic to ensure that the random numbers generated are nonperiodic. To further secure the PUF-CPRNG, the feedback values of CPRNG are fed from PUF. An hardware architecture for the proposed methodology has been designed, and the proof of concept implementation was carried out using Xilinx Virtex-7 field-programmable gate array (FPGA). The proposed PUF-CPRNG passes the statistical test NIST 800-22, ENT, and correlation analysis.
Srisubha Kalanadhabhatta, Kiran Kumar Anumandla, S. Ashish Reddy, Amit Acharyya
IEEE Trans. Very Large Scale Integr. Syst.5
2019 A Framework for TSV Based 3D-IC to Analyze Aging and TSV Thermo-Mechanical Stress on Soft Errors
abstract
The CMOS aging, transient effects, and TSV thermomechanical stress degrade the resilience of 3D-ICs. The transients effects lead to soft errors and aggravated with the CMOS Bias temperature instability (BTI). In this paper, we analyze detrimental transient and BTI effect on soft error rate (SER) in 3D-ICs. However, TSV thermomechanical stress presents a considerable benefit by enhancing the critical charge (Qc) and reduce the SER due to decrease in the threshold voltage and increase in mobility of carriers in transistor present out of keep-out-zone and useful range. Therefore we propose a framework to evaluate the effect of transient, BTI, and TSV thermomechanical stress on critical charge and SER in 3D-ICs. Subsequently, through HSPICE simulation we show that for a lifetime of ten years and on the topmost layer of stacked 3D-IC, the reduction in SER of NAND gate by 5.12% - 9.05% and in 6T SRAM 2.51% - 4.76% and 3.77% - 5.64% decrease for storing 0 and 1 respectively.
Raviteja P. Reddy, Amit Acharyya, S. Saqib Khursheed
ITC-Asia2
2019 Simplex FastICA: An Accelerated and Low Complex Architecture Design Methodology for $n$ D FastICA
abstract
This paper proposes an n-dimensional Simplex FastICA (FICA), an accelerated and low complex architectural design methodology for FICA to attain high computation speed targeted for resource-constrained applications. This is achieved by exploiting the algorithmic redundancies of FICA that make the proposed Simplex FICA faster than the conventional FICA without adding any extra architectural complexities. The proposed methodology has been verified and validated by applying it for separating 6D EEG signals. Subsequently, the corresponding hardware has been designed using Verilog HDL. It is synthesized using UMC 90-nm technology resulting in 0.5703-mW power and 0.495 mm2 area at 1.08 V. The computation time saving for 4D-12D FICA was computed by varying the number of iterations of convergence for the nth stage from 2 to 4 for 4096 samples. The average percentage of computation time saved for nth stage achieved with the proposed methodology in comparison to the state of the art varies from 99.783% to 99.891%. The overall average percentage of computation time saving for the proposed design with the aforementioned specifications varies from 9.14% to 16.6%.
Swati Bhardwaj 0001, Shashank Raghuraman, Amit Acharyya
IEEE Trans. Very Large Scale Integr. Syst.3
2018 BiometricNet: Deep Learning based Biometric Identification using Wrist-Worn PPG
abstract
Rapid advances in semiconductor fabrication technology have enabled the proliferation of miniaturized body-worn sensors capable of long term pervasive biomedical signal monitoring. In this paper, we present a novel deep learning-based framework (BiometricNET) on biometric identification using data collected from wrist-worn Photoplethysmography (PPG) signals in ambulatory environments. We have formulated a completely personalized data-driven approach, using a four-layer deep neural network - employing two convolution neural network (CNN) layers in conjunction with two long short-term memory (LSTM) layers, followed by a dense output layer for modelling the temporal sequence inherent within the pulsatile signal representative of cardiac activity. The proposed network configuration was evaluated on the TROIKA dataset collected from 12 subjects involved in physical activity, achieved an average five-fold cross-validation accuracy of 96%.
Luke R. Everson, Dwaipayan Biswas, Madhuri Panwar, Dimitrios Rodopoulos, Amit Acharyya, Chris H. Kim, Chris Van Hoof, Mario Konijnenburg, Nick Van Helleputte
ISCAS5
2018 Modified Huffman based compression methodology for Deep Neural Network Implementation on Resource Constrained Mobile Platforms
abstract
Modern Deep Neural Network (DNN) architectures produce high accuracy across applications, however incur high computational complexity and memory requirements, making it challenging for execution on resource constrained mobile platforms. Driven by application requirements, there has been a shift in execution paradigm of Deep Nets from cloud based computation to sensor/mobile platforms. The limited memory available onboard a mobile platform, necessitates an effective mechanism for storage of network parameters (viz. weights) generated offline post-training. Hence, we propose a modified Huffman encoding-decoding technique, with dynamic usage of net layers, executed on-the-fly in parallel, which can be applied on a memory constrained multicore environment. To the best of our knowledge, this is the first study on applying compression based on multiple bit pattern sequences, to achieve a maximum compression rate of 64 percent and a single module decompression time of about 0.33 seconds without trading-off accuracy.
Chandrajit Pal, Sunil Pankaj, Wasim Akram, Amit Acharyya, Dwaipayan Biswas
ISCAS4
2018 Runtime Performance and Power Optimization of Parallel Disparity Estimation on Many-Core Platforms
abstract
This article investigates the use of many-core systems to execute the disparity estimation algorithm, used in stereo vision applications, as these systems can provide flexibility between performance scaling and power consumption. We present a learning-based runtime management approach that achieves a required performance threshold while minimizing power consumption through dynamic control of frequency and core allocation. Experimental results are obtained from a 61-core Intel Xeon Phi platform for the aforementioned investigation. The same performance can be achieved with an average reduction in power consumption of 27.8% and increased energy efficiency by 30.04% when compared to Dynamic Voltage and Frequency Scaling control alone without runtime management.
Charles Leech, Charan Kumar Vala, Amit Acharyya, Sheng Yang 0003, Geoff V. Merrett, Bashir M. Al-Hashimi
ACM Trans. Embed. Comput. Syst.3
2017 Coordinate Rotation-Based Low Complexity K-Means Clustering Architecture
abstract
In this brief, we propose a low-complexity architectural implementation of the K-means-based clustering algorithm used widely in mobile health monitoring applications for unsupervised and supervised learning. The iterative nature of the algorithm computing the distance of each data point from a respective centroid for a successful cluster formation until convergence presents a significant challenge to map it onto a low-power architecture. This has been addressed by the use of a 2-D Coordinate Rotation Digital Computer-based low-complexity engine for computing the n-dimensional Euclidean distance involved during clustering. The proposed clustering engine was synthesized using the TSMC 130-nm technology library, and a place and route was performed following which the core area and power were estimated as 0.36 mm2and 9.21 mW at 100 MHz, respectively, making the design applicable for low-power real-time operations within a sensor node.
Bhagyaraja Adapa, Dwaipayan Biswas, Swati Bhardwaj 0001, Shashank Raghuraman, Amit Acharyya, Koushik Maharatna
IEEE Trans. Very Large Scale Integr. Syst.5
2017 Low-Complexity Methodology for Complex Square-Root Computation
abstract
In this brief, we propose a low-complexity methodology to compute a complex square root using only a circular coordinate rotation digital computer (CORDIC) as opposed to the state-of-the-art techniques that need both circular as well as hyperbolic CORDICs. Subsequently, an architecture has been designed based on the proposed methodology and implemented on the ASIC platform using the UMC 180-nm Technology node with 1.0 V at 5 MHz. Field programmable gate array (FPGA) prototyping using Xilinx' Virtex-6 (XC6v1x240t) has also been carried out. After thorough theoretical analysis and experimental validations, it can be inferred that the proposed methodology reduces 21.15% slice look up tables (on FPGA platform) and saves 20.25% silicon area overhead and decreases 19% power consumption (on ASIC platform) when compared with the state-of-the-art method without compromising the computational speed, throughput, and accuracy.
Suresh Mopuri, Amit Acharyya
IEEE Trans. Very Large Scale Integr. Syst.2
2017 A Cost-Effective Fault Tolerance Technique for Functional TSV in 3-D ICs
abstract
Regular and redundant through-silicon via (TSV) interconnects are used in fault tolerance techniques of 3-D IC. However, the fabrication process of TSVs results in defects that reduce the yield and reliability of TSVs. On the other hand, each TSV is associated with a significant amount of on-chip area overhead. Therefore, unlike the state-of-the-art fault tolerance architectures, here we propose the time division multiplexing access (TDMA)-based fault tolerance technique without using any redundant TSVs, which reduces the area overhead and enhances the yield. In the proposed technique, by means of TDMA, we reroute the signal through defect-free TSV. Subsequently, an architecture based on the proposed technique has been designed, evaluated, and validated on logic-on-logic 3-D IWLS'05 benchmark circuits using 130-nm technology node. The proposed technique is found to reduce the area overhead by 28.70%-40.60%, compared to the state-of-the-art architectures and results in a yield of 98.9%-99.8%.
Raviteja P. Reddy, Amit Acharyya, S. Saqib Khursheed
IEEE Trans. Very Large Scale Integr. Syst.2
2015 An accurate clustering algorithm for fast protein-profiling using SCICA on MALDI-TOF
abstract
In this paper we propose an accurate clustering algorithm as the necessary step of the Single Channel Independent Component Analysis (SCICA) in the context of the fast extraction of protein profiles from the mass spectra (MALDI-TOF) data. In general K-means clustering is employed for clustering of the basis vectors. However given its iterative and statistical nature, convergence to the same clusters for the same data sets is not always guaranteed making it inaccurate, especially in protein-profiling where reliability of the bio-marker based disease detection and diagnosis depend immensely on the reliability of the clustering algorithm. Furthermore the proposed algorithm does not involve any arithmetic computations helping expedite the entire SCICA process.
Amit Acharyya, Mavuduru Neehar, Ganesh R. Naik
ISCAS1
2014 A new VLSI IC design automation methodology with reduced NRE costs and time-to-market using the NPN class Representation and functional symmetry
abstract
In the VLSI IC design, the number of incremental and iterative steps in the design automation methodology will decide the non-recurring-engineering (NRE) costs and time-to-market (TTM). Since these are the major driving factors of the IC design, many algorithms were proposed in the last few decades to minimize/optimize the number of design steps in the conventional VLSI IC Design methodology. However the frontend and backend designs have to be carried separately, which has limited the further minimization of the number of design steps. Here we propose a new unconventional design automation methodology, which reduces the NRE costs and TTM by merging the frontend and backend designs partially. It maps the input RTL description directly to their corresponding physical designs (derived using the existing CAD tools and stored in a pre-computed library) without any limitation on the Boolean function's input size. We have exploited the functional symmetry and negationpermutation- negation (NPN) class representations to decoct the library size and number of comparisons. The functional symmetry reduced the number of required pre-computed circuits in our experiments from 1031 to 222 (464.4% reduction in the memory size) and helps in maintaining the regularity in the design, which is a major concern for engineering change order.
Basireddy Karunakar Reddy, Srinivas Sabbavarapu, Amit Acharyya
ISCAS3
2014 A Reconfigurable High Speed Architecture Design for Discrete Hilbert Transform
abstract
This letter proposes a high-speed and reconfigurable Discrete Hilbert Transform architecture design methodology targeting the real-time applications including Cyber-Physical systems, Internet of Things or Remote Health-Monitoring where the same chip-set needs to be used for various purposes under real-time scenario. By using this architecture we are able to get Discrete Hilbert Transform for any given M-point by re-using N-point Discrete Hilbert Transform as a kernel. Here N and M are multiple of 4 and N respectively. Subsequently we provide the architecture design details and compare the proposed architecture with the conventional state-of-the-art architecture. Thorough theoretical analysis and experimental comparison results show that the proposed design is twice as fast and reconfigurability is also achieved simultaneously.
P. Sreenivasa Reddy, Suresh Mopuri, Amit Acharyya
IEEE Signal Process. Lett.3
2014 Development of an Automated Updated Selvester QRS Scoring System Using SWT-Based QRS Fractionation Detection and Classification
abstract
The Selvester score is an effective means for estimating the extent of myocardial scar in a patient from low-cost ECG recordings. Automation of such a system is deemed to help implementing low-cost high-volume screening mechanisms of scar in the primary care. This paper describes, for the first time to the best of our knowledge, an automated implementation of the updated Selvester scoring system for that purpose, where fractionated QRS morphologies and patterns are identified and classified using a novel stationary wavelet transform (SWT)-based fractionation detection algorithm. This stage informs the two principal steps of the updated Selvester scoring scheme--the confounder classification and the point awarding rules. The complete system is validated on 51 ECG records of patients detected with ischemic heart disease. Validation has been carried out using manually detected confounder classes and computation of the actual score by expert cardiologists as the ground truth. Our results show that as a stand-alone system it is able to classify different confounders with 94.1% accuracy whereas it exhibits 94% accuracy in computing the actual score. When coupled with our previously proposed automated ECG delineation algorithm, that provides the input ECG parameters, the overall system shows 90% accuracy in confounder classification and 92% accuracy in computing the actual score and thereby showing comparable performance to the stand-alone system proposed here, with the added advantage of complete automated analysis without any human intervention.
Valentina Bono, Evangelos B. Mazomenos, Taihai Chen, James Rosengarten, Amit Acharyya, Koushik Maharatna, John M. Morgan, Nick Curzen
IEEE J. Biomed. Health Informatics5
2013 Accurate and reliable 3-lead to 12-lead ECG reconstruction methodology for remote health monitoring applications
abstract
Standard 12-lead (S12) system and Mason-Likar 12-lead (ML12) system despite of being most acceptable systems for clinical usage are not the preferred lead systems for remote monitoring (RM) applications. Usually RM applications involve wireless transmission of signals and a 2-3 lead system is preferred for bandwidth and storage limitations and data transmission time. Generally, ECG compression techniques are applied for the same, however, compression ratio (CR) depends on the number of channels and decreases with the increase in number of channels. Thus, it facilitates the usage of a 2-3 lead system. However, a reduced lead (RL) system with 2-3 leads may be inadequate for the information desired by the cardiologists who are accustomed to S12 or ML12 system pertaining to its decades old usage. In this paper, we attempt to provide solution to both technical and non-technical limitations of RM applications. We reconstruct S12 and ML12 systems from Reduced 3-lead (R3L) system comprising of basis leads I, II, V2using personalized or patient-specific transformation. Two separate investigations have been carried out for S12 and ML12 with their corresponding R3L systems comprising of their respective basis leads. PhysioNet PTBDB and INCARTDB after wavelet based preprocessing were used in this investigation. R2statistics, correlation (rx) and regression (bx) coefficients were used to evaluate reconstructed signal against the original signal and the mean values obtained were 96.53%, 0.982 and 0.968 (S12) and 96.53%, 0.982 and 0.968 (ML12) respectively. R3L system reduces number of leads and electrodes from 12 and 10 to 3 and 5 respectively, lowers bandwidth and storage requirements, data transmission time and increases CR. The study shows that basis leads obtained from S12 outperforms the basis leads of ML12 for reconstruction of precordial leads.
Sidharth Maheshwari, Amit Acharyya, Pachamuthu Rajalakshmi, Paolo Emilio Puddu, Michele Schiariti
Healthcom2
2013 A Low-Complexity ECG Feature Extraction Algorithm for Mobile Healthcare Applications
abstract
This paper introduces a low-complexity algorithm for the extraction of the fiducial points from the Electrocardiogram (ECG). The application area we consider is that of remote cardiovascular monitoring, where continuous sensing and processing takes place in low-power, computationally constrained devices, thus the power consumption and complexity of the processing algorithms should remain at a minimum level. Under this context, we choose to employ the Discrete Wavelet Transform (DWT) with the Haar function being the mother wavelet, as our principal analysis method. From the modulus-maxima analysis on the DWT coefficients, an approximation of the ECG fiducial points is extracted. These initial findings are complimented with a refinement stage, based on the time-domain morphological properties of the ECG, which alleviates the decreased temporal resolution of the DWT. The resulting algorithm is a hybrid scheme of time and frequency domain signal processing. Feature extraction results from 27 ECG signals from QTDB, were tested against manual annotations and used to compare our approach against the state-of-the art ECG delineators. In addition, 450 signals from the 15-lead PTBDB are used to evaluate the obtained performance against the CSE tolerance limits. Our findings indicate that all but one CSE limits are satisfied. This level of performance combined with a complexity analysis, where the upper bound of the proposed algorithm, in terms of arithmetic operations, is calculated as 2:423N + 214 additions and 1:093N + 12 multiplications for N 861 or 2:553N + 102 additions and 1:093N +10 multiplications for N > 861 (N being the number of input samples), reveals that the proposed method achieves an ideal trade-off between computational complexity and performance, a key requirement in remote CVD monitoring systems.
Evangelos B. Mazomenos, Dwaipayan Biswas, Amit Acharyya, Taihai Chen, Koushik Maharatna, James Rosengarten, John M. Morgan, Nick Curzen
IEEE J. Biomed. Health Informatics3
2011 Simplified logic design methodology for fuzzy membership function based robust detection of maternal modulus maxima location: A low complexity Fetal ECG extraction architecture for mobile health monitoring systems
abstract
This paper proposes a simplified logic design methodology for the fuzzy membership function used for robust and reliable detection of modulus-maxima locations in wavelet domain for fetal ECG extraction from the abdominal composite ECG signal. This simplification is achieved by exploiting the inherent time-position information of the wavelet coefficients decomposed at different resolution levels. Subsequently, a low complexity VLSI architecture for Fetal ECG extraction is presented which is designed using the recently proposed memory- efficient, multiplierless Discrete Wavelet Transform method. The generic memory model within this architecture will provide the flexibility to configure the on-chip memory with any type of orthonormal wavelets suitable for different applications. Total synthesized cell area of the proposed architecture is 14.2 mm2and power consumption is 101.5 μW at 1.2 V @ 1 MHz frequency using 0.13 μm standard cell technology. The proposed architecture is targeted for the personalized health monitoring applications within a mobile home-care medical device in the resource constrained environment.
Amit Acharyya, Koushik Maharatna, Bashir M. Al-Hashimi, Hasitha Tudugalle
ISCAS1