Darshika G. Perera

dblp:64/4273 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0001-9106-4381ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2025 FPGA-Based Hardware Architecture for Sequence Alignment by Genetic Algorithm
abstract
With the advent of next-gen DNA sequencing technologies, there has been a massive growth in sequencing data and demand for analysis. The sequence of macromolecules such as DNA, RNA, and proteins are fundamental to the study of biology and medicine. Among many multiple sequence alignment techniques, Sequence Alignment by Genetic Algorithm (SAGA) generates high-quality alignments, but is computationally intensive, leading to low performance. In this paper, we propose an FPGA-based hardware architecture for SAGA to address the complexity and performance issues of SAGA.
Laura H. Garcia, Arkan Alkamil, Mokhles A. Mohsin, Johannes Menzel, Darshika G. Perera
ISCAS5
2023 Edge Computing-based Adaptive Machine Learning Model for Dynamic IoT Environment
abstract
With the advent of IoT and smart systems, edge computing coupled with machine learning (ML) techniques are becoming imperative to locally process and analyze the heterogeneous data generated from various IoT devices in real-time. The most common problem in dynamic IoT environment is performance degradation, mainly due to virtual concept drift (VCD). The issue of VCD often incurs, in dynamic IoT environment, when statistical properties of the input features change over time and make existing ML models obsolete or degrade models' performance and efficiency. Thus, the need for adaptive ML models. To facilitate this endeavor, our main objective is to create an adaptive ML model for resource-constrained edge computing devices to address the VCD issues in dynamic IoT environment. In this paper, we present a proof-of-concept adaptive ML model for edge computing based on a CNN single classifier with a real-time transfer learning method using fine-tuning. We also present our problem formulation and preliminary experimental results for problem validation.
Darshika G. Perera
ISCAS2
2023 Optimizing Density-Based Ant Colony Stream Clustering Using FPGA-Based Hardware Accelerator
abstract
In the era of IoT, a massive amount of data will be generated from various sensors and corresponding IoT devices. Density-based Ant Colony Stream Clustering (ACSC) is one of the best solutions for big data processing for real-world applications, due to its many inherent traits. Also, FPGAs are one of the best avenues to support/accelerate complex algorithms, such as ACSC. In this paper, we introduce an FPGA-based hardware accelerator for ACSC, which achieves maximum speedups of 603 and 2.5 vs. its software counterparts on embedded processor and PC, respectively, without compromising cluster accuracy. No similar FPGA-based ACSC hardware accelerator exists in the literature.
Jeremy R. Graf, Darshika G. Perera
ISCAS2
2023 A Systolic Array Architecture for SVM Classifier for Machine Learning on Embedded Devices
abstract
With the proliferation of embedded computing, many machine learning (ML) applications have found their way into embedded devices. The SVM classifier is deemed suitable for real-world ML applications, which often comprise large-scale multi-dimensional data. In this paper, we introduce an FPGA-based systolic array architecture to support and accelerate SVM classifier on resource-constrained embedded devices. We also introduce a unique system-level architecture to further enhance speedup and to facilitate real-time processing. Our systolic array achieves a maximum speedup of 107x compared to its software counterparts, and a maximum classification accuracy of 98.5%.
Srikanth Ramadurgam, Darshika G. Perera
ISCAS2
2020 A Quad Joint Relational Feature for 3D Skeletal Action Recognition with Circular CNNs
abstract
To deal with the limitations of human action recognition systems that apply deep neural networks (DNNs) to 3D skeletal feature maps, we propose an improved set of features that enable better pattern discrimination when using a spectrally enriched circular convolutional neural network (CCNN). These new features exploit the local relationships between joint movements based on 3D quadrilaterals constructed for all possible sets of four joints. Next, we compute the volumes of these time-varying quadrilaterals, by generating color-coded images, named spatio-temporal quad-joint relative volume feature maps (QjRVMs). To preserve the pixel frequency distribution while training a DNN, which is otherwise lost due to vanishing gradients and random dropouts, we propose a new architecture CCNNs. CCNNs use cyclic multi-resolution filters in a four-stream architecture, requiring only batch normalization and ReLU operations to identify multiple pixel pattern variations simultaneously. Applying the proposed CCNN to QjRVM images illustrates that combining multi-resolution features enhances the overall classification accuracy. Finally, we evaluate our proposed human action framework using our own 102-class, 5-subject action dataset, created using 3D motion capture technology, named KLHA3D-102. We also evaluate our framework using 3 publicly available datasets: CMU, HDM05, and NTU RGB-D.
P. V. V. Kishore, Darshika G. Perera, Maddala Teja Kiran Kumar, Dande Anil Kumar, Kiran Kumar Eepuri
ISCAS2
2020 An Optimized FPGA-Based Hardware Accelerator for Physics-Based EKF for Battery Cell Management
abstract
Battery technology is the cornerstone of HEVs. Model Predictive Control coupled with physics-based model (PBM) is an effective technique for battery management systems (BMS). In this case, extended Kalman filter (EKF) is used as the state observer for the highly non-linear PBM. Thus far, the sheer computational complexity of PBM hinders it from being utilized for portable BMS. In this paper, we introduce a novel and efficient FPGA-based hardware accelerator for physics-based EKF, to address the computational complexity of PBM.
Anne K. Madsen, Michael Scott Trimboli, Darshika G. Perera
ISCAS3
2017 An efficient embedded multi-ported memory architecture for next-generation FPGAs
abstract
In recent years, there has been a dramatic increase in utilization of FPGAs to enhance the speed-performance of many real-time compute and data intensive applications on embedded platforms. FPGA-based designs leverage parallelism in computations to achieve high speed-performance. Parallel computations require multi-ported memories to provide any number of ports for simultaneous multiple read/write (R/W) operations. Although several multi-ported memories are proposed in the literature, these designs become complex due to the extra logic and routing used for techniques/architectures to provide an arbitrary number of R/W ports. In this research work, we introduce a novel and efficient multi-ported memory architecture utilizing simple dual-port BRAMs, to provide an arbitrary number of R/W ports. Apart from the BRAMs, our proposed multi-ported memory design only consists of the Decision Making Modules and a counter, thus simplifying the design process. The R/W operations within our architecture are also straightforward. Experiments are performed to evaluate the feasibility and efficiency of our multi-ported memory architecture. We also evaluate our architecture with the most recently proposed multi-ported memory designs, implemented using LVT and XOR techniques, from the existing literature. FPGA manufacturers could employ our multi-ported memory architecture to accelerate real-time compute/data intensive applications with their next-generation FPGAs. Due to lower design complexity compared to the existing designs, our simplified memory architecture would enable seamless integration to the existing FPGA-based CAD tools with minimal design cost.
S. Navid Shahrouzi, Darshika G. Perera
ASAP2
2009 Similarity Computation Using Reconfigurable Embedded Hardware
abstract
Advances in portable devices and location-aware applications have necessitated the research in sophisticated yet small-footprint hardware and software in embedded systems, while the proliferation of the Web and distributed database systems has led to new data mining applications. We are investigating the utilization of reconfigurable hardware, due to its flexibility and performance, for data mining applications in portable and embedded computing. In this work, we introduce a reconfigurable hardware solution using field programmable gate array (FPGA) for similarity matrix computation, a commonly used data structure to represent the computed similarity among a set of feature vectors. Our hardware design can be dynamically reconfigured to accommodate three different similarity measures. A space-time cost analysis of the proposed multiplexer-based approach is presented. Experiments performed on the implemented reconfigurable hardware show encouraging and promising results that warrant further investigation in dynamically reconfigurable FPGA-based hardware for data mining applications.
Darshika G. Perera, Kin Fun Li
DASC1
2008 Parallel Computation of Similarity Measures Using an FPGA-Based Processor Array
abstract
An enormous amount of data needs to be processed in many data mining applications. In addition to algorithmic development, hardware support is imperative to improve the effectiveness and efficiency of these applications. We are investigating various hardware architectural design techniques and methodologies to support data mining at the chip level. In this work, we focus on the design of an FPGA-based processor array for the computation of similarity matrix, a commonly used data structure to represent the similarity among a set of feature vectors, with each matrix element representing the computed similarity measure between two vectors. An algorithm is developed to assign computation efficiently to the array of processing elements. Theoretical performance metrics are derived and compared to the experimental results. Performance gains using the processor array over software implementations are also presented and discussed.
Darshika G. Perera, Kin Fun Li
AINA1