EDBT 2026 Demo / reviewers in the wild / expert
Darshika G. Perera
dblp:64/4273
· DBLP profile ↗
9ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0001-9106-4381ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FPGA-Based Hardware Architecture for Sequence Alignment by Genetic AlgorithmabstractWith the advent of next-gen DNA sequencing technologies, there has been a massive growth in sequencing data and demand for analysis. The sequence of macromolecules such as DNA, RNA, and proteins are fundamental to the study of biology and medicine. Among many multiple sequence alignment techniques, Sequence Alignment by Genetic Algorithm (SAGA) generates high-quality alignments, but is computationally intensive, leading to low performance. In this paper, we propose an FPGA-based hardware architecture for SAGA to address the complexity and performance issues of SAGA. Laura H. Garcia, Arkan Alkamil, Mokhles A. Mohsin, Johannes Menzel, Darshika G. Perera |
ISCAS | 5 |
| 2023 | Edge Computing-based Adaptive Machine Learning Model for Dynamic IoT EnvironmentabstractWith the advent of IoT and smart systems, edge computing coupled with machine learning (ML) techniques are becoming imperative to locally process and analyze the heterogeneous data generated from various IoT devices in real-time. The most common problem in dynamic IoT environment is performance degradation, mainly due to virtual concept drift (VCD). The issue of VCD often incurs, in dynamic IoT environment, when statistical properties of the input features change over time and make existing ML models obsolete or degrade models' performance and efficiency. Thus, the need for adaptive ML models. To facilitate this endeavor, our main objective is to create an adaptive ML model for resource-constrained edge computing devices to address the VCD issues in dynamic IoT environment. In this paper, we present a proof-of-concept adaptive ML model for edge computing based on a CNN single classifier with a real-time transfer learning method using fine-tuning. We also present our problem formulation and preliminary experimental results for problem validation. Darshika G. Perera |
ISCAS | 2 |
| 2023 | Optimizing Density-Based Ant Colony Stream Clustering Using FPGA-Based Hardware AcceleratorabstractIn the era of IoT, a massive amount of data will be generated from various sensors and corresponding IoT devices. Density-based Ant Colony Stream Clustering (ACSC) is one of the best solutions for big data processing for real-world applications, due to its many inherent traits. Also, FPGAs are one of the best avenues to support/accelerate complex algorithms, such as ACSC. In this paper, we introduce an FPGA-based hardware accelerator for ACSC, which achieves maximum speedups of 603 and 2.5 vs. its software counterparts on embedded processor and PC, respectively, without compromising cluster accuracy. No similar FPGA-based ACSC hardware accelerator exists in the literature. Jeremy R. Graf, Darshika G. Perera |
ISCAS | 2 |
| 2023 | A Systolic Array Architecture for SVM Classifier for Machine Learning on Embedded DevicesabstractWith the proliferation of embedded computing, many machine learning (ML) applications have found their way into embedded devices. The SVM classifier is deemed suitable for real-world ML applications, which often comprise large-scale multi-dimensional data. In this paper, we introduce an FPGA-based systolic array architecture to support and accelerate SVM classifier on resource-constrained embedded devices. We also introduce a unique system-level architecture to further enhance speedup and to facilitate real-time processing. Our systolic array achieves a maximum speedup of 107x compared to its software counterparts, and a maximum classification accuracy of 98.5%. Srikanth Ramadurgam, Darshika G. Perera |
ISCAS | 2 |
| 2020 | A Quad Joint Relational Feature for 3D Skeletal Action Recognition with Circular CNNsabstractTo deal with the limitations of human action recognition systems that apply deep neural networks (DNNs) to 3D skeletal feature maps, we propose an improved set of features that enable better pattern discrimination when using a spectrally enriched circular convolutional neural network (CCNN). These new features exploit the local relationships between joint movements based on 3D quadrilaterals constructed for all possible sets of four joints. Next, we compute the volumes of these time-varying quadrilaterals, by generating color-coded images, named spatio-temporal quad-joint relative volume feature maps (QjRVMs). To preserve the pixel frequency distribution while training a DNN, which is otherwise lost due to vanishing gradients and random dropouts, we propose a new architecture CCNNs. CCNNs use cyclic multi-resolution filters in a four-stream architecture, requiring only batch normalization and ReLU operations to identify multiple pixel pattern variations simultaneously. Applying the proposed CCNN to QjRVM images illustrates that combining multi-resolution features enhances the overall classification accuracy. Finally, we evaluate our proposed human action framework using our own 102-class, 5-subject action dataset, created using 3D motion capture technology, named KLHA3D-102. We also evaluate our framework using 3 publicly available datasets: CMU, HDM05, and NTU RGB-D. P. V. V. Kishore, Darshika G. Perera, Maddala Teja Kiran Kumar, Dande Anil Kumar, Kiran Kumar Eepuri |
ISCAS | 2 |
| 2020 | An Optimized FPGA-Based Hardware Accelerator for Physics-Based EKF for Battery Cell ManagementabstractBattery technology is the cornerstone of HEVs. Model Predictive Control coupled with physics-based model (PBM) is an effective technique for battery management systems (BMS). In this case, extended Kalman filter (EKF) is used as the state observer for the highly non-linear PBM. Thus far, the sheer computational complexity of PBM hinders it from being utilized for portable BMS. In this paper, we introduce a novel and efficient FPGA-based hardware accelerator for physics-based EKF, to address the computational complexity of PBM. Anne K. Madsen, Michael Scott Trimboli, Darshika G. Perera |
ISCAS | 3 |
| 2017 | An efficient embedded multi-ported memory architecture for next-generation FPGAsabstractIn recent years, there has been a dramatic increase in utilization of FPGAs to enhance the speed-performance of many real-time compute and data intensive applications on embedded platforms. FPGA-based designs leverage parallelism in computations to achieve high speed-performance. Parallel computations require multi-ported memories to provide any number of ports for simultaneous multiple read/write (R/W) operations. Although several multi-ported memories are proposed in the literature, these designs become complex due to the extra logic and routing used for techniques/architectures to provide an arbitrary number of R/W ports. In this research work, we introduce a novel and efficient multi-ported memory architecture utilizing simple dual-port BRAMs, to provide an arbitrary number of R/W ports. Apart from the BRAMs, our proposed multi-ported memory design only consists of the Decision Making Modules and a counter, thus simplifying the design process. The R/W operations within our architecture are also straightforward. Experiments are performed to evaluate the feasibility and efficiency of our multi-ported memory architecture. We also evaluate our architecture with the most recently proposed multi-ported memory designs, implemented using LVT and XOR techniques, from the existing literature. FPGA manufacturers could employ our multi-ported memory architecture to accelerate real-time compute/data intensive applications with their next-generation FPGAs. Due to lower design complexity compared to the existing designs, our simplified memory architecture would enable seamless integration to the existing FPGA-based CAD tools with minimal design cost. S. Navid Shahrouzi, Darshika G. Perera |
ASAP | 2 |
| 2009 | Similarity Computation Using Reconfigurable Embedded HardwareabstractAdvances in portable devices and location-aware applications have necessitated the research in sophisticated yet small-footprint hardware and software in embedded systems, while the proliferation of the Web and distributed database systems has led to new data mining applications. We are investigating the utilization of reconfigurable hardware, due to its flexibility and performance, for data mining applications in portable and embedded computing. In this work, we introduce a reconfigurable hardware solution using field programmable gate array (FPGA) for similarity matrix computation, a commonly used data structure to represent the computed similarity among a set of feature vectors. Our hardware design can be dynamically reconfigured to accommodate three different similarity measures. A space-time cost analysis of the proposed multiplexer-based approach is presented. Experiments performed on the implemented reconfigurable hardware show encouraging and promising results that warrant further investigation in dynamically reconfigurable FPGA-based hardware for data mining applications. Darshika G. Perera, Kin Fun Li |
DASC | 1 |
| 2008 | Parallel Computation of Similarity Measures Using an FPGA-Based Processor ArrayabstractAn enormous amount of data needs to be processed in many data mining applications. In addition to algorithmic development, hardware support is imperative to improve the effectiveness and efficiency of these applications. We are investigating various hardware architectural design techniques and methodologies to support data mining at the chip level. In this work, we focus on the design of an FPGA-based processor array for the computation of similarity matrix, a commonly used data structure to represent the similarity among a set of feature vectors, with each matrix element representing the computed similarity measure between two vectors. An algorithm is developed to assign computation efficiently to the array of processing elements. Theoretical performance metrics are derived and compared to the experimental results. Performance gains using the processor array over software implementations are also presented and discussed. Darshika G. Perera, Kin Fun Li |
AINA | 1 |