EDBT 2026 Demo / reviewers in the wild / expert
Duvindu Piyasena
dblp:228/2441
· DBLP profile ↗
7ranked-venue papers
5as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hardware Accelerator for Feature Matching with Binary Search TreeabstractFeature matching is an essential step for autonomous robot to localize itself during navigation. However, it is often difficult to achieve the matching in real-time due to limited on-board computing resources. We present an FPGA implementation for stream-processing based feature matching called Binary Search Tree Matcher (BSTM) that utilizes a balanced binary search tree. To improve the matching precision, we integrated a ratio-test (RT) outlier rejection mechanism in our design. Our proposed FPGA-based BSTM is scalable and resource efficient, and significantly faster than the best-performing hardware implementation in the literature. When compared to the commonly used Linear Exhaustive Search (LES) method running on the FPGA, our proposed design is approximately 12X faster. Miyuru Thathsara, Siew-Kei Lam, Damith Kawshan, Duvindu Piyasena |
ISCAS | 4 |
| 2021 | Edge Accelerator for Lifelong Deep Learning using Streaming Linear Discriminant AnalysisabstractLifelong deep learning models are expected to continuously adapt and acquire new knowledge in dynamic environments. This capability is essential for numerous vision tasks in robotics and drones, and the models must be deployed on the edge to achieve real-time performance. We propose a FPGA accelerator of a streaming classifier for lifelong deep learning, which is based on streaming linear discriminant analysis (SLDA). When combined with a frozen Convolutional Neural Network (CNN) model, the proposed system is capable of class incremental lifelong learning for object classification. Duvindu Piyasena, Siew-Kei Lam, Meiqing Wu |
FCCM | 1 |
| 2021 | Accelerating Continual Learning on Edge FPGAabstractReal-time edge AI systems operating in dynamic environments must learn quickly from streaming input samples without needing to undergo offline model training. We propose an FPGA accelerator for continual learning based on streaming linear discriminant analysis (SLDA), which is capable of class-incremental object classification. The proposed SLDA accelerator employs application-specific parallelism, efficient data reuse, resource sharing, and approximate computing to achieve high performance and power efficiency. Additionally, we introduce a new variant of SLDA and discuss the accuracy-efficiency trade-offs. The proposed SLDA accelerator is combined with a Convolutional Neural Network (CNN). which is implemented on Xilinx DPU to achieve full continual learning capability at nearly the same latency as inference. Experiments based on popular datasets for continual learning, CoRE50 and CUB200, demonstrate that the proposed SLDA accelerator outperforms the embedded CPU and GPU counterparts, in terms of speed and energy efficiency. Duvindu Piyasena, Siew-Kei Lam, Meiqing Wu |
FPL | 1 |
| 2020 | Dynamically Growing Neural Network Architecture for Lifelong Deep Learning on the EdgeabstractConventional deep learning models are trained once and deployed. However, models deployed in agents operating in dynamic environments need to constantly acquire new knowledge, while preventing catastrophic forgetting of previous knowledge. This ability is commonly referred to as lifelong learning. In this paper, we address the performance and resource challenges for realizing lifelong learning on edge devices. We propose a FPGA based architecture for a Self-Organization Neural Network (SONN), that in combination with a Convolutional Neural Network (CNN) can perform class-incremental lifelong learning for object classification. The proposed SONN architecture is capable of performing unsupervised learning on input features from the CNN by dynamically growing neurons and connections. In order to meet the tight constraints of edge computing, we introduce efficient scheduling methods to maximize resource reuse and parallelism, as well as approximate computing strategies. Experiments based on the Core50 dataset for continuous object recognition from video sequences demonstrated that the proposed FPGA architecture significantly outperforms CPU and GPU based implementations. Duvindu Piyasena, Miyuru Thathsara, Sathursan Kanagarajah, Siew-Kei Lam, Meiqing Wu |
FPL | 1 |
| 2019 | Reducing Dynamic Power in Streaming CNN Hardware Accelerators by Exploiting Computational RedundanciesabstractConvolutional neural networks (CNNs) have achieved tremendous successes in various application domains such as computer vision. However, current implementations are characterized by large memory requirements and accesses, which pose an impediment towards their deployment on low cost embedded devices with fast runtime requirements. Recently, FPGA based streaming CNN hardware accelerators have been reported for alleviating these memory bottlenecks. However, these implementations suffer from large number of convolution operations which incur high power consumption. In this paper, we investigate methods to exploit the redundancies in the activation layers in order to reduce the dynamic power. We propose a computationally efficient approximation method to reduce the overall convolution operations with marginal accuracy loss. Experimental results of our FPGA implementation based on image classification datasets show that the proposed method leads to considerable power savings. Duvindu Piyasena, Rukshan Wickramasinghe, Debdeep Paul, Siew-Kei Lam, Meiqing Wu |
FPL | 1 |
| 2019 | Lowering Dynamic Power of a Stream-based CNN Hardware AcceleratorabstractCustom hardware accelerators of Convolutional Neural Networks (CNN) provide a promising solution to meet real-time constraints for a wide range of applications on low-cost embedded devices. In this work, we aim to lower the dynamic power of a stream-based CNN hardware accelerator by reducing the computational redundancies in the CNN layers. In particular, we investigate the redundancies due to the downsampling effect of max pooling layers which are prevalent in state-of-the-art CNNs, and propose an approximation method to reduce the overall computations. The experimental results show that the proposed method leads to lower dynamic power without sacrificing accuracy. Duvindu Piyasena, Rukshan Wickramasinghe, Debdeep Paul, Siew-Kei Lam, Meiqing Wu |
MMSP | 1 |
| 2018 | A Novel Low-Complexity VLSI Architecture for an EEG Feature Extraction PlatformabstractThis paper presents a novel low-complexity VLSI architecture for an EEG Feature Extraction Platform for wearable health monitoring systems. It integrates a processing core unit, preprocessing FIR filter block and an FFT accelerator. The platform supports 8 scalp EEG channels and it can be scaled up to 16 or 32 EEG channels. In this paper, we analyze several architectural optimizations for hardware implementation of features such as Power Spectral Density, Discrete Wavelet Transform, Autoregression etc. The proposed architecture is easily scalable to new feature extractors. An FFT accelerator is proposed to improve the flexibility of the platform. We provide a comparison with software models for each feature extractor and also validate the platform by implementing a seizure detection algorithm on CHB-MIT Scalp EEG database with an accuracy of over 95%. Dakila Serasinghe, Duvindu Piyasena, Ajith Pasqual |
DSD | 2 |