EDBT 2026 Demo / reviewers in the wild / expert
Hasitha Muthumala Waidyasooriya
dblp:59/2309
· DBLP profile ↗
15ranked-venue papers
10as first author
6since 2021 · last 2025
0000-0001-5108-9891ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large-Scale AGV Routing Based on Multi-FPGA SQA AccelerationabstractEnhancing the efficiency, safety, and speed of large-scale Automated Guided Vehicle (AGV) systems is critical to increasing the productivity of logistics warehouses. Studies using the latest quantum annealers such as "D-Wave Advantage" with over 5000 qubits have shown the potential of quantum annealing (QA) to rapidly optimize AGV routing. However, applying QA to complex and large-scale AGV routing problems is a challenging task due to insufficient consideration of intricate operational conditions, and also due to the insufficient number of qubits in quantum annealers. This paper proposes a refined combinatorial optimization problem that minimizes the total travel time of thousands of AGVs while enhancing safety and efficiency by avoiding collisions. To solve such large-scale optimization problems with thousands of variables, we also propose a novel system architecture containing a Simulated Quantum Annealing (SQA) accelerator using multiple FPGAs. The proposed SQA accelerator is capable of processing problems with over 50,000 variables, which could be a few tens to several hundred times larger than the problems processed on the "D-Wave Advantage". It addresses multiple combinatorial optimization problems across multiple FPGAs concurrently while processing each problem in a high degree of parallelism. We demonstrate the accurate operation of the proposed SQA accelerator using a real-world large-scale AGV system with over 1000 AGVs. According to the experimental results, we observed faster processing speed and better quality results over existing SQA solvers. Thinh NguyenQuang, Kosuke Matsuyama, Keisuke Shimizu, Hiroki Sugano, Eiji Kurimoto, Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Masayuki Ohzeki |
ASP-DAC | 6 |
| 2024 | Performance evaluation of Word2vec accelerators exploiting spatial and temporal parallelism on DDR/HBM-based FPGAsabstractAbstract Word embedding is a technique for representing words as vectors in a way that captures their semantic and syntactic relationships. The processing time of one of the most popular word embedding technique Word2vec is very large due to the huge data size. We evaluate the performance of a power-efficient FPGA-based accelerator designed using OpenCL. We achieved up to 18.7 times speed-up compared to single-core CPU implementation with the same accuracy. The proposed accelerator consumes less than 83 W of power and it is the most power-efficient one compared to many top-end CPU and GPU-based accelerators. Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
J. Supercomput. | 1 |
| 2022 | Word2Vec FPGA Accelerator Based on Spatial and Temporal Parallelism
Hasitha Muthumala Waidyasooriya, Shutaro Ishihara, Masanori Hariyama |
PDCAT | 1 |
| 2022 | FPGA-Accelerated Searchable Encrypted Database Management Systems for Cloud ServicesabstractThe use of database management systems (DBMSs) as a cloud service is rapidly expanding. Cloud DBMSs offer many advantages, such as easier management, lower costs, and greater scalability. However, there are still security concerns regarding attacks from adversaries. DBMSs that use searchable encryption have been investigated with regard to ensuring their security. Because searchable encryption allows query execution over encrypted data in the cloud, sensitive data can be securely stored there in the cloud. On the other hand, encrypted query processing is slower than query processing on plaintext data. In this article, we use a field-programmable gate array (FPGA) to accelerate query processing in a searchable encrypted DBMS. We also propose a new cache function to shorten the access time to database tables in a DBMS. According to an evaluation using basic queries, the proposed system has achieved up to 110.7 times speed-up compared with the central processing unit (CPU) processing of a single core. In addition, the proposed system can process queries faster than the plaintext processing on a CPU when processing large amounts of data. Mitsuhiro Okada 0002, Takayuki Suzuki, Naoya Nishio, Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | Design space exploration for an FPGA-based quantum annealing simulator with interaction-coefficient-generators
Chia-Yin Liu, Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
J. Supercomput. | 2 |
| 2022 | Temporal and spatial parallel processing of simulated quantum annealing on a multicore CPU
Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
J. Supercomput. | 1 |
| 2019 | FPGA-Based Acceleration of Word2vec using OpenCLabstractWord2vec is a word embedding method that converts words into vectors in such a way that the semantically and syntactically relevant words are close to each other in the vector space. The processing time of Word2vec is very large due to the huge data size. We propose a power efficient FPGA-based accelerator designed using OpenCL. We achieved 13.4 times speed-up compared to single-core CPU implementation with only 53W of power consumption. The proposed FPGA-based accelerator has the highest power-efficiency compared to existing top-end GPU-based accelerators. Taisuke Ono, Tomoki Shoji, Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Yuichiro Aoki, Yuki Kondoh, Yaoko Nakagawa |
ISCAS | 3 |
| 2019 | OpenCL-based design of an FPGA accelerator for quantum annealing simulation
Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Masamichi J. Miyama, Masayuki Ohzeki |
J. Supercomput. | 1 |
| 2017 | Architecture of an FPGA accelerator for LDA-based inferenceabstractLatent Dirichlet allocation (LDA) based topic inference is a data classification method, that is used efficiently for extremely large data sets. However, the processing time is very large due to the serial computational behavior of the Markov Chain Monte Carlo method used for the topic inference. We propose a pipelined hardware architecture and memory allocation scheme to accelerate LDA using parallel processing. The proposed architecture is implemented on a reconfigurable hardware called FPGA (field programmable gate array), using OpenCL design environment. According to the experimental results, we achieved maximum speed-up of 2.38 times, while maintaining the same quality compared to the conventional CPU-based implementation. Taisuke Ono, Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Tsukasa Ishigaki |
SNPD | 2 |
| 2017 | OpenCL-Based FPGA-Platform for Stencil Computation and Its Optimization MethodologyabstractStencil computation is widely used in scientific computations and many accelerators based on multicore CPUs and GPUs have been proposed. Stencil computation has a small operational intensity so that a large external memory bandwidth is usually required for high performance. FPGAs have the potential to solve this problem by utilizing large internal memory efficiently. However, a very large design, testing and debugging time is required to implement an FPGA architecture successfully. To solve this problem, we propose an FPGA-platform using C-like programming language called open computing language (OpenCL). We also propose an optimization methodology to find the optimal architecture for a given application using the proposed FPFA-platform. According to the experimental results, we achieved 119 - 237 Gflop/s of processing power and higher processing speed compared to conventional GPU and multicore CPU implementations. Hasitha Muthumala Waidyasooriya, Yasuhiro Takei, Shunsuke Tatsumi, Masanori Hariyama |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | FPGA-based deep-pipelined architecture for FDTD acceleration using OpenCLabstractAcceleration of the FDTD (finite-difference time-domain) computation is very important for the electromagnetic simulations. Conventional FDTD acceleration methods using multicore CPUs and CPUs have the common problem of memory-bandwidth limitation due to a large amount of parallel data access. Although FPGAs have the potential to solve this problem, very long design, testing and debugging time is required to implement an architecture successfully. To solve this problem, we propose an FPGA architecture designed using C-like programming language called OpenCL (open computing language). Therefore, the design time is very small and extensive knowledge about hardware-design is not required. We implemented the proposed architecture on an FPGA and achieved over 114 GFLOPS of processing power. We also achieved more than 13 times and 4 times speed-up compared to CPU and GPU implementations respectively. Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
ICIS | 1 |
| 2016 | Architecture of an FPGA accelerator for molecular dynamics simulation using OpenCLabstractMolecular dynamics (MD) simulations are very important to study physical properties of the atoms and molecules. However, a huge amount of processing time is required to simulate a few nano-seconds of an actual experiment. Although the hardware acceleration using FPGAs provides promising results, huge design time and hardware design skills are required to implement an accelerator successfully. In this paper, we propose an FPGA accelerator designed using C-based OpenCL. We achieved over 4.6 times of speed-up compared to CPU-based processing, by using only 36% of the Stratix V FPGA resources. Maximum of 18.4 times speed-up is possible by using 80% of the FPGA resources. Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Kota Kasahara |
ICIS | 1 |
| 2016 | Hardware-Acceleration of Short-Read Alignment Based on the Burrows-Wheeler TransformabstractThe alignment of millions of short DNA fragments to a large genome is a very important aspect of the modern computational biology. However, software-based DNA sequence alignment takes many hours to complete. This paper proposes an FPGA-based hardware accelerator to reduce the alignment time. We apply a data encoding scheme that reduces the data size by 96 percent, and propose a pipelined hardware decoder to decode the data. We also design customized data paths to efficiently use the limited bandwidth of the DDR3 memories. The proposed accelerator can align a few hundred million short DNA fragments in an hour by using 80 processing elements in parallel. The proposed accelerator has the same mapping quality compared to the software-based methods. Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | FPGA implementation of heterogeneous multicore platform with SIMD/MIMD custom acceleratorsabstractHeterogeneous multi-core architecture with CPUs and accelerators attract many attentions since they can achieve power-efficient computing in various areas from low-power embedded processing to high-performance computing. Since the optimal architecture is different from application to applications, it is important to explore suitable architectures for different applications. In this paper, we propose an FPGA-based heterogeneous multi-core platform with custom accelerators for power-efficient computing. Our platform allows to select the most suitable accelerator according to the requirements of an application. Moreover, we optimize the number of ALUs, memory and interconnection network of the selected accelerators to increase the performance and to reduce the power consumption. Experimental results with simple media processing applications times power-efficient compared to the GPU. Hasitha Muthumala Waidyasooriya, Yasuhiro Takei, Masanori Hariyama, Michitaka Kameyama |
ISCAS | 1 |
| 2011 | Memory Allocation Exploiting Temporal Locality for Reducing Data-Transfer Bottlenecks in Heterogeneous Multicore ProcessorsabstractHigh performance and low-power very large-scale integrations are required to implement complex media processing applications on mobile devices. Heterogeneous multicore processors are a promising way to achieve this objective. They contain multiple accelerator cores and CPU cores to increase the processing speed. Since media processing applications access a huge amount of data, fast address generation is very important. To increase the address generation speed, accelerator cores contain address generation units (AGUs). To reduce the power consumption, the AGUs have limited hardware resources such as adders and counters. Therefore, the AGUs generate simple addressing patterns where the address increases linearly in each clock cycle. Media processing applications frequently encounter addressing patterns where the same data are accessed in different time slots. To implement such addressing patterns, the same data have to be allocated into multiple memory addresses in such a way that those addresses can be generated by the AGUs. Allocation of the same data in multiple addresses is called the “data-duplication.” The data-duplication increases the data-transfer time and also the total processing time significantly. To remove such data-transfer bottlenecks, this paper proposes a memory allocation method that exploits the temporal and spatial locality of the memory access in media processing applications. We evaluate the proposed method using media processing applications to validate its effectiveness. According to the results, the proposed method reduces the total processing time by 14% to more than 85% compared to previous works. Hasitha Muthumala Waidyasooriya, Yosuke Ohbayashi, Masanori Hariyama, Michitaka Kameyama |
IEEE Trans. Circuits Syst. Video Technol. | 1 |