EDBT 2026 Demo / reviewers in the wild / expert
Masanori Hariyama
dblp:80/3876
· DBLP profile ↗
26ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-1464-8807ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large-Scale AGV Routing Based on Multi-FPGA SQA AccelerationabstractEnhancing the efficiency, safety, and speed of large-scale Automated Guided Vehicle (AGV) systems is critical to increasing the productivity of logistics warehouses. Studies using the latest quantum annealers such as "D-Wave Advantage" with over 5000 qubits have shown the potential of quantum annealing (QA) to rapidly optimize AGV routing. However, applying QA to complex and large-scale AGV routing problems is a challenging task due to insufficient consideration of intricate operational conditions, and also due to the insufficient number of qubits in quantum annealers. This paper proposes a refined combinatorial optimization problem that minimizes the total travel time of thousands of AGVs while enhancing safety and efficiency by avoiding collisions. To solve such large-scale optimization problems with thousands of variables, we also propose a novel system architecture containing a Simulated Quantum Annealing (SQA) accelerator using multiple FPGAs. The proposed SQA accelerator is capable of processing problems with over 50,000 variables, which could be a few tens to several hundred times larger than the problems processed on the "D-Wave Advantage". It addresses multiple combinatorial optimization problems across multiple FPGAs concurrently while processing each problem in a high degree of parallelism. We demonstrate the accurate operation of the proposed SQA accelerator using a real-world large-scale AGV system with over 1000 AGVs. According to the experimental results, we observed faster processing speed and better quality results over existing SQA solvers. Thinh NguyenQuang, Kosuke Matsuyama, Keisuke Shimizu, Hiroki Sugano, Eiji Kurimoto, Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Masayuki Ohzeki |
ASP-DAC | 7 |
| 2024 | Performance evaluation of Word2vec accelerators exploiting spatial and temporal parallelism on DDR/HBM-based FPGAsabstractAbstract Word embedding is a technique for representing words as vectors in a way that captures their semantic and syntactic relationships. The processing time of one of the most popular word embedding technique Word2vec is very large due to the huge data size. We evaluate the performance of a power-efficient FPGA-based accelerator designed using OpenCL. We achieved up to 18.7 times speed-up compared to single-core CPU implementation with the same accuracy. The proposed accelerator consumes less than 83 W of power and it is the most power-efficient one compared to many top-end CPU and GPU-based accelerators. Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
J. Supercomput. | 2 |
| 2022 | Word2Vec FPGA Accelerator Based on Spatial and Temporal Parallelism
Hasitha Muthumala Waidyasooriya, Shutaro Ishihara, Masanori Hariyama |
PDCAT | 3 |
| 2022 | FPGA-Accelerated Searchable Encrypted Database Management Systems for Cloud ServicesabstractThe use of database management systems (DBMSs) as a cloud service is rapidly expanding. Cloud DBMSs offer many advantages, such as easier management, lower costs, and greater scalability. However, there are still security concerns regarding attacks from adversaries. DBMSs that use searchable encryption have been investigated with regard to ensuring their security. Because searchable encryption allows query execution over encrypted data in the cloud, sensitive data can be securely stored there in the cloud. On the other hand, encrypted query processing is slower than query processing on plaintext data. In this article, we use a field-programmable gate array (FPGA) to accelerate query processing in a searchable encrypted DBMS. We also propose a new cache function to shorten the access time to database tables in a DBMS. According to an evaluation using basic queries, the proposed system has achieved up to 110.7 times speed-up compared with the central processing unit (CPU) processing of a single core. In addition, the proposed system can process queries faster than the plaintext processing on a CPU when processing large amounts of data. Mitsuhiro Okada 0002, Takayuki Suzuki, Naoya Nishio, Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
IEEE Trans. Cloud Comput. | 5 |
| 2022 | Design space exploration for an FPGA-based quantum annealing simulator with interaction-coefficient-generators
Chia-Yin Liu, Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
J. Supercomput. | 3 |
| 2022 | Temporal and spatial parallel processing of simulated quantum annealing on a multicore CPU
Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
J. Supercomput. | 2 |
| 2019 | FPGA-Based Acceleration of Word2vec using OpenCLabstractWord2vec is a word embedding method that converts words into vectors in such a way that the semantically and syntactically relevant words are close to each other in the vector space. The processing time of Word2vec is very large due to the huge data size. We propose a power efficient FPGA-based accelerator designed using OpenCL. We achieved 13.4 times speed-up compared to single-core CPU implementation with only 53W of power consumption. The proposed FPGA-based accelerator has the highest power-efficiency compared to existing top-end GPU-based accelerators. Taisuke Ono, Tomoki Shoji, Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Yuichiro Aoki, Yuki Kondoh, Yaoko Nakagawa |
ISCAS | 4 |
| 2019 | OpenCL-based design of an FPGA accelerator for quantum annealing simulation
Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Masamichi J. Miyama, Masayuki Ohzeki |
J. Supercomput. | 2 |
| 2017 | Architecture of an FPGA accelerator for LDA-based inferenceabstractLatent Dirichlet allocation (LDA) based topic inference is a data classification method, that is used efficiently for extremely large data sets. However, the processing time is very large due to the serial computational behavior of the Markov Chain Monte Carlo method used for the topic inference. We propose a pipelined hardware architecture and memory allocation scheme to accelerate LDA using parallel processing. The proposed architecture is implemented on a reconfigurable hardware called FPGA (field programmable gate array), using OpenCL design environment. According to the experimental results, we achieved maximum speed-up of 2.38 times, while maintaining the same quality compared to the conventional CPU-based implementation. Taisuke Ono, Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Tsukasa Ishigaki |
SNPD | 3 |
| 2017 | OpenCL-Based FPGA-Platform for Stencil Computation and Its Optimization MethodologyabstractStencil computation is widely used in scientific computations and many accelerators based on multicore CPUs and GPUs have been proposed. Stencil computation has a small operational intensity so that a large external memory bandwidth is usually required for high performance. FPGAs have the potential to solve this problem by utilizing large internal memory efficiently. However, a very large design, testing and debugging time is required to implement an FPGA architecture successfully. To solve this problem, we propose an FPGA-platform using C-like programming language called open computing language (OpenCL). We also propose an optimization methodology to find the optimal architecture for a given application using the proposed FPFA-platform. According to the experimental results, we achieved 119 - 237 Gflop/s of processing power and higher processing speed compared to conventional GPU and multicore CPU implementations. Hasitha Muthumala Waidyasooriya, Yasuhiro Takei, Shunsuke Tatsumi, Masanori Hariyama |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | FPGA-based deep-pipelined architecture for FDTD acceleration using OpenCLabstractAcceleration of the FDTD (finite-difference time-domain) computation is very important for the electromagnetic simulations. Conventional FDTD acceleration methods using multicore CPUs and CPUs have the common problem of memory-bandwidth limitation due to a large amount of parallel data access. Although FPGAs have the potential to solve this problem, very long design, testing and debugging time is required to implement an architecture successfully. To solve this problem, we propose an FPGA architecture designed using C-like programming language called OpenCL (open computing language). Therefore, the design time is very small and extensive knowledge about hardware-design is not required. We implemented the proposed architecture on an FPGA and achieved over 114 GFLOPS of processing power. We also achieved more than 13 times and 4 times speed-up compared to CPU and GPU implementations respectively. Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
ICIS | 2 |
| 2016 | Architecture of an FPGA accelerator for molecular dynamics simulation using OpenCLabstractMolecular dynamics (MD) simulations are very important to study physical properties of the atoms and molecules. However, a huge amount of processing time is required to simulate a few nano-seconds of an actual experiment. Although the hardware acceleration using FPGAs provides promising results, huge design time and hardware design skills are required to implement an accelerator successfully. In this paper, we propose an FPGA accelerator designed using C-based OpenCL. We achieved over 4.6 times of speed-up compared to CPU-based processing, by using only 36% of the Stratix V FPGA resources. Maximum of 18.4 times speed-up is possible by using 80% of the FPGA resources. Hasitha Muthumala Waidyasooriya, Masanori Hariyama, Kota Kasahara |
ICIS | 2 |
| 2016 | Hardware-Acceleration of Short-Read Alignment Based on the Burrows-Wheeler TransformabstractThe alignment of millions of short DNA fragments to a large genome is a very important aspect of the modern computational biology. However, software-based DNA sequence alignment takes many hours to complete. This paper proposes an FPGA-based hardware accelerator to reduce the alignment time. We apply a data encoding scheme that reduces the data size by 96 percent, and propose a pipelined hardware decoder to decode the data. We also design customized data paths to efficiently use the limited bandwidth of the DDR3 memories. The proposed accelerator can align a few hundred million short DNA fragments in an hour by using 80 processing elements in parallel. The proposed accelerator has the same mapping quality compared to the software-based methods. Hasitha Muthumala Waidyasooriya, Masanori Hariyama |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Asynchronous Domino Logic Pipeline Design Based on Constructed Critical Data PathabstractThis paper presents a high-throughput and ultralow-power asynchronous domino logic pipeline design method, targeting to latch-free and extremely fine-grain OR gate-level design. The data paths are composed of a mixture of dual-rail and single-rail domino gates. Dual-rail domino gates are limited to construct a stable critical data path. Based on this critical data path, the handshake circuits are greatly simplified, which offers the pipeline high throughput as well as low power consumption. Moreover, the stable critical data path enables the adoption of single-rail domino gates in the noncritical data paths. This further saves a lot of power by reducing the overhead of logic circuits. An 8×8 array style multiplier is used for evaluating the proposed pipeline method. Compared with a bundled-data asynchronous domino logic pipeline, the proposed pipeline, respectively, saves up to 60.2% and 24.5% of energy in the best case and the worst case when processing different data patterns. Zhengfan Xia, Masanori Hariyama, Michitaka Kameyama |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | FPGA implementation of heterogeneous multicore platform with SIMD/MIMD custom acceleratorsabstractHeterogeneous multi-core architecture with CPUs and accelerators attract many attentions since they can achieve power-efficient computing in various areas from low-power embedded processing to high-performance computing. Since the optimal architecture is different from application to applications, it is important to explore suitable architectures for different applications. In this paper, we propose an FPGA-based heterogeneous multi-core platform with custom accelerators for power-efficient computing. Our platform allows to select the most suitable accelerator according to the requirements of an application. Moreover, we optimize the number of ALUs, memory and interconnection network of the selected accelerators to increase the performance and to reduce the power consumption. Experimental results with simple media processing applications times power-efficient compared to the GPU. Hasitha Muthumala Waidyasooriya, Yasuhiro Takei, Masanori Hariyama, Michitaka Kameyama |
ISCAS | 3 |
| 2012 | Dual-rail/single-rail hybrid logic design for high-performance asynchronous circuitabstractThis paper presents a fine-grain pipelined asynchronous circuit that uses a mixture of dual-rail and single-rail logic. Dual-rail logic is limited to construct a stable critical path. Based on this critical path, the handshake control circuit is greatly simplified, which improves the performance of speed and power consumption. On the other hand, non-critical paths are composed of single-rail logic which has small logic overhead and the entire pipelined circuit has no intermediate registers or latches. To evaluate the proposed design method, an array style multiplier is designed and simulated in a 65nm design rule. The multiplier works as high as 4.35G data-set/s. Compared to the classical synchronous circuit, the proposed circuit has no active power consumption when there are no data operation. Even the circuits work at peak speed, the proposed circuit still reduces the power consumption by 35%. Zhengfan Xia, Shota Ishihara, Masanori Hariyama, Michitaka Kameyama |
ISCAS | 3 |
| 2011 | An implementation of an asychronous FPGA based on LEDR/four-phase-dual-rail hybrid architectureabstractThis paper presents an asynchronous FPGA that combines four-phase dual-rail encoding and LEDR (Level-Encoded Dual-Rail) encoding. Four-phase dual-rail encoding is used for small area and low power of function units, while LEDR encoding for high throughput and low power of data transfer. The proposed FPGA is fabricated in the e-Shuttle 65nm CMOS process and operates at 870 MHz. Compared to the synchronous FPGA, the power consumption is reduced by 38% for the workload of 15%. Yoshiya Komatsu, Shota Ishihara, Masanori Hariyama, Michitaka Kameyama |
ASP-DAC | 3 |
| 2011 | Memory Allocation Exploiting Temporal Locality for Reducing Data-Transfer Bottlenecks in Heterogeneous Multicore ProcessorsabstractHigh performance and low-power very large-scale integrations are required to implement complex media processing applications on mobile devices. Heterogeneous multicore processors are a promising way to achieve this objective. They contain multiple accelerator cores and CPU cores to increase the processing speed. Since media processing applications access a huge amount of data, fast address generation is very important. To increase the address generation speed, accelerator cores contain address generation units (AGUs). To reduce the power consumption, the AGUs have limited hardware resources such as adders and counters. Therefore, the AGUs generate simple addressing patterns where the address increases linearly in each clock cycle. Media processing applications frequently encounter addressing patterns where the same data are accessed in different time slots. To implement such addressing patterns, the same data have to be allocated into multiple memory addresses in such a way that those addresses can be generated by the AGUs. Allocation of the same data in multiple addresses is called the “data-duplication.” The data-duplication increases the data-transfer time and also the total processing time significantly. To remove such data-transfer bottlenecks, this paper proposes a memory allocation method that exploits the temporal and spatial locality of the memory access in media processing applications. We evaluate the proposed method using media processing applications to validate its effectiveness. According to the results, the proposed method reduces the total processing time by 14% to more than 85% compared to previous works. Hasitha Muthumala Waidyasooriya, Yosuke Ohbayashi, Masanori Hariyama, Michitaka Kameyama |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | A Low-Power FPGA Based on Autonomous Fine-Grain Power GatingabstractThis paper presents a field-programmable gate array (FPGA) based on lookup table level fine-grain power gating with small overheads. The power gating technique implemented in the proposed architecture can directly detect the activity of each look-up-table easily by exploiting features of asynchronous architectures. Moreover, detecting the data arrival in advance prevents the delay increase for waking-up and the power consumption of unnecessary power switching. Since the power gating technique has small overheads, the granularity size of a power-gated domain is as fine as a single two-input and one-output lookup table. The proposed FPGA is fabricated using the ASPLA 90-nm CMOS process with dual threshold voltages. We use an image processing application called “template matching” for evaluation. Since the proposed FPGA is suitable for processing where the workload changes dynamically, an adaptive algorithm where a small computational kernel is employed. Compared to a synchronous FPGA and an asynchronous FPGA without power gating, the power consumption is reduced respectively by 38% and 15% at 85°C. Shota Ishihara, Masanori Hariyama, Michitaka Kameyama |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | A low-power FPGA based on autonomous fine-grain power-gatingabstractThis is the first implementation of an FPGA based on autonomous fine-grain power-gating. To cut the power consumption of clock network and detect the activity of the cell efficiently, asynchronous architecture is full exploited. The proposed FPGA is fabricated in a 90nm CMOS process with dual threshold voltages. It is more efficient in power than the synchronous FPGA at less than 30% utilization. Shota Ishihara, Masanori Hariyama, Michitaka Kameyama |
ASP-DAC | 2 |
| 2009 | Optimal Periodic Memory Allocation for Image Processing With Multiple WindowsabstractOne major issue in designing image processors is to design a memory system that supports parallel access with a simple interconnection network. This paper presents an efficient memory allocation to minimize the number of memory modules and processing elements with a parallel access capability when multiple windows with arbitrary shapes are specified. This paper also presents an efficient search method based on regularity of window-type image processing. We give some practical examples including a stereo-matching processor for acquiring 3-D information, and an optical-flow processor for motion estimation. These examples show that the numbers of memory modules are reduced to 2.7% and 10%, respectively, in comparison with a basic approach. It is also shown that the search time is less than 1 ms for practical image sizes and window sizes. Yasuhiro Kobayashi, Masanori Hariyama, Michitaka Kameyama |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | FPGA implementation of a vehicle detection algorithm using three-dimensional informationabstractThis paper presents a vehicle detection algorithm using 3-dimensional(3D) information and its FPGA implementation. For high-speed acquisition of 3D information, feature- based stereo matching is employed to reduce search area. Our algorithm consists of some tasks with high degree of column- level parallelism. Based on the parallelism, we propose area- efficient VLSI architecture with local data transfer between memory modules and processing elements. Images are equally divided into blocks with some columns, and a block is allocated to a PE. Each PE performs the processing in parallel. The proposed architecture is implemented on FPGA (Altera Stratix EP1S40F1020C7). For specifications of image size 640 times 480, 100 frames/sec, and operating frequency 100 MHz, only 11,000 logic elements (< 30%) are required for 30 PEs. Masanori Hariyama, Kensaku Yamashita, Michitaka Kameyama |
IPDPS | 1 |
| 2006 | Architecture of a multi-context FPGA using a hybrid multiple-valued/binary context switching signalabstractMulti-context FPGAs have multiple memory bits per configuration bit forming configuration planes for fast switching between contexts. Large amount of memory causes significant overhead in area and power consumption. This paper presents two key technologies. The first is a floating-gate-MOS functional pass gate that merges storage and switching functions area - efficiently. The second is the use of a hybrid multiple-valued/binary context switching signal that eliminates redundancy of a conventional multi-context (MC) switch with high scalability. The transistor count of the proposed MC-switch is reduced to 7% in comparison with that of a SRAM-based one. Yoshihiro Nakatani, Masanori Hariyama, Michitaka Kameyama |
IPDPS | 2 |
| 2005 | Genetic Approach to Minimizing Energy Consumption of VLSI Processors Using Multiple Supply VoltagesabstractThis paper presents an efficient search method for a scheduling and module selection problem using multiple supply voltages so as to minimize dynamic energy consumption under time and area constraints. The proposed algorithm is based on a genetic algorithm so that it can find near-optimal solutions in a short time for large-size problems, n efficient search can be achieved by crossover that prevents generating nonvalid individuals and a local search is also utilized in the algorithm. Experimental results for large-size problems with 1,000 operations demonstrate that the proposed method can achieve significant energy reduction up to 50 percent and can find a near-optimal solution (within 2.8 percent from the lower bound of optimal solutions) in 10 minutes. On the other hand, the ILP-based method cannot find any feasible solution in one hour for the large-size problem, even if a state-of-art mathematical programming solver is used. Masanori Hariyama, Tetsuya Aoyama, Michitaka Kameyama |
IEEE Trans. Computers | 1 |
| 2001 | VLSI Processor for Reliable Stereo Matching Based on Adaptive Window-Size SelectionabstractStereo vision is a well known method to acquire 3D information. One important problem in stereo vision is to establish reliable correspondence between images. Another problem is that the correspondence search is time-consuming. This paper presents a reliable stereo-matching algorithm and a new parallel VLSI processor architecture for stereo matching. One commonly-used method to establish correspondence between images is the SAD (sum of absolute differences) method. A window size is iteratively enlarged to select as small a window for each pixel as possible that can avoid ambiguity based on uniqueness of a minimum of an SAD graph. This process is called a global search. Next, the estimate of the corresponding pixel obtained by the global search is iteratively refined by shrinking the window size. To avoid ambiguity with a small window size, the correspondence estimate obtained by the global search is efficiently used. The proposed algorithm has regular data flow based on iterations of SAD computation so that it is suitable for parallel processing. Masanori Hariyama, Toshiki Takeuchi, Michitaka Kameyama |
ICRA | 1 |
| 1998 | Design of a Collision Detection VLSI Processor Based on Minimization of Area-Time ProductsabstractThis paper presents the design of a new high-performance VLSI processor based on a systematic methodology for area minimization under a time constraint. A VLSI-oriented algorithm based on regular iterations of coordinate transformation and matching operation are introduced. The VLSI-processor consists of several identical clusters which has a CAM for parallel matching operation and PEs for parallel coordinate transformation. Under a condition of 100% utilization of PEs and a CAM, area minimization of the VLSI-processor is attributed to minimization of area-time products of a CAM and a PE. A multiport CAM (MCAM) and a PE based on bit-serial pipelined architecture can be efficiently employed for the minimization. The result shows that the total area can be reduced by about 30% in comparison with a straightforward design and that the performance is several ten thousand times higher than that of a general-purpose processor. Masanori Hariyama, Michitaka Kameyama |
ICRA | 1 |