Keisuke Sugiura

dblp:132/1848 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 A Sub-Microsecond HuffYUV-Based FPGA Accelerator for Real-Time Video Compression over High-Bandwidth Networks
abstract
This paper proposes a sub-microsecond FPGA-based video (de)compression system using the HuffYUV lossless format, designed to meet the ultra-low-latency demands of future high-speed networks such as IOWN and Beyond 5G. Unlike conventional codecs like H.264, HEVC, and JPEG XS, which introduce tens to hundreds of microseconds of latency, the proposed architecture achieves an end-to-end latency of just 1.1µs on Full HD 60fps video. By employing a bit-parallel Huffman decoding scheme, the system enables deterministic latency and efficient pipelining. Implemented on a Xilinx Versal FPGA (VPK120), the system consumes only a small portion of resources (e.g., 18.6% BRAMs) and achieves a throughput required to process 4K videos. Evaluation with a complete video pass-through system demonstrates its real-time performance and practical deployability. The architecture is highly suitable for edge applications requiring extreme responsiveness, including remote musical performance, telemedicine, and interactive XR.
Takeo Kurosawa, Keisuke Sugiura, Yoshiki Yamaguchi, Ryouhei Tsugami, Toshihito Fujiwara, Tatsuya Fukui, Satoshi Narikawa
CCNC2
2025 Marker-Based Recognition for Autonomous Micro-Drone Flight: An FPGA-Optimized Feasibility Study
abstract
Drones are nowadays essential in construction and structural maintenance, providing high-resolution images for structural assessments and precise interventions. However, complex maneuvering in confined spaces and reduced communication with external positioning systems, such as GPS, pose challenges for conventional drones. Micro-drones, with their small size that minimizes failure impact and enhances adaptability in restricted or hazardous environments, have gained increasing research interest for expanding drone applications across various fields. This study presents an energy-efficient design that incorporates an optimized combination of basic color separation and a custom downsized version of the Sobel operator to detect markers placed along the trajectory of the drone in real time. The flight instructions stem from the marker identification to fine-tune altitude, horizontal position, and orientation of the micro-drone before moving forward to the next marker. The proposed strategy demonstrated accurate marker detection under favorable lighting conditions 97% of the time, while significantly reducing BlockRAM usage.
Diego Marcelo Ramirez Jove, Keisuke Sugiura, Yoshiki Yamaguchi
VLSI-SoC2
2025 FPGA-accelerated Correspondence-free Point Cloud Registration with PointNet Features
abstract
Point cloud registration serves as a basis for vision and robotic applications including 3D reconstruction and mapping. Despite significant improvements on the quality of results, recent deep learning approaches are computationally expensive and power-hungry, making them difficult to deploy on resource-constrained edge devices. To tackle this problem, in this article, we propose a fast, accurate, and robust registration for low-cost embedded FPGAs. Based on a parallel and pipelined PointNet feature extractor, we develop custom accelerator cores namely PointLKCore and ReAgentCore, for two different learning-based methods. They are both correspondence-free and computationally efficient as they avoid the costly feature matching step involving nearest-neighbor search. The proposed cores are implemented on the Xilinx ZCU104 board and evaluated using both synthetic and real-world datasets, showing substantial improvements in the tradeoff between runtime and registration quality. They run 58.27–63.21x faster than ARM Cortex-A53 CPU and offer 1.54–10.91x speedups over Intel Xeon CPU and Nvidia Jetson boards, while consuming less than 1 W and achieving 133.63–267.57x energy efficiency compared to Nvidia GeForce GPU. The proposed cores are more robust to noise and large initial misalignments than the classical methods and quickly find reasonable solutions in less than 4–12 ms, demonstrating real-time performance.
Keisuke Sugiura, Hiroki Matsutani
ACM Trans. Reconfigurable Technol. Syst.1
2024 An Integrated FPGA Accelerator for Deep Learning-Based 2D/3D Path Planning
abstract
Path planning is a crucial component for realizing the autonomy of mobile robots. However, due to limited computational resources on mobile robots, it remains challenging to deploy state-of-the-art methods and achieve real-time performance. To address this, we propose P3Net (PointNet-based Path Planning Networks), a lightweight deep-learning-based method for 2D/3D path planning, and design an IP core (P3NetCore) targeting FPGA SoCs (Xilinx ZCU104). P3Net improves the algorithm and model architecture of the recently-proposed MPNet. P3Net employs an encoder with a PointNet backbone and a lightweight planning network in order to extract robust point cloud features and sample path points from a promising region. P3NetCore is comprised of the fully-pipelined point cloud encoder, batched bidirectional path planner, and parallel collision checker, to cover most part of the algorithm. On the 2D (3D) datasets, P3Net with the IP core runs 30.52–186.36x and 7.68–143.62x (15.69–93.26x and 5.30–45.27x) faster than ARM Cortex CPU and Nvidia Jetson while only consuming 0.255W (0.809W), and is up to 1278.14x (455.34x) power-efficient than the workstation. P3Net improves the success rate by up to 28.2% and plans a near-optimal path, leading to a significantly better tradeoff between computation and solution quality than MPNet and the state-of-the-art sampling-based methods.
Keisuke Sugiura, Hiroki Matsutani
IEEE Trans. Computers1
2023 An Efficient Accelerator for Deep Learning-based Point Cloud Registration on FPGAs
abstract
Point cloud registration is the basis for many robotic applications such as odometry and Simultaneous Localization And Mapping (SLAM), which are increasingly important for autonomous mobile robots. The limitation of computational resources and power budgets on such robots motivates us to study the resource-efficient registration method on low-cost edge devices. In this paper, we propose an FPGA-based novel pipeline for 3D point cloud registration built upon a recent deep learning-based method, PointNetLK. Based on the profiling results, we focus on the PointNet feature extraction as it becomes a major bottleneck; we improve its scalability and memory-efficiency by consuming each input point one-by-one in a pipelined manner instead of processing the whole point cloud at once. We then design a fully-parallelized and pipelined accelerator consisting of a custom PointNet IP core, which fits within both low-cost and mid-range FPGAs (e.g., Avnet Ultra96v2 and Xilinx ZCU104). Experimental results show that our proposed pipeline achieves up to 21.34x and 69.60x faster registration speed than the vanilla PointNetLK and ICP, respectively, while only consuming 722mW and maintaining the same level of accuracy.
Keisuke Sugiura, Hiroki Matsutani
PDP1
2022 P3Net: PointNet-based Path Planning on FPGA
abstract
Path planning is of crucial importance for au-tonomous mobile robots, and comes with a wide range of real-world applications including transportation, surveillance, and rescue. Currently, its high computational complexity is a major bottleneck for the application on such resource-limited robots. As a promising and effective solution to tackle this issue, in this paper, we propose a novel learning-based method for 2D/3D path planning, P3Net (PointNet-based Path Planning Network), along with its resource-efficient implementation targeting Xilinx ZCU104 boards. Our proposal is built upon two improvements to the recently proposed MPNet: we use a parameter-efficient PointNet-based encoder network to extract high-fidelity obstacle features from a point cloud, in conjunction with a lightweight planning network to iteratively plan a path. Experimental results using 2D/3D datasets demonstrate that our FPGA-based P3Net performs significantly better than MPNet and even comparable to the state-of-the-art sampling-based methods such as BIT*. P3Net is able to plan near-optimal paths 6.24x-9.34x faster than MPNet, and eventually improves the success rate by up to 24.45%, while reducing the parameter size by 5.43x-32.32x. This enables the subsecond real-time performance in many cases and opens up a new research direction for the edge-based efficient path planning.
Keisuke Sugiura, Hiroki Matsutani
FPT1
2022 dsODENet: Neural ODE and Depthwise Separable Convolution for Domain Adaptation on FPGAs
abstract
High-performance deep neural network (DNN)-based systems are in high demand in edge environments. Due to its high computational complexity, it is challenging to deploy DNNs on edge devices with strict limitations on computational resources. In this paper, we derive a compact while highly-accurate DNN model, termed dsODENet, by combining recently-proposed parameter reduction techniques: Neural ODE (Ordinary Differential Equation) and DSC (Depthwise Separable Convolution). Neural ODE exploits a similarity between ResNet and ODE, and shares most of weight parameters among multiple layers, which greatly reduces the memory consumption. We apply dsODENet to a domain adaptation as a practical use case with image classification datasets. We also propose a resource-efficient FPGA-based design for dsODENet, where all the parameters and feature maps except for pre- and post-processing layers can be mapped onto onchip memories. It is implemented on Xilinx ZCU104 board and evaluated in terms of domain adaptation accuracy, training speed, FPGA resource utilization, and speedup rate compared to a software counterpart. The results demonstrate that dsODENet achieves comparable or slightly better domain adaptation accuracy compared to our baseline Neural ODE implementation, while the total parameter size without pre- and post-processing layers is reduced by 54.2% to 79.8%. Our FPGA implementation accelerates the inference speed by 27.9 times.
Hiroki Kawakami, Hirohisa Watanabe, Keisuke Sugiura, Hiroki Matsutani
PDP3
2021 A unified accelerator design for LiDAR SLAM algorithms for low-end FPGAs
abstract
A fast and reliable LiDAR (Light Detection and Ranging) SLAM (Simultaneous Localization and Mapping) system is the growing need for autonomous mobile robots, which are used for a variety of tasks such as indoor cleaning, navigation, and transportation. To bridge the gap between the limited processing power on such robots and the high computational requirement of the SLAM system, in this paper we propose a unified accelerator design for 2D SLAM algorithms on resource-limited FPGA devices. As scan matching is the heart of these algorithms, the proposed FPGA-based accelerator utilizes scan matching cores on the programmable logic part and users can switch the SLAM algorithms to adapt to performance requirements and environments without modifying and re-synthesizing the logic part. We integrate the accelerator into two representative SLAM algorithms, namely particle filter-based and graph-based SLAM. They are evaluated in terms of resource utilization, processing speed, and quality of output results with various real-world datasets, highlighting their algorithmic characteristics. Experiment results on a Pynq-Z2 board demonstrate that scan matching is accelerated by 13.67–14.84x, improving the overall performance of particle filter-based and graph-based SLAM by 4.03–4.67x and 3.09–4.00x respectively, while maintaining the accuracy comparable to their software counterparts and even state-of-the-art methods.
Keisuke Sugiura, Hiroki Matsutani
FPT1