Xiuyuan Qi

dblp:336/5250 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0006-5377-7106ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 VSLAM-BA: Algorithm and Hardware Co-Design for High Performance and Energy-Efficient Visual SLAM Backend Hardware Accelerator
abstract
Visual Simultaneous Localization and Mapping (VSLAM) is a key localization technology for emerging applications such as autonomous driving and uncrewed aerial vehicles (UAVs). Compared with VSLAM frontend, VSLAM backend plays a more important role as it is employed to improve the localization accuracy. However, the VSLAM backend usually uses Bundle Adjustment (BA) as its core optimization method which is well-known for its large scale of problem construction, high computational complexity and high serialization of data processing, making it difficult to achieve high performance and energy efficiency on platforms such as CPUs or GPUs. Although there are some VSLAM backend accelerators proposed recently for addressing the above issues, they did not well exploit the data regularity and computational characteristics, resulting in limited performance/energy efficiency improvements or degraded accuracy. In this work, we propose VSLAM-BA which is a high performance and energy-efficient VSLAM backend accelerator with algorithm-hardware co-design. On the algorithm level, a keyframe-split-based Schur elimination scheme is proposed to reduce latency, power consumption and memory storage while maintaining accuracy. On the hardware level, a column-folding-based computing architecture is proposed to boost performance and energy efficiency. A loading-sensitive matrix-computing technique with an adaptive task scheduler is proposed to reduce the latency and energy consumption. Further, a recyclable computing technique with point-aware solver is proposed to reduce the memory and energy consumption. The experimental results show that the proposed VSLAM-BA achieves the highest performance (380 fps) and the highest energy efficiency (0.51 mJ per frame) with high accuracy and low memory storage, compared with the SOTA designs. The proposed accelerator can work with different VSLAM frontend for backend optimization of localization accuracy.
Ye Liu 0011, Xiuyuan Qi, Shuang Hao 0005, Zili Huang, Neng Zhao, Ruixin Mao, Sixu Li, Ang Hu, Yu Long 0005, Shanshan Liu 0001, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 An Ultra-High Performance and Scalable Optical Flow Hardware Accelerator Based on FPGA for Autonomous Driving
abstract
Optical flow plays an extremely important role in the field of computer vision and extremely high real-time performance is required especially in autonomous driving. Traditional optical flow methods generally improve the accuracy of optical flow through the image pyramid technique. Nevertheless, the incorporation of the image pyramid elevates the computational complexity. Moreover, the data relationships between pyramid layers result in strong data dependencies, which renders it difficult to accelerate via parallel processing and makes it challenging to fulfill real-time demands in practical scenarios. To address this issue, in this paper, we propose an ultra-high performance and scalable optical flow hardware accelerator based on FPGA with several techniques, including an adaptive optical flow computation technique based on dynamic direction prediction to reduce computation without accuracy degradation, a highly scalable computing architecture with configurable numbers of PEs to improve the flexibility and hardware utilization under different hardware resource constraints, and a reconfigurable pyramid-layer pipeline technique to improve performance and reduce memory size. The proposed hardware accelerator was implemented and evaluated on a Xilinx FPGA ZCU104 achieving ultra-high performance (405 FPS) while maintaining high accuracy (AEE 1.02) compared with SOTA hardware accelerators.
Ye Liu 0011, Shuang Hao 0005, Xiuyuan Qi, Zili Huang, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 An FPGA-based Ultra-High Performance and Scalable Optical Flow Hardware Accelerator for Autonomous Driving
abstract
Optical flow plays an extremely important role in the field of computer vision and extremely high real-time performance is required especially in autonomous driving. Traditional optical flow methods generally improve the accuracy of optical flow through the image pyramid technique. However, the introduction of the image pyramid increases the computational complexity. Additionally, the data relationships between pyramid layers result in strong data dependencies, making it difficult to accelerate using parallel processing, and it is challenging to meet real-time requirements in practical scenarios. To address this issue, in this paper, we propose an FPGA-based ultra-high performance and scalable optical flow hardware accelerator with several techniques, including an adaptive optical flow computation technique based on dynamic direction prediction to reduce computation without accuracy degradation, a highly scalable computing architecture with configurable numbers of PEs to improve the flexibility and hardware utilization under different hardware resource constraints, and a reconfigurable pyramid-layer pipeline technique to improve performance and reduce memory size. The proposed hardware accelerator was implemented and evaluated on a Xilinx FPGA ZCU104 achieving ultra-high performance (405 FPS) while maintaining high accuracy (AEE 0.64).
Ye Liu 0011, Shuang Hao 0005, Zili Huang, Xiuyuan Qi, Yu Long 0005, Jun Zhou 0017
ISCAS6
2024 A High-Performance ORB Accelerator with Algorithm and Hardware Co-design for Visual Localization
abstract
Vision-based localization plays an important role in numerous emerging applications, including mobile robots, UAVs, AR/VR, etc. The ORB algorithm is a very classic algorithm in visual localization, typically consisting of feature point extraction and matching. Owing to the extensive computational load involved in these processes, real-time performance remains difficult in applications based on embedded devices. In this work, we present a high-performance ORB accelerator through the algorithm and hardware co-design, including a three-levels parallel processing architecture to improve performance, a pixel-level RS-BRIEF descriptor generator to reduce the computational complexity, and an on-chip sliding storage bucket matching and overwriting strategy for storage optimization. The proposed work has been implemented on the ZCU104 FPGA board achieving a high-performance (108 FPS) and a low average trajectory error at the cost of slightly fewer hardware resources compared with some SOTA works.
Xiuyuan Qi, Ye Liu 0011, Shuang Hao 0005, Zherong Liu, Jun Zhou 0011
ISCAS1
2022 MobileSP: An FPGA-Based Real-Time Keypoint Extraction Hardware Accelerator for Mobile VSLAM
abstract
Keypoint extraction is a key technique for Visual Simultaneous Localization and Mapping (VSLAM). Recently, Convolutional Neural Network (CNN) has been used in the keypoint extraction for improving the accuracy. As one of the state-of-the-art CNN based keypoint extraction techniques, the SuperPoint ranked top in the CVPR2020 image matching challenge. However, the use of complex CNN makes it difficult to meet the real-time performance on a mobile platform with limited resource such as mobile robots and wearable Augmented Reality (AR) devices. In this work, based on the SuperPoint, we proposed an FPGA-based real-time keypoint extraction hardware accelerator through algorithm-hardware co-design for mobile VSLAM applications, which is named as MobileSP. Several algorithm and hardware level design techniques have been proposed to reduce the computation and improve the processing speed while maintaining high accuracy, including a partially shared detection & description encoding architecture, a pre-sorting based Non-Maximum Suppression (NMS) engine and a software-hardware hybrid pipeline computing technique. The design has been implemented and evaluated on a ZCU104 FPGA board. It achieves real-time performance of 42 fps with low Absolute Trajectory Error (ATE) of 1.82 cm simultaneously, outperforming several state-of-the-art designs.
Ye Liu 0011, Xiuyuan Qi, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.5