Yu Long 0005

dblp:13/102-5 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 VSLAM-BA: Algorithm and Hardware Co-Design for High Performance and Energy-Efficient Visual SLAM Backend Hardware Accelerator
abstract
Visual Simultaneous Localization and Mapping (VSLAM) is a key localization technology for emerging applications such as autonomous driving and uncrewed aerial vehicles (UAVs). Compared with VSLAM frontend, VSLAM backend plays a more important role as it is employed to improve the localization accuracy. However, the VSLAM backend usually uses Bundle Adjustment (BA) as its core optimization method which is well-known for its large scale of problem construction, high computational complexity and high serialization of data processing, making it difficult to achieve high performance and energy efficiency on platforms such as CPUs or GPUs. Although there are some VSLAM backend accelerators proposed recently for addressing the above issues, they did not well exploit the data regularity and computational characteristics, resulting in limited performance/energy efficiency improvements or degraded accuracy. In this work, we propose VSLAM-BA which is a high performance and energy-efficient VSLAM backend accelerator with algorithm-hardware co-design. On the algorithm level, a keyframe-split-based Schur elimination scheme is proposed to reduce latency, power consumption and memory storage while maintaining accuracy. On the hardware level, a column-folding-based computing architecture is proposed to boost performance and energy efficiency. A loading-sensitive matrix-computing technique with an adaptive task scheduler is proposed to reduce the latency and energy consumption. Further, a recyclable computing technique with point-aware solver is proposed to reduce the memory and energy consumption. The experimental results show that the proposed VSLAM-BA achieves the highest performance (380 fps) and the highest energy efficiency (0.51 mJ per frame) with high accuracy and low memory storage, compared with the SOTA designs. The proposed accelerator can work with different VSLAM frontend for backend optimization of localization accuracy.
Ye Liu 0011, Xiuyuan Qi, Shuang Hao 0005, Zili Huang, Neng Zhao, Ruixin Mao, Sixu Li, Ang Hu, Yu Long 0005, Shanshan Liu 0001, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.14
2025 An Ultra-High Performance and Scalable Optical Flow Hardware Accelerator Based on FPGA for Autonomous Driving
abstract
Optical flow plays an extremely important role in the field of computer vision and extremely high real-time performance is required especially in autonomous driving. Traditional optical flow methods generally improve the accuracy of optical flow through the image pyramid technique. Nevertheless, the incorporation of the image pyramid elevates the computational complexity. Moreover, the data relationships between pyramid layers result in strong data dependencies, which renders it difficult to accelerate via parallel processing and makes it challenging to fulfill real-time demands in practical scenarios. To address this issue, in this paper, we propose an ultra-high performance and scalable optical flow hardware accelerator based on FPGA with several techniques, including an adaptive optical flow computation technique based on dynamic direction prediction to reduce computation without accuracy degradation, a highly scalable computing architecture with configurable numbers of PEs to improve the flexibility and hardware utilization under different hardware resource constraints, and a reconfigurable pyramid-layer pipeline technique to improve performance and reduce memory size. The proposed hardware accelerator was implemented and evaluated on a Xilinx FPGA ZCU104 achieving ultra-high performance (405 FPS) while maintaining high accuracy (AEE 1.02) compared with SOTA hardware accelerators.
Ye Liu 0011, Shuang Hao 0005, Xiuyuan Qi, Zili Huang, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.10
2025 PRADO: A Low-Latency and Energy-Efficient 6DoF Pose Refinement Accelerator With Domain-Specific Explorations
abstract
Six degrees of freedom (6DoF) pose estimation is a critical technique for applications involving humanoid robotics, autonomous driving, and virtual and augmented reality (VR/AR). Pose refinement plays a pivotal role in the 6DoF pose estimation, significantly enhancing accuracy by iteratively refining the initial pose derived from a single-shot pose estimation (SSPE) process. However, this iterative approach is time-and energy-consuming which is not suitable for resource-and power-constrained edge-devices. In this paper, we propose PRADO, an energy-efficient 6DoF pose refinement accelerator for edge applications. Several domain-specific explorations aimed at reducing processing latency and energy-consumption while maintaining high accuracy have been proposed, including an adaptive sparse correspondences sampling (ASCS) architecture to reduce redundant correspondences process, a hybrid static & dynamic pruning (HSDP) technique with co-designed hardware architecture to reduce computational complexity, and a cosine similarity-based adaptive early exit (CSAE2) technique to eliminate unnecessary iterations. Implemented and evaluated on the ZCU102 FPGA board, PRADO achieves the lowest processing latency of 61.2ms and energy consumption of 198.6mJ, and highest accuracy of 0.7112, outperforming state-of-the-art designs. Implemented with 28nm CMOS technology, PRADO achieves an even lower processing latency of 18.1ms and energy consumption of 12.6mJ.
Haojie Wei, Le Jin, Junhao Zeng, Liwei Zhuang, Yu Long 0005, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.9
2024 An FPGA-based Ultra-High Performance and Scalable Optical Flow Hardware Accelerator for Autonomous Driving
abstract
Optical flow plays an extremely important role in the field of computer vision and extremely high real-time performance is required especially in autonomous driving. Traditional optical flow methods generally improve the accuracy of optical flow through the image pyramid technique. However, the introduction of the image pyramid increases the computational complexity. Additionally, the data relationships between pyramid layers result in strong data dependencies, making it difficult to accelerate using parallel processing, and it is challenging to meet real-time requirements in practical scenarios. To address this issue, in this paper, we propose an FPGA-based ultra-high performance and scalable optical flow hardware accelerator with several techniques, including an adaptive optical flow computation technique based on dynamic direction prediction to reduce computation without accuracy degradation, a highly scalable computing architecture with configurable numbers of PEs to improve the flexibility and hardware utilization under different hardware resource constraints, and a reconfigurable pyramid-layer pipeline technique to improve performance and reduce memory size. The proposed hardware accelerator was implemented and evaluated on a Xilinx FPGA ZCU104 achieving ultra-high performance (405 FPS) while maintaining high accuracy (AEE 0.64).
Ye Liu 0011, Shuang Hao 0005, Zili Huang, Xiuyuan Qi, Yu Long 0005, Jun Zhou 0017
ISCAS9
2023 A High Accuracy and Low Power CNN-Based Environmental Sound Classification Processor
abstract
The environmental sound classification (ESC) has attracted increasing attention as the environmental sound contains a wealth of information that can be used to detect particular events. However, so far, most of the existing work in ESC still remains in the stage of algorithm design and the design of ESC processor has not been thoroughly investigated. The existing ESC processor designs have issues in meeting low power consumption and high accuracy simultaneously due to the lack of joint-optimization between algorithm and hardware, and very few work has demonstrated a complete ESC system containing all the necessary modules. In this work, a high accuracy and low power CNN-based ESC processor has been proposed, featuring: 1) a big-small CNN-based reconfigurable ESC processing hardware architecture to reduce the power consumption and hardware overhead while maintaining high classification accuracy. 2) a Mel feature adaptation engine reusing the neural network processing unit to further reduce the power consumption. 3) an event-driven ESC processing technique to reduce the inference time and the power consumption. The design has been implemented on a Kintex-7 FPGA and achieves low power consumption of 0.313W with high accuracy of 84.5% for the ESC-50 dataset, outperforming other state-of-the-art ESC processors.
Lujie Peng, Junyu Yang, Longke Yan, Xiben Jiao, Jianbiao Xiao, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.9
2022 MobileSP: An FPGA-Based Real-Time Keypoint Extraction Hardware Accelerator for Mobile VSLAM
abstract
Keypoint extraction is a key technique for Visual Simultaneous Localization and Mapping (VSLAM). Recently, Convolutional Neural Network (CNN) has been used in the keypoint extraction for improving the accuracy. As one of the state-of-the-art CNN based keypoint extraction techniques, the SuperPoint ranked top in the CVPR2020 image matching challenge. However, the use of complex CNN makes it difficult to meet the real-time performance on a mobile platform with limited resource such as mobile robots and wearable Augmented Reality (AR) devices. In this work, based on the SuperPoint, we proposed an FPGA-based real-time keypoint extraction hardware accelerator through algorithm-hardware co-design for mobile VSLAM applications, which is named as MobileSP. Several algorithm and hardware level design techniques have been proposed to reduce the computation and improve the processing speed while maintaining high accuracy, including a partially shared detection & description encoding architecture, a pre-sorting based Non-Maximum Suppression (NMS) engine and a software-hardware hybrid pipeline computing technique. The design has been implemented and evaluated on a ZCU104 FPGA board. It achieves real-time performance of 42 fps with low Absolute Trajectory Error (ATE) of 1.82 cm simultaneously, outperforming several state-of-the-art designs.
Ye Liu 0011, Xiuyuan Qi, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.7
2022 ULECGNet: An Ultra-Lightweight End-to-End ECG Classification Neural Network
abstract
ECG classification is a key technology in intelligent electrocardiogram (ECG) monitoring. In the past, traditional machine learning methods such as support vector machine (SVM) and K-nearest neighbor (KNN) have been used for ECG classification, but with limited classification accuracy. Recently, the end-to-end neural network has been used for ECG classification and shows high classification accuracy. However, the end-to-end neural network has large computational complexity including a large number of parameters and operations. Although dedicated hardware such as field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) can be developed to accelerate the neural network, they result in large power consumption, large design cost, or limited flexibility. In this work, we have proposed an ultra-lightweight end-to-end ECG classification neural network that has extremely low computational complexity (∼8.2k parameters & ∼227k multiplication/addition operations) and can be squeezed into a low-cost microcontroller (MCU) such as MSP432 while achieving 99.1% overall classification accuracy. This outperforms the state-of-the-art ECG classification neural network. Implemented on MSP432, the proposed design consumes only 0.4 mJ and 3.1 mJ per heartbeat classification for normal and abnormal heartbeats respectively for real-time ECG classification.
Jianbiao Xiao, Jiahao Liu 0006, Huanqi Yang, Ning Wang 0070, Zhen Zhu 0005, Yu Long 0005, Liang Chang 0002, Jun Zhou 0017
IEEE J. Biomed. Health Informatics8