Yun Li 0015

dblp:87/6284-15 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0003-1753-7317ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUs
abstract
The 3D Non-Local Means (NLM) algorithm has become a crucial preprocessing technique for 3D image datasets due to its effectiveness in denoising while preserving fine details. This method has been proven to be highly efficient in high-demand tasks within industrial applications such as medical imaging and remote sensing. The 3D NLM algorithm computes the filtered value for each voxel by calculating the weighted average of all voxels within a 3D search window, where the weights are determined by the similarity between pairs of 3D template windows. Therefore, the computational burden becomes significant, especially in embedded GPUs with limited computational power and memory resources. To address this issue, we propose an efficient GPU parallel kernel to minimize redundant computations and memory accesses. The kernel integrates three nested reuse strategies to handle redundant computations in three dimensions: for columns, we leverage the fast data exchange mechanism to reuse column computation results via on-chip registers; for rows, we use a sliding window strategy, utilizing GPU global memory as an intermediary to store and reuse similarity values between filtered rows; and for channels, we introduce a zigzag scanning strategy that enables simultaneous computation across multiple channels and employs on-chip registers to facilitate channel computation reuse. Experimental results demonstrate that our kernel achieves an average speedup of 7.7x on the embedded Jetson AGX Xavier platform across a range of 3D image datasets compared to existing methods, showcasing exceptional performance.
Xiang Li 0110, Qiong Chang, Yun Li 0015, Jun Miyazaki
ACM Trans. Archit. Code Optim.3
2024 An Optimized GPU Implementation for GIST Descriptor
abstract
The GIST descriptor is a classic feature descriptor primarily used for scene categorization and recognition tasks. It drives a bank of Gabor filters, which respond to edges and textures at various scales and orientations to capture the spatial structures in an image. Compared to other scene recognition algorithms that rely on detailed object detection, GIST has lower computational complexity, allowing it to be widely applied. However, its internal multi-scale and multi-orientation Gabor filters also mean that systems based on it cannot be executed fast enough. This article proposes an optimized GPU kernel for the GIST descriptor. It fully takes advantage of the symmetry of Gabor filters and proposes different optimization strategies for both oblique and orthogonal orientations. Extensive experiments demonstrate that the proposed kernel is adaptable to images of various scales and different GPUs. Compared to the cuFFT library, our kernel achieves 12.09× and 3.86× acceleration on an RTX 3080 GPU and a Jetson AGX Xavier GPU, respectively.
Xiang Li 0110, Qiong Chang, Aolong Zha, Shijie Chang, Yun Li 0015, Jun Miyazaki
ACM Trans. Archit. Code Optim.5
2024 TinyStereo: A Tiny Coarse-to-Fine Framework for Vision-Based Depth Estimation on Embedded GPUs
abstract
Stereo vision, a popular depth estimation technology in computing vision, finds wide-ranging applications in embedded systems, including robotics vision and autonomous driving. These applications demand both high accuracy and fast processing speeds. To address hardware limitations, most current embedded systems rely on nonlearning algorithms for fast matching, sacrificing accuracy. Some recent studies have explored using convolutional neural networks (CNNs) to improve matching accuracy, but the computational load of existing learning-based systems hampers real-world applicability. This article presents significant contributions: 1) a novel stereo matching framework that greatly enhances accuracy on real-time embedded platforms and 2) a two-pronged approach combining a nonlearning-based algorithm and a lightweight super-resolution residual neural network (sRRNet). The nonlearning-based algorithm yields a low-resolution disparity map, while the lightweight sRRNet generates a high-resolution disparity map. Experimental results on benchmark data demonstrate that the proposed method achieves a low matching error rate of 5.17% and a real-time processing speed of 51 fps using the embedded Jetson AGX GPU. The proposed method outperforms all existing real-time embedded systems.
Qiong Chang, Aolong Zha, Meng Joo Er, Yongqing Sun, Yun Li 0015
IEEE Trans. Syst. Man Cybern. Syst.6
2023 StereoVAE: A lightweight stereo-matching system using embedded GPUs
abstract
We propose a lightweight system for stereo-matching using embedded graphic processing units (GPUs). The proposed system overcomes the trade-off between accuracy and processing speed in stereo matching, thus further improving the matching accuracy while ensuring real-time processing. The basic idea is to construct a tiny neural network based on a variational autoencoder (VAE) to achieve the upscaling and refinement a small size of coarse disparity map. This map is initially generated using a traditional matching method. The proposed hybrid structure maintains the advantage of low computational complexity found in traditional methods. Additionally, it achieves matching accuracy with the help of a neural network. Extensive experiments on the KITTI 2015 benchmark dataset demonstrate that our tiny system exhibits high robustness in improving the accuracy of coarse disparity maps generated by different algorithms, while running in real-time on embedded GPUs.
Qiong Chang, Xiang Li 0110, Xin Liu 0020, Yun Li 0015, Jun Miyazaki
ICRA5
2023 Multi-directional Sobel operator kernel on GPUs
Qiong Chang, Xiang Li 0110, Yun Li 0015, Jun Miyazaki
J. Parallel Distributed Comput.3