Yun Li 0015

dblp:87/6284-15 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0003-1753-7317ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
GPUs and heterogeneous computing · 100%
Artificial intelligence
2 papers
Image recognition and object detection · 36% 3D vision · 32% Efficient and distributed learning · 32%
Computer graphics and multimedia
2 papers
Image and video processing · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU kernel optimization
1.622025
3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUs · ACM Trans. Archit. Code Optim. 2025
An Optimized GPU Implementation for GIST Descriptor · ACM Trans. Archit. Code Optim. 2024
GPUs and heterogeneous computing
embedded GPU
1.122025
3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUs · ACM Trans. Archit. Code Optim. 2025
StereoVAE: A lightweight stereo-matching system using embedded GPUs · ICRA 2023
Image and video processing › image restoration
image denoising
0.912025
3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUs · ACM Trans. Archit. Code Optim. 2025
Image and video processing › image restoration › image denoising › patch-based denoising
non-local means
0.912025
3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUs · ACM Trans. Archit. Code Optim. 2025
Computer vision › Image recognition and object detection
scene recognition
0.812024
An Optimized GPU Implementation for GIST Descriptor · ACM Trans. Archit. Code Optim. 2024
Machine learning › Efficient and distributed learning
model compression
0.712023
StereoVAE: A lightweight stereo-matching system using embedded GPUs · ICRA 2023
Computer vision › 3D vision › stereo vision
stereo matching
0.712023
StereoVAE: A lightweight stereo-matching system using embedded GPUs · ICRA 2023
Medical and health informatics
medical imaging
0.312025
3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUs · ACM Trans. Archit. Code Optim. 2025
Image and video processing › image filtering › directional filtering
gabor filtering
0.212024
An Optimized GPU Implementation for GIST Descriptor · ACM Trans. Archit. Code Optim. 2024

Methods — techniques the papers use, named apart from their topics

zigzag scanning · 2.6sliding window · 2.6nested reuse strategies · 2.6filter symmetry exploitation · 2.3GPU parallelization · 2.3variational autoencoder · 1.3neural network · 1.3
YearPublicationVenuePosition
2025 3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUs
abstract
The 3D Non-Local Means (NLM) algorithm has become a crucial preprocessing technique for 3D image datasets due to its effectiveness in denoising while preserving fine details. This method has been proven to be highly efficient in high-demand tasks within industrial applications such as medical imaging and remote sensing. The 3D NLM algorithm computes the filtered value for each voxel by calculating the weighted average of all voxels within a 3D search window, where the weights are determined by the similarity between pairs of 3D template windows. Therefore, the computational burden becomes significant, especially in embedded GPUs with limited computational power and memory resources. To address this issue, we propose an efficient GPU parallel kernel to minimize redundant computations and memory accesses. The kernel integrates three nested reuse strategies to handle redundant computations in three dimensions: for columns, we leverage the fast data exchange mechanism to reuse column computation results via on-chip registers; for rows, we use a sliding window strategy, utilizing GPU global memory as an intermediary to store and reuse similarity values between filtered rows; and for channels, we introduce a zigzag scanning strategy that enables simultaneous computation across multiple channels and employs on-chip registers to facilitate channel computation reuse. Experimental results demonstrate that our kernel achieves an average speedup of 7.7x on the embedded Jetson AGX Xavier platform across a range of 3D image datasets compared to existing methods, showcasing exceptional performance.
Xiang Li 0110, Qiong Chang, Yun Li 0015, Jun Miyazaki
ACM Trans. Archit. Code Optim.3
2024 An Optimized GPU Implementation for GIST Descriptor
abstract
The GIST descriptor is a classic feature descriptor primarily used for scene categorization and recognition tasks. It drives a bank of Gabor filters, which respond to edges and textures at various scales and orientations to capture the spatial structures in an image. Compared to other scene recognition algorithms that rely on detailed object detection, GIST has lower computational complexity, allowing it to be widely applied. However, its internal multi-scale and multi-orientation Gabor filters also mean that systems based on it cannot be executed fast enough. This article proposes an optimized GPU kernel for the GIST descriptor. It fully takes advantage of the symmetry of Gabor filters and proposes different optimization strategies for both oblique and orthogonal orientations. Extensive experiments demonstrate that the proposed kernel is adaptable to images of various scales and different GPUs. Compared to the cuFFT library, our kernel achieves 12.09× and 3.86× acceleration on an RTX 3080 GPU and a Jetson AGX Xavier GPU, respectively.
Xiang Li 0110, Qiong Chang, Aolong Zha, Shijie Chang, Yun Li 0015, Jun Miyazaki
ACM Trans. Archit. Code Optim.5
2024 TinyStereo: A Tiny Coarse-to-Fine Framework for Vision-Based Depth Estimation on Embedded GPUs
abstract
Stereo vision, a popular depth estimation technology in computing vision, finds wide-ranging applications in embedded systems, including robotics vision and autonomous driving. These applications demand both high accuracy and fast processing speeds. To address hardware limitations, most current embedded systems rely on nonlearning algorithms for fast matching, sacrificing accuracy. Some recent studies have explored using convolutional neural networks (CNNs) to improve matching accuracy, but the computational load of existing learning-based systems hampers real-world applicability. This article presents significant contributions: 1) a novel stereo matching framework that greatly enhances accuracy on real-time embedded platforms and 2) a two-pronged approach combining a nonlearning-based algorithm and a lightweight super-resolution residual neural network (sRRNet). The nonlearning-based algorithm yields a low-resolution disparity map, while the lightweight sRRNet generates a high-resolution disparity map. Experimental results on benchmark data demonstrate that the proposed method achieves a low matching error rate of 5.17% and a real-time processing speed of 51 fps using the embedded Jetson AGX GPU. The proposed method outperforms all existing real-time embedded systems.
Qiong Chang, Aolong Zha, Meng Joo Er, Yongqing Sun, Yun Li 0015
IEEE Trans. Syst. Man Cybern. Syst.6
2023 StereoVAE: A lightweight stereo-matching system using embedded GPUs
abstract
We propose a lightweight system for stereo-matching using embedded graphic processing units (GPUs). The proposed system overcomes the trade-off between accuracy and processing speed in stereo matching, thus further improving the matching accuracy while ensuring real-time processing. The basic idea is to construct a tiny neural network based on a variational autoencoder (VAE) to achieve the upscaling and refinement a small size of coarse disparity map. This map is initially generated using a traditional matching method. The proposed hybrid structure maintains the advantage of low computational complexity found in traditional methods. Additionally, it achieves matching accuracy with the help of a neural network. Extensive experiments on the KITTI 2015 benchmark dataset demonstrate that our tiny system exhibits high robustness in improving the accuracy of coarse disparity maps generated by different algorithms, while running in real-time on embedded GPUs.
Qiong Chang, Xiang Li 0110, Xin Liu 0020, Yun Li 0015, Jun Miyazaki
ICRA5
2023 Multi-directional Sobel operator kernel on GPUs
Qiong Chang, Xiang Li 0110, Yun Li 0015, Jun Miyazaki
J. Parallel Distributed Comput.3