Yaobin Wang

dblp:17/6158 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AES-SpMM: Balancing accuracy and speed by adaptive edge sampling strategy to accelerate SpMM in GNNs
Yingchen Song, Yaobin Wang, Pingping Tang
Neurocomputing2
2026 SpineSiamSwin: An IoMT-Driven Siamese Swin Transformer Model Based on Transfer Learning for Intelligent Diagnosis of Spinal Diseases
abstract
The incidence of spinal diseases has been steadily increasing in recent years, with affected populations becoming progressively younger. With the rapid development of artificial intelligence and medical image processing technologies, automated intelligent diagnostic algorithms based on spinal X-ray images have been continuously developed. However, current deep learning–based diagnostic methods for spinal diseases face limitations such as the limited availability of labeled data and subtle inter-class feature differences. To address these challenges in spinal X-ray image diagnosis, this study proposes a novel Siamese Swin Transformer model based on transfer learning, within the Internet of Medical Things (IoMT) framework—referred to as the SpineSiamSwin model. The model adopts a Siamese network structure and leverages metric learning loss functions to maximize the distance between different classes in the embedding space, thereby enhancing the model’s sensitivity to subtle structural differences between disease categories. Extensive experiments on real-world spinal X-ray datasets demonstrate the effectiveness and practical applicability of the SpineSiamSwin model.
Changgong Lan, Jiajie Lin, Yaobin Wang, Zhiqiang Wu 0001, Hao Wu 0144
IEEE Internet Things J.4
2025 DCI: An Efficient Workload-Aware Dual-Cache Allocation GNN Inference Acceleration System
Yaobin Wang, Yingchen Song, Qingfeng Wang 0004, Jun Huang 0005
Euro-Par (2)2
2025 MH-SpGEMM: Efficient Sparse General Matrix-Matrix Multiplication on Modern GPUs via Masking and Hashing Cooperative Optimization
abstract
Sparse General Matrix-Matrix Multiplication (SpGEMM) is a core operation in high-performance computing applications such as algebraic multigrid solvers, machine learning, and graph processing. Traditional hash-based SpGEMM methods face several challenges, including frequent hash table accesses during the symbolic phase, high-overhead sorting operations in the numeric phase, and imbalanced workload distribution across threads. To address these issues, we propose a highly optimized SpGEMM method-MH-SpGEMM. Unlike conventional approaches, MH-SpGEMM introduces a mask-based storage format for matrix$B$during the symbolic phase and uses bitwise OR operations to efficiently calculate the number of nonzeros in the result matrix. In the numeric phase, a bitonic sort algorithm is employed for rows with many nonzeros to improve sorting efficiency. Furthermore, MH-SpGEMM incorporates a finer-grained load balancing strategy. We conduct a comprehensive evaluation of MHSpGEMM against four state-of-the-art SpGEMM methods (HSMU-SpGEMM, OpSparse, nsparse, and cuSPARSE) on two NVIDIA GPUs(Ada Lovelace, Ampere). The results demonstrate that MH-SpGEMM consistently outperforms existing methods in terms of performance.
Yaobin Wang, Qiong Yu
ICCD2
2025 HR-SpMM: Adaptive Row Partitioning and Hybrid Kernel Design for Sparse Matrix Multiplication
abstract
Sparse Matrix-Matrix Multiplication (SpMM) plays a critical role in high-performance computing and applications like Graph Neural Networks (GNNs).However, due to the sparsity and irregularity of real-world data, optimizing SpMM performance on modern GPUs has remained a significant challenge.Existing methods often involve trade-offs between load balancing and hardware utilization, making it difficult to efficiently handle both long and short rows in sparse matrices.To address these issues, we propose HR-SpMM, a lightweight framework based on adaptive row partitioning and hybrid kernel design.HR-SpMM divides sparse matrix rows into two categories: long rows and short rows, leveraging Tensor Cores and CUDA Cores, respectively, to optimize computational efficiency.Long rows are further partitioned into fixed-size blocks to fully align with the hardware characteristics of Tensor Cores, while short rows adopt a flexible
Qi Wang 0113, Yaobin Wang, Pingping Tang
ICS2
2025 Multi-view isosurface similarity analysis for transfer function design in direct volume rendering
Yaobin Wang
Comput. Graph.5
2025 Efficient yet secure: An archive knowledge graph-enhanced native sparse attention network for lightweight privacy-preserving recommendation
Chenxi Ma, Yaobin Wang, Limei Sun
Knowl. Based Syst.3
2024 A Data Augmentation Approach for Well Log Interpretation
Yaobin Wang, Guanwen Zhang, Wei Zhou 0020
ICPR (25)2
2024 PIK-Convolution: Step Convolution Acceleration based on Multi-GPU Architecture
abstract
SimpleConvolution is the most important and time-consuming part of convolutional neural networks (CNN) for image processing. Each slide of the window in two-dimensional convolution will cause repeated memory access. To solve above problems, we propose an algorithm integrated with the hardware level, parallel in kernel Convolution(PIK-Convolution), which reduces the granularity of a SingleConvolution and avoid the problem of storage access pressure caused by spatial locality. Unlike Image to column(im2col) algorithm, we avoid the problem of non-adjacent data in memory directly from hardware perspective. With MGPUSim, a multi-GPU architecture platform, we implement our algorithm. A large number of experimental results show that the proposed algorithm performs well on the new architecture. Compared with the unoptimized algorithm, the overall efficiency of the new algorithm is optimized by 2.2×-7.45×, and the memory access efficiency is an average of 30× of the unoptimized algorithm. Compared to the state-of-the-art convolution acceleration, our algorithm has a performance optimization of 1.5×-2.2.
Yutao Peng, Yaobin Wang, Tianhai Wang, Yunxin Xu, Pingping Tang
ISPA2
2024 P-AES: Advanced Encryption Standard Parallel Optimization on MGPUSim
abstract
Advanced Encryption Standard (AES), as one of the most popular encryption algorithms, has been widely studied on single GPU and CPU. However, the research on multi-GPU platforms is not deep enough, and with the rapid increase in data size, it is difficult for single-GPU platforms to meet the demand for high-performance computing. To solve this problem, this paper presents a novel AES parallelization study "P-AES" on MGPUSim, including its parallel execution mechanism as well as architectural design. In addition, we modify the instruction set of MGPUSim along with thread optimization to save memory overhead, increase data access speed, and improve system performance by loading S-boxes and key extension arrays from global memory to shared memory. The experimental results show that: (1) P-AES performs well on MGPUSim, achieving an average speedup of 1.5×-1.7× compared to the AES implementation using CUDA kernels, and a speedup of 2.75× compared to M-AES (the unoptimized AES on MGPUSim). (2) By conducting encryption experiments on plaintext data of varying sizes, we found that P-AES exhibits high stability and throughput. Compared to the modern sliced AES, the P-AES algorithm achieves a throughput of up to 812 Gbps, which is 1.3 times that of their implementation and 2.74 times that of M-AES.
Jiawei Qin, Yaobin Wang
ISPA2
2024 An Efficient Sampling-Based SpMM Kernel for Balancing Accuracy and Speed in GNN Inference
abstract
How to coordinate the design of sampling and Sparse-dense Matrix Multiplication (SpMM) is important in Graph Neural Network (GNN) acceleration. However, existing methods have an imbalance between accuracy and speed in performing GNN inference tasks due to irrational sampling strategies. To solve this problem, we propose an adaptive edge sampling strategy SpMM kernel. It considers the relationship between the number of non-zero elements in each matrix row and the shared memory width. The edge sampling scheme is adaptively selected according to the different situations of each row. Our method reduces the graph size by adaptive edge sampling to fit into the GPU’s shared memory, which decreases the computational cost and increases the data locality ultimately achieving a balance between accuracy and speed in GNN inference. We conducted experiments on NVIDIA RTX 4060 Ti GPU using representative GNN model and datasets. Experimental results show that our designed kernel outperforms the cuSPARSE SpMM kernel and GE-SpMM by up to 20.2× and 17.3× respectively, with less than 1% accuracy loss. Compared to ES-SpMM, it reduces the accuracy loss by 4.7% on average and achieves an average 1.26× speedup.
Yingchen Song, Yaobin Wang, Chaoyu Xiong, Tianhai Wang, Pingping Tang
ISPA2
2024 Implementation and Optimization of 8×8 Block Discrete Cosine Transform on MGPUSim
abstract
Discrete cosine transform for 8×8 block(DCT8×8) is widely used in image compression due to its high signal decorrelation rate. Current research for DCT is mainly focused on CPU and single GPU platforms, and the exploration of multi-GPU architectures is still insufficient. With the rapid increase of image data volume, the DCT8×8 algorithm on CPU and single GPU architectures can no longer meet the demand of efficient computation. Therefore, optimizing DCT8×8 algorithms on multi-GPU architectures becomes particularly important. To address this challenge, this paper explores the porting and optimization strategies of DCT algorithms on the new MGPUSim architecture.First, we extend the instruction implementation of MG-PUSim to accomplish the porting of the original DCT8×8 algorithm (O-DCT), and then propose a kernel implementation strategy for optimizing the DCT8×8 algorithm by setting the size of the workgroup appropriately so that the kernel function can take advantage of the symmetry of the transformation matrix to reduce the redundant computation, which we refer to as the M-DCT. We conducted experiments on multiple grayscale images with different resolutions. The experimental results show that O-DCT performs well on MGPUSim, achieving an average 1.93× speedup ratio compared to the original DCT8×8 kernel (kernel1) in the CUDA SDK. Further experimental results show that M-DCT improves the execution efficiency on MGPUSim by 1.76× relative to O-DCT, and also achieves an average of 1.3× acceleration ratio compared to the optimized DCT8×8 kernel (kernel2) in the CUDA SDK, which improves the computational performance of the DCT8×8 algorithm on multi-GPU architectures. It provides new ideas and methods for future research on image compression on multi-GPU architectures.
Yaobin Wang, Jiawei Qin, Guotang Bi
ISPA2
2024 UTR: A UNet-like transformer for efficient unsupervised medical image registration
Lianjin Xiong, Ning Li 0030, Yaobin Wang, Yangsong Zhang 0001
Image Vis. Comput.4
2023 The Optimization and Parallelization of Two-Dimensional Zigzag Scanning on the Matrix
Yaobin Wang, Lijuan Peng, Guangwei Li, Xiaolin Jia
ICANN (4)2
2022 Local-Whole-Focus: Identifying Breast Masses and Calcified Clusters on Full-Size Mammograms
abstract
The detection of breast masses and calcified clusters on mammograms is critical for early diagnosis and treatment to improve the survivals of breast cancer patients. In this study, we propose a local-whole-focus pipeline to automatically identify breast masses and calcified clusters on full-size mammograms, from local breast tissues to the whole mammograms, and then focusing on the lesion areas. We first train a deep model to learn the fine features of breast masses and calcified clusteres on local breast tissues, and then transfer the well-trained deep model to identify breast masses and calcified clusteres on full-size mammograms with image-level annotations. We also highlight the areas of the breast masses and calcified clusteres in mammograms to visualize the identification results. We evaluated the proposed local-whole-focus pipeline on a public dataset CBIS-DDSM (Curated Breast Imaging Subset of Digital Database for Screening Mammography) and a private dataset MY-Mammo (Mianyang central hospital mammograms). The experiment results showed the DenseNet embedded with squeeze-and-excitation (SE) blocks achieved competitive results on the identification of breast masses and calcified clusteres on full-size mammograms. The highlight areas of the breast masses and calcified clusteres on the entire mammograms could also explain model decision making, which are important in practical medical applications.
Jun Huang 0005, Qingfeng Wang 0004, Zhiqin Liu, Yaobin Wang
BIBM6
2022 The Parallelization and Optimization of K-means Algorithm Based on MGPUSim
Zhangbin Mo, Yaobin Wang, Qingming Zhang, Guangbing Zhang, Mingfeng Guo
ICANN (4)2
2022 Rgs-SpMM: Accelerate Sparse Matrix-Matrix Multiplication by Row Group Splitting Strategy on the GPU
Mingfeng Guo, Yaobin Wang, Jun Huang 0005, Qingfeng Wang 0004, Mu Xu
NPC2
2020 Procedure and Loop Level Speculative Parallelism Analysis in HPEC
Yaobin Wang, Deqing Bu, Manasah Musariri
ICA3PP (1)2
2019 Higher-order Transfer Learning for Pulmonary Nodule Attribute Prediction in Chest CT Images
abstract
Attributes like texture, lobulation, malignancy, etc., are commonly used to describe the phenotype of a pulmonary nodule in computed tomography (CT) image, which can provide useful medical knowledge for the identification of early stage lung cancer. There may exist certain relations among these attributes, and some attributes may naturally imply or boost others that have been less comprehensively exploited in previous studies. In this paper, we explicitly model the relations among 11 attributes of nodules by way of transfer learning and extract a meta-structure that captures the transferabilities across deep features of these attributes. Specifically, a higher-order transfer learning scheme is proposed by involving three phases, i.e., semantic attribute-specific modeling, semantic attributes transfer modeling and pathologic attribute generalizing, to explore the strongest association across various attributes and to boost the nodule attribute predictions in chest CT images. The proposed approach has been evaluated on the 2632 nodules in the public Lung Image Database Consortium and Image Database Resource Initiative (LIDC-IDRI) dataset. The experimental results suggest that our higher-order transfer approach shows the superior predictive performance not only in the most of the semantic attributes compared with the schemes of learning from scratch and the first-order transfer but also for the pathologic attribute compared with the related studies. In addition, we demonstrate an attribute transfer graph to reveal which attributes combination can supply the most useful information to boost the predictive performance of target attributes.
Qingfeng Wang 0004, Jun Huang 0005, Zhiqin Liu, Jie-Zhi Cheng, Qiyu Liu, Yaobin Wang, Xuehai Zhou, Chao Wang 0003
BIBM7
2016 Parallelizing Back Propagation Neural Network on Speculative Multicores
abstract
Applications typically exhibit extremely different performance characteristics depending on the accelerator. Back propagation neural network (BPNN) has been parallelized into different platforms. However, it has not yet been explored on speculative multicore architecture thoroughly. This paper presents a study of parallelizing BPNN on a speculative multicore architecture, including its speculative execution model, hardware design and programming model. The implementation was analyzed with seven well-known benchmark data sets. Furthermore, it trades off several important design factors in coming speculative multicore architecture. The experimental results show that: (1) the BPNN performs well on speculative multicore platform. It can achieve similar speedup (17.7x to 57.4x) compared with graphics processors (GPU) while provides a more friendly programmability. (2) 64 cores' computing resources can be used efficiently and 4k is the proper speculative buffer capacity in the model.
Yaobin Wang, Hong An, Zhiqin Liu, Dongmei Zhao
ICPADS1
2015 Parallelizing Block Cryptography Algorithms on Speculative Multicores
Yaobin Wang, Hong An, Zhiqin Liu, Qingfeng Wang 0004
ICA3PP (1)1
2015 Optimization and Analysis of Parallel Back Propagation Neural Network on GPU Using CUDA
Yaobin Wang, Pingping Tang, Hong An, Zhiqin Liu
ICONIP (3)1
2010 Dynamic Resource Tuning for Flexible Core Chip Multiprocessors
Yongqing Ren, Hong An, Yaobin Wang
ICA3PP (2)5
2009 The Mapping Framework and Optimizing Strategy for Block Cryptography Algorithms on Cell Broadband Engine
abstract
The Cell Broadband Engine is a typical heterogeneous chip multiprocessor which provides potential high performance for computing-intensive applications. Our researches focus on how to use Cell to speed up block cryptography applications. In this paper, we propose a mapping framework for block cryptography working in ECB mode and corresponding optimizing strategy. We take four algorithms(RC5, 3DES, AES, and Twofish) as benchmark and implement these four algorithms using Cell programming language. In order to enhance the performance, we present an optimizing strategy and evaluate the effects of the optimizing methods including compiler optimization, dual buffering, vectorization, and loop unrolling. The experiments indicate that all these four algorithms can obtain 5-20 times speedup compared with traditional processors, which shows that our mapping framework and optimizing strategy are effective for the block cryptography algorithms.
Mu Xu, Hong An, Gu Liu, Yaobin Wang, Ping Yao, Xiurui Hao, Wenting Han
PDCAT4
2007 Balancing Thread Partition for Efficiently Exploiting Speculative Thread-Level Parallelism
Yaobin Wang, Hong An, Yongqing Ren
APPT1