Yong Wang 0057

dblp:84/2694-57 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-8326-2924ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Hierarchical Hybrid Transformer for Sparse-View X-Ray 3D Reconstruction
abstract
With the continuous advancement of medical imaging technology, sparse-view X-ray 3D reconstruction has found widespread applications in low-dose imaging and rapid scanning. However, traditional reconstruction methods have limitations in capturing details and expressing complex structures. To address this issue, we propose a sparse-view X-ray 3D reconstruction method based on a hierarchical hybrid Transformer. This research uses NeRF as the core framework, combining sparse-view X-ray images. First, a module combining focal attention and multi-scale mixing is employed to partition focal regions, extract multi-scale features, and perform channel attention-based fusion, effectively integrating key regions and hierarchical structural information in the image. This enables the collaborative representation of fine details and global semantics. To address complex anatomical structures such as bones in chest, jaw and foot, as well as soft tissues like aneurism and pancreas, a hierarchical hybrid Transformer module is introduced to model dependencies between local and global regions, enhancing the ability to capture fine details and represent overall structure. Experimental results show that the proposed method achieves significant performance improvements in sparse-view X-ray 3D reconstruction tasks across various medical structures.
Pengbo Zhou, Yong Wang 0057, Wuyang Shui
IEEE Signal Process. Lett.4
2026 Dual-Branch Feature Fusion for Sparse-View X-Ray 3D Reconstruction
abstract
With the rapid development of medical imaging, sparse-view X-ray 3D reconstruction has become an essential technique for addressing low-dose X-ray imaging challenges. However, due to sparse angular sampling, traditional reconstruction methods often face challenges in handling complex bone and soft tissue structures, leading to information loss and insufficient detail capture. To address these issues, this paper proposes a sparse-view X-ray 3D reconstruction method based on Neural Radiance Fields (NeRF) with a dual-branch feature fusion framework. By synergistically extracting local and global features, this approach enhances the reconstruction of intricate bone and soft tissue structures. For specific applications in regions like the pelvis and aneurism, the method employs depthwise separable convolutions in the local branch to efficiently capture X-ray image details, enhancing the reconstruction of complex bone structures. In the global branch, a pooling Transformer with window mechanisms and hybrid positional encoding is introduced to capture the global features of soft tissue structures like aneurism. Experimental results demonstrate the superiority of this method on multiple medical imaging datasets, particularly in reconstructing complex bone regions and recovering details of soft tissue structures, compared to traditional methods and existing deep learning models.
Yong Wang 0057, Guohua Geng, Wen Tang 0004
IEEE Signal Process. Lett.1
2024 Low-Overlap Point Cloud Registration With Transformer
abstract
In real-world scenarios, due to factors like sensor noise, point cloud data often exhibits low overlap, posing challenges for traditional registration methods. To address this issue, we propose a low-overlap point cloud registration with Transformer. This algorithm employs a dynamic positional encoding strategy that adaptively computes position encodings for each point based on its distribution. This enables better capturing of richer spatial relationships between point clouds and facilitates adaptation to diverse point cloud distributions across various scenes. Furthermore, we combine the mechanisms of self-attention and graph convolutions. The self-attention mechanism captures global dependencies among points, while the graph convolutions capture local neighborhood information between points. Lastly, in the context of cross-attention, adaptive weights are introduced during the attention calculation process. This involves multiplying attention scores by adaptive weights, enhancing the model's ability to focus on crucial registration areas. In scenarios with low overlap, this algorithm significantly enhances the success rate of successful registrations. It achieves notable improvements and attains a new state-of-the-art performance in the 3DLoMatch benchmark test, reaching a registration recall rate of 71.7%.
Yong Wang 0057, Pengbo Zhou, Guohua Geng, Qi Zhang 0091
IEEE Signal Process. Lett.1
2024 Neighborhood Multi-Compound Transformer for Point Cloud Registration
abstract
Point cloud registration is a critical issue in 3D reconstruction and computer vision, particularly challenging in cases of low overlap and different datasets, where algorithm generalization and robustness are pressing challenges. In this paper, we propose a point cloud registration algorithm called Neighborhood Multi-compound Transformer (NMCT). To capture local information, we introduce Neighborhood Position Encoding for the first time. By employing a nearest neighbor approach to select spatial points, this encoding enhances the algorithm’s ability to extract relevant local feature information and local coordinate information from dispersed points within the point cloud. Furthermore, NMCT utilizes the Multi-compound Transformer as the interaction module for point cloud information. In this module, the Spatial Transformer phase engages in local-global fusion learning based on Neighborhood Position Encoding, facilitating the extraction of internal features within the point cloud. The Temporal Transformer phase, based on Neighborhood Position Encoding, performs local position-local feature interaction, achieving local and global interaction between two point cloud. The combination of these two phases enables NMCT to better address the complexity and diversity of point cloud data. The algorithm is extensively tested on different datasets (3DMatch, ModelNet, KITTI, MVP-RG), demonstrating outstanding generalization and robustness.
Yong Wang 0057, Pengbo Zhou, Guohua Geng, Kang Li 0005, Ruoxue Li
IEEE Trans. Circuits Syst. Video Technol.1
2024 MATR: Multicompound Adaptive Transformer for Point Cloud Registration
abstract
Point cloud registration plays a key role in the fields of computer vision, particularly in scenarios with low overlap, large scenes, different datasets, where difficulties, such as difficulty matching, scale changes and geometric deformations, local feature loss are commonly encountered. In this article, we propose a point cloud registration algorithm named multicompound adaptive transformer, which introduces adaptive position encoding, dynamically adjusting the local coordinates and feature information of scattered points within the point cloud through an adaptive threshold enhancement mechanism. Simultaneously, the multicompound transformer is introduced. In the spatial transformer stage, it accomplishes the local position-local feature interaction of individual point clouds through adaptive position encoding. Then, in the temporal transformer stage, it achieves local–local interaction and local–global information interaction between two point clouds through a dual-branch multiscale transformer. Through experiments on different datasets, we validate the algorithm's superior generalization performance in scenarios with low overlap, large scenes, and different datasets.
Yong Wang 0057, Pengbo Zhou, Guohua Geng, Kang Li 0005
IEEE Trans. Ind. Informatics1
2023 SparseFormer: Sparse transformer network for point cloud classification
Yong Wang 0057, Pengbo Zhou, Guohua Geng, Qi Zhang 0091
Comput. Graph.1
2022 Encrypted speech retrieval based on long sequence Biohashing
Yi-Bo Huang 0001, Yong Wang 0057
Multim. Tools Appl.2
2021 A high security BioHashing encrypted speech retrieval algorithm based on feature fusion
Yi-Bo Huang 0001, Yong Wang 0057, Yi-rong Xie
Multim. Tools Appl.3
2021 Multi-format speech BioHashing based on energy to zero ratio and improved LP-MMSE parameter fusion
Yong Wang 0057, Yi-Bo Huang 0001
Multim. Tools Appl.1
2020 Multi-format speech BioHashing based on spectrogram
Yi-Bo Huang 0001, Yong Wang 0057, Wei-zhao Zhang, Manhong Fan
Multim. Tools Appl.2