Yun-Dong Wu

dblp:31/5864 · also Yundong Wu · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0009-3554-5549ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 PSGF: Progressive Semantic-Guided Fusion for Ambiguity-Aware 3D Visual Grounding
Zhuangzhi Liu, Yunjing Yi, Yang Luo 0002, Yun-Dong Wu, Jinhe Su
ICIC (9)5
2025 Edge First: Edge-Guided Geometry for Superior 3D Roof Wireframe Reconstruction
abstract
Roof wireframe reconstruction has shown great success in 3D building reconstruction due to its lightweight nature and straightforward representation. However, previous methods consider all roof points, which result in edge redundancy and omissions. In this paper, we propose a novel and streamlined Edge-guided Geometric wireframe reconstruction framework, named EDGE. We find that points distributed along the roof edges make a significant contribution to the precise geometric structure of wireframe. Therefore, we design an edge point extractor (EPE) to capture the spatial relationship between points and edges, filtering out internal plane points. Moreover, we discover that the previous edge detectors rely solely on corner points, leading to error accumulation. To address this, we present the Hybrid Edge Detector (HED) feeding corner points with edge contextual features, which not only enhances edge completeness but also mitigates edge redundancy. Comprehensive experiments demonstrate EDGE outperforms existing wireframe reconstruction methods with 0.86 Corner F1-score and 0.71 Edge F1-score on Building3D dataset, striking the significant improvement of accuracy between corner and edge. Notably, our EDGE achieves a significant improvement of over 11% in Edge Recall, demonstrating the effectiveness and robustness of the proposed method.
Qiaoqiao Hao, Ting Han 0001, Yujun Liu 0005, Shangfeng Huang, Duxin Zhu, Jinhe Su, Yun-Dong Wu, Guo-Rong Cai
ICASSP7
2025 A Stereo-Wise Masking Strategy for Weakly Supervised Point Cloud Semantic Segmentation
abstract
Weakly supervised point cloud semantic segmentation (WSPCSS) has gained attention for reducing reliance on densely annotated data. However, inefficiencies in utilizing sparse annotations hinder comprehensive understanding of complex scenes. Inspired by masked autoencoder (MAE) techniques in image processing, researchers have adapted these methods to WSPCSS. Yet, current 3D masking strategies often fail to capture intricate geometric properties, resulting in generated to-be-filled content that inaccurately represents the underlying 3D scene structure. To address these limitations, this study proposes a novel stereo-wise masking strategy, which extends 2D plane masking into 3D space to generate coherent and semantically rich masked regions with contextual relevance. These regions serve as high-quality learning targets, enabling the model to better comprehend complex point cloud structures. Experimental results demonstrate that, at a 0.01 % annotation density, the proposed method achieves improvements in mIoU by 1.9 % and 4.72 % on the indoor datasets S3DIS and ScanNet V2, respectively, compared to previous methods. Furthermore, at a 0.1% annotation density on the forest dataset For-Instance, the method exhibits a 0.29 % improvement. These results substantiate the effectiveness and stability of the stereo-wise masking strategy.
Guoqing Jiang, Yun-Dong Wu, Jinhe Su
IJCNN4
2025 Semantic Uncertainty-Awared for Semantic Segmentation of Remote Sensing Images
abstract
ABSTRACT Remote sensing image segmentation is crucial for applications ranging from urban planning to environmental monitoring. However, traditional approaches struggle with the unique challenges of aerial imagery, including complex boundary delineation and intricate spatial relationships. To address these limitations, we introduce the semantic uncertainty‐aware segmentation (SUAS) method, an innovative plug‐and‐play solution designed specifically for remote sensing image analysis. SUAS builds upon the rotated multi‐scale interaction network (RMSIN) architecture and introduces the prompt refinement and uncertainty adjustment module (PRUAM). This novel component transforms original textual prompts into semantic uncertainty‐aware descriptions, particularly focusing on the ambiguous boundaries prevalent in remote sensing imagery. By incorporating semantic uncertainty, SUAS directly tackles the inherent complexities in boundary delineation, enabling more refined segmentations. Experimental results demonstrate SUAS's effectiveness, showing improvements over existing methods across multiple metrics. SUAS achieves consistent enhancements in mean intersection‐over‐union (mIoU) and precision at various thresholds, with notable performance in handling objects with irregular and complex boundaries—a persistent challenge in aerial imagery analysis. The results indicate that SUAS's plug‐and‐play design, which leverages semantic uncertainty to guide the segmentation task, contributes to improved boundary delineation accuracy in remote sensing image analysis.
Xiangfeng Qiu, Youcheng Yang, Yun-Dong Wu, Jinhe Su
IET Image Process.6
2025 CSFNet: Cross-Modal Semantic Focus Network for Semantic Segmentation of Large-Scale Point Clouds
abstract
Semantic segmentation of large-scale point clouds is an indispensable component of outdoor scene perception, providing essential 3-D semantic insights for applications in scene reconstruction, urban planning, autonomous driving, and more. However, the discriminative capability of point clouds features declines with increasing distance from the sensor, causing current methods to usually perform poorly in segmenting distant objects. To overcome this challenge and improve the differentiation between classes with similar geometric features, we propose the cross-modal semantic focus network (CSFNet). Firstly, we design a multiscale feature dynamic fusion (MDF) module to leverage multiscale image features, thereby enriching the feature representation of point clouds with additional images color and texture information. Then, in order to extract the distinguishing features of distant and different categories of objects more efficiently, we propose a semantic focus module (SFM) that employs a multiclass contrastive learning strategy to enhance feature discrimination. Finally, we introduce cross-modal knowledge distillation (KD) to augment the model’s comprehension of point clouds. Extensive experiments conducted on the SemanticKITTI and nuScenes datasets demonstrate the effectiveness of our method. Notably, our method achieves superior segmentation accuracy across multiple classes at various distances compared to current methods.
Yang Luo 0002, Ting Han 0001, Yujun Liu 0005, Jinhe Su, Yiping Chen 0002, Yun-Dong Wu, Guo-Rong Cai
IEEE Trans. Geosci. Remote. Sens.7
2021 Convertible Sparse Convolution for Point Cloud Instace Segmentation
abstract
Instance segmentation based on 3D point cloud is a key step in scene understanding. It is widely used in indoor robot navigation, outdoor autonomous driving, and other fields. But research in this area is still in its infancy. Instance segmentation not only needs to predict the semantic label of each point but also the instance label of each point. Therefore, semantic segmentation can be considered the basis of instance segmentation to some extent. Based on this motivation, we designed a voxel-based branch based on convertible sparse convolution and residual optimization modules. We design a point-based branch so that the network can maintain high-resolution representation. Then the two branches are combined to optimize the semantic segmentation results. Breadth-first search (BFS) performs well in indoor point clouds and is simple to operate. Therefore, we use this clustering operation to group the points of the same instance to obtain the instance segmentation result. The proposed method was tested on the indoor dataset Scan-Net v2 and achieved relatively good instance segmentation precision.
Jing Du 0007, Guo-Rong Cai, Zongyue Wang, Jinhe Su, Yun-Dong Wu
IGARSS5
2021 Multi-Scale Cascade Guided Object Detection in Aerial Images
abstract
Object detection in aerial images has received increasing attention during the last few years. Scale variation is one of the main challenges in large scene aerial images. Existing object detection pipelines usually detect objects of different scale objects at multiple scale layers. However, the conventional detection approaches with multi-scale density predictions could cause duplicate detections of the same object. In this paper, we proposed a multi-scale cascade guided detection framework (MCGNet) to address these issues by guiding the different scales in detector focus on different scale objects. In particular, we proposed a multi-scale cascade module to predict the different scale objects with an explicit constraint in the loss function. Experiments on benchmark DOTA show promising performance of MCGNet compared with other detectors. Code will be released at https://github.com/jason-su/MCGNET.
Jiajia Liao, Yingchao Piao, Guo-Rong Cai, Yun-Dong Wu, Jinhe Su
IGARSS4
2021 Phosphate binding sites prediction in phosphorylation-dependent protein-protein interactions
abstract
MOTIVATION: Phosphate binding plays an important role in modulating protein-protein interactions, which are ubiquitous in various biological processes. Accurate prediction of phosphate binding sites is an important but challenging task. Small size and diversity of phosphate binding sites lead to a substantial challenge for developing accurate prediction methods. RESULTS: Here, we present the phosphate binding site predictor (PBSP), a novel and accurate approach to identifying phosphate binding sites from protein structures. PBSP combines an energy-based ligand-binding sites identification method with reverse focused docking using a phosphate probe. We show that PBSP outperforms not only general ligand binding sites predictors but also other existing phospholigand-specific binding sites predictors. It achieves ∼95% success rate for top 10 predicted sites with an average Matthews correlation coefficient value of 0.84 for successful predictions. PBSP can accurately predict phosphate binding modes, with average position error of 1.4 and 2.4 Å in bound and unbound datasets, respectively. Lastly, visual inspection of the predictions is conducted. Reasons for failed predictions are further analyzed and possible ways to improve the performance are provided. These results demonstrate a novel and accurate approach to phosphate binding sites identification in protein structures. AVAILABILITY AND IMPLEMENTATION: The software and benchmark datasets are freely available at http://web.pkusz.edu.cn/wu/PBSP/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zheng-Chang Lu, Yun-Dong Wu
Bioinform.3
2020 CRiSP: accurate structure prediction of disulfide-rich peptides with cystine-specific sequence alignment and machine learning
abstract
MOTIVATION: High-throughput sequencing discovers many naturally occurring disulfide-rich peptides or cystine-rich peptides (CRPs) with diversified bioactivities. However, their structure information, which is very important to peptide drug discovery, is still very limited. RESULTS: We have developed a CRP-specific structure prediction method called Cystine-Rich peptide Structure Prediction (CRiSP), based on a customized template database with cystine-specific sequence alignment and three machine-learning predictors. The modeling accuracy is significantly better than several popular general-purpose structure modeling methods, and our CRiSP can provide useful model quality estimations. AVAILABILITY AND IMPLEMENTATION: The CRiSP server is freely available on the website at http://wulab.com.cn/CRISP. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zi-Lin Liu, Jing-Hao Hu, Yun-Dong Wu
Bioinform.4
2020 IDRMutPred: predicting disease-associated germline nonsynonymous single nucleotide variants (nsSNVs) in intrinsically disordered regions
abstract
MOTIVATION: Despite of the lack of folded structure, intrinsically disordered regions (IDRs) of proteins play versatile roles in various biological processes, and many nonsynonymous single nucleotide variants (nsSNVs) in IDRs are associated with human diseases. The continuous accumulation of nsSNVs resulted from the wide application of NGS has driven the development of disease-association prediction methods for decades. However, their performance on nsSNVs in IDRs remains inferior, possibly due to the domination of nsSNVs from structured regions in training data. Therefore, it is highly demanding to build a disease-association predictor specifically for nsSNVs in IDRs with better performance. RESULTS: We present IDRMutPred, a machine learning-based tool specifically for predicting disease-associated germline nsSNVs in IDRs. Based on 17 selected optimal features that are extracted from sequence alignments, protein annotations, hydrophobicity indices and disorder scores, IDRMutPred was trained using three ensemble learning algorithms on the training dataset containing only IDR nsSNVs. The evaluation on the two testing datasets shows that all the three prediction models outperform 17 other popular general predictors significantly, achieving the ACC between 0.856 and 0.868 and MCC between 0.713 and 0.737. IDRMutPred will prioritize disease-associated IDR germline nsSNVs more reliably than general predictors. AVAILABILITY AND IMPLEMENTATION: The software is freely available at http://www.wdspdb.com/IDRMutPred. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jing-Bo Zhou, Yao Xiong, Zhi-Qiang Ye, Yun-Dong Wu
Bioinform.5
2019 WDSPdb: an updated resource for WD40 proteins
abstract
SUMMARY: The WD40-repeat proteins are a large family of scaffold molecules that assemble complexes in various cellular processes. Obtaining their structures is the key to understanding their interaction details. We present WDSPdb 2.0, a significantly updated resource providing accurately predicted secondary and tertiary structures and featured sites annotations. Based on an optimized pipeline, WDSPdb 2.0 contains about 600 thousand entries, an increase of 10-fold, and integrates more than 37 000 variants from sources of ClinVar, Cosmic, 1000 Genomes, ExAC, IntOGen, cBioPortal and IntAct. In addition, the web site is largely improved for visualization, exploring and data downloading. AVAILABILITY AND IMPLEMENTATION: http://www.wdspdb.com/wdsp/ or http://wu.scbb.pkusz.edu.cn/wdsp/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jing-Bo Zhou, Nuo-Si Wu, Zhi-Qiang Ye, Yun-Dong Wu
Bioinform.7
2019 Cover patches: A general feature extraction strategy for spoofing detection
abstract
Summary Face anti‐spoofing has attracted many attentions in security applications, such as mobile payment and entrance guard. Until now, face anti‐spoofing technique is still a challenging task. Mainstream image‐based spoofing algorithms usually use global motion or texture information to distinguish whether an input face is live or fake. However, the performance of these methods are sensitive in light changes, or images acquired from different sensors. The main reason is that spoofed face image always has slight different texture in local areas, such as landmark or salient region of face. To this end, this paper proposes a novel multi‐patches feature extraction strategy to detect spoofing. First, a set of patches with specific combination scheme is selected to cover the face image. Second, features such as hand‐crafted Gray Level Co‐occurrence Matrix (GLCM), Local Binary Patterns (LBP), or deep features are extracted from these patches. Third, all features are combined as the global descriptor of the face image, then fed into an SVM classifier to verify the anti‐spoofing detection. Experimental results show that the proposed strategy can effectively enhance the performance, concerning with the accuracy of spoofed face detection in four widely used anti‐spoofing databases.
Guo-Rong Cai, Songzhi Su, Chengcai Leng, Jipeng Wu, Yun-Dong Wu, Shaozi Li
Concurr. Comput. Pract. Exp.5
2018 Combining 2D and 3D features to improve road detection based on stereo cameras
abstract
Road detection is a fundamental component of autonomous driving systems since it provides validspace and candidate regions of objects for driving decision. The core of roaddetection methods is extracting effective and discriminative features. Sincetwo‐dimensional (2D) and 3D features are complementary, the authors propose arobust multi‐feature combination and optimisation framework for stereo imagepairs, called Feature++. First, several 2D and 3D features such as Gabor andplane are, respectively, extracted after the generation of 2D super‐pixel and a3D depth image from stereo matching. Second, the combined features are fed intoa three‐layer shallow neural network classifier to decide whether a super‐pixelis road region or not. Finally, the classified results are further refined usingfully connected conditional random field (CRF), taking the content informationinto consideration. We extensively evaluate the performance of four 2D features,four 3D features, and their combinations. Experiments conducted on the KITTIROAD benchmark show that (i) the combinations of 2D and 3D features greatlyimprove the road detection performance and (ii) using CRF as a refinement stepis necessary. Overall, their proposed ‘Feature + +’ method outperforms mostmanually designed features, and is comparable with state‐of‐the‐art methods thatare based on deep learning methods.
Guo-Rong Cai, Songzhi Su, Wenli He, Yun-Dong Wu, Shaozi Li
IET Comput. Vis.4
2016 Spectral-spatial co-clustering of hyperspectral image data based on bipartite graph
Wei Liu 0005, Shaozi Li, Xianming Lin, Yun-Dong Wu, Rongrong Ji
Multim. Syst.4
2015 Novel Graph Cuts Method for Multi-Frame Super-Resolution
abstract
In this letter, we propose a new graph cuts multi-frame super resolution method. The method is carried out in 3 steps. First, we project each high-resolution pixel p onto the low-resolution images and select low-resolution pixels which fall within the zone of influence of p. Second, we weigh the contribution of the low-resolution pixels via a soft switching function and add them to construct a virtual low resolution pixel. The high resolution image is then recovered after minimizing a Maximum a posteriori Markov Random Field (MAP-MRF) energy function. This is done by approximating our energy function to make it graph representable and minimize it with a graph cuts α-expansion algorithm. Experimental results show that our approach outperforms state-of-the-art methods.
Dongxiao Zhang, Pierre-Marc Jodoin, Cuihua Li, Yun-Dong Wu, Guo-Rong Cai
IEEE Signal Process. Lett.4
2013 Perspective-SIFT: An efficient tool for low-altitude remote sensing image registration
Guo-Rong Cai, Pierre-Marc Jodoin, Shaozi Li, Yun-Dong Wu, Songzhi Su, Zhenkun Huang
Signal Process.4