VLDB 2026 Research / reviewers in the wild / expert
Yingfu Wu
dblp:295/7043
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModalabstractMOTIVATION: Advances in spatial transcriptomics (ST) technologies have made it possible to jointly acquire gene expression and histological image information while preserving spatial coordinates. This breakthrough presents unprecedented opportunities for the precise dissection of spatial heterogeneity in complex tissues. However, existing computational methods remain limited in their capacity for effective integration and synergistic modelling of multimodal ST data. RESULTS: We propose SpatialModal, a multimodal graph learning framework that learns robust joint representations by combining a hierarchical representation strategy with a dual-level contrastive learning mechanism. We perform extensive validation of SpatialModal across diverse ST datasets spanning human and mouse tissues. The results demonstrate that SpatialModal effectively reveals intricate brain architectures in humans and mice, dissects tumour microenvironment heterogeneity in breast cancer, delineates Alzheimer's disease patterns, and characterizes spatiotemporal developmental trajectories within the embryonic heart, underscoring its capability to decipher the spatial heterogeneity of biological tissues. Furthermore, SpatialModal exhibits remarkable versatility and robustness, maintaining superior efficacy even on unimodal datasets devoid of histological images, thereby ensuring its broad applicability across diverse ST platforms. AVAILABILITY AND IMPLEMENTATION: SpatialModal is implemented in Python and is freely available at https://github.com/xingyili/SpatialModal. The source code used in this study has been archived on Zenodo at DOI: https://doi.org/10.5281/zenodo.21264356. All datasets used in this study are publicly available at https://doi.org/10.5281/zenodo.18220735. Xingyi Li 0003, Dongmin Zhao, Xiangting Jia, Gaoyuan Du, Jialuo Xu, Yingfu Wu, Junnan Zhu, Xuequn Shang 0001 |
Bioinform. | 8 |
| 2025 | CREATE: a novel attention-based framework for efficient classification of transposable elementsabstractTransposable elements (TEs) are DNA sequences that can move within a genome. They constitute a substantial portion of the eukaryotic genome and play essential roles in gene regulation and genome evolution. Accurate classification of these repetitive elements is crucial for investigating their potential impact on the genome. Over the past few decades, several alignment-based tools have been developed to annotate TE types. While these methods rely heavily on prior knowledge and are often computationally expensive, machine learning-based approaches have been proposed to overcome these limitations. However, most of these approaches fail to capture the multiscale features of TEs, resulting in suboptimal performance. Here, we propose a novel framework called CREATE, which simultaneously integrates the global pattern distribution and the local sequence profile of TEs using Convolutional neural networks and Recurrent neural nEtworks with an Attention mechanism for efficient TE classification. Due to the hierarchical structure of TE groups, we trained nine classifiers corresponding to parent nodes within the class hierarchy. We further applied a top-down hierarchical classification strategy to achieve a more complete classification of unknown TEs. Comprehensive experiments demonstrate that CREATE outperforms existing TE-type annotation methods and achieves superior performance in hierarchical classification tasks. In conclusion, CREATE exhibits great potential for improving the accuracy of TE annotation. The source code and demo data are available at https://github.com/yangqi-cs/CREATE. Yingfu Wu, Meihong Gao, Fuhao Zhang, Xingyu Liao, Xuequn Shang 0001 |
Briefings Bioinform. | 3 |
| 2022 | scHiCSC: A Novel Single-Cell Hi-C Clustering Framework by Contact-Weight-Based Smoothing and Feature FusionabstractSingle-cell Hi-C technology is utilized to obtain chromosome interaction information at the single-cell level and further study the differences in genome structures between different cell types. However, there are few accurate and efficient clustering methods for single-cell Hi-C data, with the following manifestations: The clustering efficacy on the dataset with a large number of cells is not very satisfactory, and it is difficult to cluster these cells with small number in the whole dataset. In this study, we propose a high-performance single-cell Hi-C clustering framework, called scHiCSC. A new smoothing method based on contact number weight is first proposed to generate cell embedding with more accurate cell features. In addition, a novel feature fusion method is proposed to further supplement the feature information of cells by fusing the chromosome structure information within cells and the distance information between cells. The experimental results show that scHiCSC has a strong generalization ability on different sizes of datasets and outperforms the existing single-cell Hi-C clustering frameworks. Moreover, scHiCSC achieves an optimal and stable clustering efficacy in the datasets with large-scale cell numbers and can cluster the cells with small number in the whole dataset. The source code of scHiCSC is freely available at https://github.com/HaoWuLab-Bioinformatics/scHiCSC. Xiangfei Zhou, Zhenqi Shi, Yingfu Wu, Hao Wu 0062 |
BIBM | 3 |
| 2022 | scHiCStackL: a stacking ensemble learning-based method for single-cell Hi-C classification using cell embeddingabstractSingle-cell Hi-C data are a common data source for studying the differences in the three-dimensional structure of cell chromosomes. The development of single-cell Hi-C technology makes it possible to obtain batches of single-cell Hi-C data. How to quickly and effectively discriminate cell types has become one hot research field. However, the existing computational methods to predict cell types based on Hi-C data are found to be low in accuracy. Therefore, we propose a high accuracy cell classification algorithm, called scHiCStackL, based on single-cell Hi-C data. In our work, we first improve the existing data preprocessing method for single-cell Hi-C data, which allows the generated cell embedding better to represent cells. Then, we construct a two-layer stacking ensemble model for classifying cells. Experimental results show that the cell embedding generated by our data preprocessing method increases by 0.23, 1.22, 1.46 and 1.61$\%$ comparing with the cell embedding generated by the previously published method scHiCluster, in terms of the Acc, MCC, F1 and Precision confidence intervals, respectively, on the task of classifying human cells in the ML1 and ML3 datasets. When using the two-layer stacking ensemble framework with the cell embedding, scHiCStackL improves by 13.33, 19, 19.27 and 14.5 over the scHiCluster, in terms of the Acc, ARI, NMI and F1 confidence intervals, respectively. In summary, scHiCStackL achieves superior performance in predicting cell types using the single-cell Hi-C data. The webserver and source code of scHiCStackL are freely available at http://hww.sdu.edu.cn:8002/scHiCStackL/ and https://github.com/HaoWuLab-Bioinformatics/scHiCStackL, respectively. Hao Wu 0062, Yingfu Wu, Haoru Zhou, Zhongli Chen, Yi Xiong 0002, Quanzhong Liu, Hongming Zhang 0002 |
Briefings Bioinform. | 2 |
| 2022 | CLNN-loop: a deep learning model to predict CTCF-mediated chromatin loops in the different cell lines and CTCF-binding sites (CBS) pair typesabstractMOTIVATION: Three-dimensional (3D) genome organization is of vital importance in gene regulation and disease mechanisms. Previous studies have shown that CTCF-mediated chromatin loops are crucial to studying the 3D structure of cells. Although various experimental techniques have been developed to detect chromatin loops, they have been found to be time-consuming and costly. Nowadays, various sequence-based computational methods can capture significant features of 3D genome organization and help predict chromatin loops. However, these methods have low performance and poor generalization ability in predicting chromatin loops. RESULTS: Here, we propose a novel deep learning model, called CLNN-loop, to predict chromatin loops in different cell lines and CTCF-binding sites (CBS) pair types by fusing multiple sequence-based features. The analysis of a series of examinations based on the datasets in the previous study shows that CLNN-loop has satisfactory performance and is superior to the existing methods in terms of predicting chromatin loops. In addition, we apply the SHAP framework to interpret the predictions of different models, and find that CTCF motif and sequence conservation are important signs of chromatin loops in different cell lines and CBS pair types. AVAILABILITY AND IMPLEMENTATION: The source code of CLNN-loop is freely available at https://github.com/HaoWuLab-Bioinformatics/CLNN-loop and the webserver of CLNN-loop is freely available at http://hwclnn.sdu.edu.cn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yingfu Wu, Haoru Zhou, Hongming Zhang 0002, Hao Wu 0062 |
Bioinform. | 2 |
| 2021 | Strategies of attack-defense game for wireless sensor networks considering the effect of confidence level in fuzzy environment
Yingfu Wu, Bingyi Kang, Hao Wu 0062 |
Eng. Appl. Artif. Intell. | 1 |