VLDB 2026 Research / reviewers in the wild / expert
Ping-An He 0001
dblp:33/1529 · also Ping-an He 0001, Pingan He 0001
· DBLP profile ↗
8ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0003-3749-1483ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Castl: robust identification of spatially variable genes in spatial transcriptomics via an ensemble-based frameworkabstractSpatially variable genes (SVGs) are essential for elucidating tissue organization within spatially resolved transcriptomics. While a number of computational methods have been developed for SVG identification, their reliance on algorithm-specific assumptions, such as predefined kernel functions or spatial neighborhood graphs, often results in substantial variability in sensitivity and inflated false discovery rates (FDRs) across heterogeneous datasets. To address this challenge, we here develop Castl, an ensemble-based framework for SVG identification that integrates multiple detection methods through statistically designed aggregation modules. Comprehensive evaluations on both simulated and real-world data demonstrate that Castl consistently identifies biologically meaningful spatial expression patterns, mitigates method-specific biases and effectively controls FDRs across various biological contexts, resolutions, and spatial technologies. This flexible, assumption-free framework offers a robust and standardized foundation for spatially informed feature discovery in complex biological systems. Yiyi Yu, Ping-An He 0001, Xiaoqi Zheng |
Briefings Bioinform. | 3 |
| 2025 | Predicting drug-target interactions based on multivariate information fusion and graph contrast learning
Siying Yang, Ping-An He 0001, Pan Zeng, Yajie Meng, Feifei Cui, Yuhua Yao, Jialiang Yang, Junlin Xu |
J. Biomed. Informatics | 2 |
| 2024 | Precision DNA methylation typing via hierarchical clustering of Nanopore current signals and attention-based neural networkabstractDecoding DNA methylation sites through nanopore sequencing has emerged as a cutting-edge technology in the field of DNA methylation research, as it enables direct sequencing of native DNA molecules without the need for prior enzymatic or chemical treatments. During nanopore sequencing, methylation modifications on DNA bases cause changes in electrical current intensity. Therefore, constructing deep neural network models to decode the electrical signals of nanopore sequencing has become a crucial step in methylation site identification. In this study, we utilized nanopore sequencing data containing diverse DNA methylation types and motif sequence diversity. We proposed a feature encoding method based on current signal clustering and leveraged the powerful attention mechanism in the Transformer framework to construct the PoreFormer model for identifying DNA methylation sites in nanopore sequencing. The model demonstrated excellent performance under conditions of multi-class methylation and motif sequence diversity, offering new insights into related research fields. Wen-Jing Yi, Jia-Ning Zhao, Wei Zhang 0256, Ping-An He 0001, Xiao-Qing Liu, Ying-Feng Zheng, Zhuoxing Shi |
Briefings Bioinform. | 6 |
| 2022 | Prediction of lncRNA-disease association based on a Laplace normalized random walk with restart algorithm on heterogeneous networksabstractBACKGROUND: More and more evidence showed that long non-coding RNAs (lncRNAs) play important roles in the development and progression of human sophisticated diseases. Therefore, predicting human lncRNA-disease associations is a challenging and urgently task in bioinformatics to research of human sophisticated diseases. RESULTS: In the work, a global network-based computational framework called as LRWRHLDA were proposed which is a universal network-based method. Firstly, four isomorphic networks include lncRNA similarity network, disease similarity network, gene similarity network and miRNA similarity network were constructed. And then, six heterogeneous networks include known lncRNA-disease, lncRNA-gene, lncRNA-miRNA, disease-gene, disease-miRNA, and gene-miRNA associations network were applied to design a multi-layer network. Finally, the Laplace normalized random walk with restart algorithm in this global network is suggested to predict the relationship between lncRNAs and diseases. CONCLUSIONS: The ten-fold cross validation is used to evaluate the performance of LRWRHLDA. As a result, LRWRHLDA achieves an AUC of 0.98402, which is higher than other compared methods. Furthermore, LRWRHLDA can predict isolated disease-related lnRNA (isolated lnRNA related disease). The results for colorectal cancer, lung adenocarcinoma, stomach cancer and breast cancer have been verified by other researches. The case studies indicated that our method is effective. Liugen Wang, Min Shang, Ping-An He 0001 |
BMC Bioinform. | 4 |
| 2021 | Understanding the unimodal distributions of cancer occurrence rates: it takes two factors for a cancer to occurabstractData from the SEER reports reveal that the occurrence rate of a cancer type generally follows a unimodal distribution over age, peaking at an age that is cancer-type specific and ranges from 30+ through 70+. Previous studies attribute such bell-shaped distributions to the reduced proliferative potential in senior years but fail to explain why some cancers have their occurrence peak at 30+ or 40+. We present a computational model to offer a new explanation to such distributions. The model uses two factors to explain the observed age-dependent cancer occurrence rates: cancer risk of an organ and the availability level of the growth signals in circulation needed by a cancer type, with the former increasing and the latter decreasing with age. Regression analyses were conducted of known occurrence rates against such factors for triple negative breast cancer, testicular cancer and cervical cancer; and all achieved highly tight fitting results, which were also consistent with clinical, gene-expression and cancer-drug data. These reveal a fundamentally important relationship: while cancer is driven by endogenous stressors, it requires sufficient levels of exogenous growth signals to happen, hence suggesting the realistic possibility for treating cancer via cleaning out the growth signals in circulation needed by a cancer. Zheng An, Renbo Tan, Ping-An He 0001, Jingjing Jing |
Briefings Bioinform. | 4 |
| 2020 | 2SigFinder: the combined use of small-scale and large-scale statistical testing for genomic island detection from a single genomeabstractBACKGROUND: Genomic islands are associated with microbial adaptations, carrying genomic signatures different from the host. Some methods perform an overall test to identify genomic islands based on their local features. However, regions of different scales will display different genomic features. RESULTS: We proposed here a novel method "2SigFinder ", the first combined use of small-scale and large-scale statistical testing for genomic island detection. The proposed method was tested by genomic island boundary detection and identification of genomic islands or functional features of real biological data. We also compared the proposed method with the comparative genomics and composition-based approaches. The results indicate that the proposed 2SigFinder is more efficient in identifying genomic islands. CONCLUSIONS: From real biological data, 2SigFinder identified genomic islands from a single genome and reported robust results across different experiments, without annotated information of genomes or prior knowledge from other datasets. 2SigHunter identified 25 Pathogenicity, 1 tRNA, 2 Virulence and 2 Repeats from 27 Pathogenicity, 1 tRNA, 2 Virulence and 2 Repeats, and detected 101 Phage and 28 HEG out of 130 Phage and 36 HEGs in S. enterica Typhi CT18, which shows that it is more efficient in detecting functional features associated with GIs. Xinnan Xu, Ping-An He 0001, Michael Q. Zhang |
BMC Bioinform. | 4 |
| 2017 | Matrix completion with side information and its applications in predicting the antigenicity of influenza virusesabstractMOTIVATION: Low-rank matrix completion has been demonstrated to be powerful in predicting antigenic distances among influenza viruses and vaccines from partially revealed hemagglutination inhibition table. Meanwhile, influenza hemagglutinin (HA) protein sequences are also effective in inferring antigenic distances. Thus, it is natural to integrate HA protein sequence information into low-rank matrix completion model to help infer influenza antigenicity, which is critical to influenza vaccine development. RESULTS: We have proposed a novel algorithm called biological matrix completion with side information (BMCSI), which first measures HA protein sequence similarities among influenza viruses (especially on epitopes) and then integrates the similarity information into a low-rank matrix completion model to predict influenza antigenicity. This algorithm exploits both the correlations among viruses and vaccines in serological tests and the power of HA sequence in predicting influenza antigenicity. We applied this model into H3N2 seasonal influenza virus data. Comparing to previous methods, we significantly reduced the prediction root-mean-square error in a 10-fold cross validation analysis. Based on the cartographies constructed from imputed data, we showed that the antigenic evolution of H3N2 seasonal influenza is generally S-shaped while the genetic evolution is half-circle shaped. We also showed that the Spearman correlation between genetic and antigenic distances (among antigenic clusters) is 0.83, demonstrating a globally high correspondence and some local discrepancies between influenza genetic and antigenic evolution. Finally, we showed that 4.4%±1.2% genetic variance (corresponding to 3.11 ± 1.08 antigenic distances) caused an antigenic drift event for H3N2 influenza viruses historically. AVAILABILITY AND IMPLEMENTATION: The software and data for this study are available at http://bi.sky.zstu.edu.cn/BMCSI/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xianhong Li, Yuhua Yao, Fayou Wang, Jiasheng Yang, Hailiang Sun 0001, Ping-An He 0001, Jialiang Yang |
Bioinform. | 11 |
| 2013 | Comparison study on statistical features of predicted secondary structures for protein structural class prediction: From content to positionabstractBACKGROUND: Many content-based statistical features of secondary structural elements (CBF-PSSEs) have been proposed and achieved promising results in protein structural class prediction, but until now position distribution of the successive occurrences of an element in predicted secondary structure sequences hasn't been used. It is necessary to extract some appropriate position-based features of the secondary structural elements for prediction task. RESULTS: We proposed some position-based features of predicted secondary structural elements (PBF-PSSEs) and assessed their intrinsic ability relative to the available CBF-PSSEs, which not only offers a systematic and quantitative experimental assessment of these statistical features, but also naturally complements the available comparison of the CBF-PSSEs. We also analyzed the performance of the CBF-PSSEs combined with the PBF-PSSE and further constructed a new combined feature set, PBF11CBF-PSSE. Based on these experiments, novel valuable guidelines for the use of PBF-PSSEs and CBF-PSSEs were obtained. CONCLUSIONS: PBF-PSSEs and CBF-PSSEs have a compelling impact on protein structural class prediction. When combining with the PBF-PSSE, most of the CBF-PSSEs get a great improvement over the prediction accuracies, so the PBF-PSSEs and the CBF-PSSEs have to work closely so as to make significant and complementary contributions to protein structural class prediction. Besides, the proposed PBF-PSSE's performance is extremely sensitive to the choice of parameter k. In summary, our quantitative analysis verifies that exploring the position information of predicted secondary structural elements is a promising way to improve the abilities of protein structural class prediction. Xiao-Qing Liu, Yuhua Yao, Yunjie Cao, Ping-An He 0001 |
BMC Bioinform. | 6 |