VLDB 2026 Research / reviewers in the wild / expert
Hao Lv 0007
dblp:27/5239-7
· DBLP profile ↗
13ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-7580-0155ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Navigating the 3D genome at single-cell resolution: techniques, computation, and mechanistic landscapesabstractThe 3D organization of the genome is critical for gene expression regulation, cellular identity, and disease progression. Traditional methods that analyze bulk genomic data often obscure cell-to-cell heterogeneity, limiting the resolution of intrinsic variability within complex biological systems. To overcome this, single-cell 3D genomics has emerged, revealing chromatin architecture at the individual cell level. Advanced experimental approaches enable genome-wide chromatin contact mapping, while computational frameworks reconstruct dynamic chromatin topologies from high-dimensional data. Building on these breakthroughs, recent advances in single-cell 3D genomics have led to transformative progress in epigenetics, linking 3D genome architecture with gene regulation, cellular identity, and disease phenotypes. This review focuses on the breakthroughs in single-cell 3D genomics, demonstrating how integrated experimental, computational, and mechanistic approaches decode chromatin architecture. These insights have deepened the understanding of genome function at the single-cell level and lay the foundation for future advances in precision medicine and topology-guided therapeutic strategies. Feitong Hong, Kaiyuan Han, Yuduo Hao, Xueqin Xie, Qiuming Chen, Yijie Wei, Xinwei Luo, Sijia Xie, Benjamin Lebeau, Crystal Ling, Hao Lv 0007, Hao Lin 0001, Fu-Ying Dao |
Briefings Bioinform. | 13 |
| 2022 | iRice-MS: An integrated XGBoost model for detecting multitype post-translational modification sites in riceabstractPost-translational modification (PTM) refers to the covalent and enzymatic modification of proteins after protein biosynthesis, which orchestrates a variety of biological processes. Detecting PTM sites in proteome scale is one of the key steps to in-depth understanding their regulation mechanisms. In this study, we presented an integrated method based on eXtreme Gradient Boosting (XGBoost), called iRice-MS, to identify 2-hydroxyisobutyrylation, crotonylation, malonylation, ubiquitination, succinylation and acetylation in rice. For each PTM-specific model, we adopted eight feature encoding schemes, including sequence-based features, physicochemical property-based features and spatial mapping information-based features. The optimal feature set was identified from each encoding, and their respective models were established. Extensive experimental results show that iRice-MS always display excellent performance on 5-fold cross-validation and independent dataset test. In addition, our novel approach provides the superiority to other existing tools in terms of AUC value. Based on the proposed model, a web server named iRice-MS was established and is freely accessible at http://lin-group.cn/server/iRice-MS. Hao Lv 0007, Yang Zhang 0125, Jia-Shu Wang, Shi-Shi Yuan, Fu-Ying Dao, Zheng-Xing Guan, Hao Lin 0001, Ke-Jun Deng |
Briefings Bioinform. | 1 |
| 2022 | PSnoD: identifying potential snoRNA-disease associations based on bounded nuclear norm regularizationabstractMany studies have proved that small nucleolar RNAs (snoRNAs) play critical roles in the development of various human complex diseases. Discovering the associations between snoRNAs and diseases is an important step toward understanding the pathogenesis and characteristics of diseases. However, uncovering associations via traditional experimental approaches is costly and time-consuming. This study proposed a bounded nuclear norm regularization-based method, called PSnoD, to predict snoRNA-disease associations. Benchmark experiments showed that compared with the state-of-the-art methods, PSnoD achieved a superior performance in the 5-fold stratified shuffle split. PSnoD produced a robust performance with an area under receiver-operating characteristic of 0.90 and an area under precision-recall of 0.55, highlighting the effectiveness of our proposed method. In addition, the computational efficiency of PSnoD was also demonstrated by comparison with other matrix completion techniques. More importantly, the case study further elucidated the ability of PSnoD to screen potential snoRNA-disease associations. The code of PSnoD has been uploaded to https://github.com/linDing-groups/PSnoD. Based on PSnoD, we established a web server that is freely accessed via http://psnod.lin-group.cn/. Hao Lv 0007, Yang Zhang 0125, Hao Lin 0001, Lin Ning 0002 |
Briefings Bioinform. | 5 |
| 2021 | iDHS-Deep: an integrated tool for predicting DNase I hypersensitive sites by deep neural networkabstractDNase I hypersensitive site (DHS) refers to the hypersensitive region of chromatin for the DNase I enzyme. It is an important part of the noncoding region and contains a variety of regulatory elements, such as promoter, enhancer, and transcription factor-binding site, etc. Moreover, the related locus of disease (or trait) are usually enriched in the DHS regions. Therefore, the detection of DHS region is of great significance. In this study, we develop a deep learning-based algorithm to identify whether an unknown sequence region would be potential DHS. The proposed method showed high prediction performance on both training datasets and independent datasets in different cell types and developmental stages, demonstrating that the method has excellent superiority in the identification of DHSs. Furthermore, for the convenience of related wet-experimental researchers, the user-friendly web-server iDHS-Deep was established at http://lin-group.cn/server/iDHS-Deep/, by which users can easily distinguish DHS and non-DHS and obtain the corresponding developmental stage ofDHS. Fu-Ying Dao, Hao Lv 0007, Hao Lin 0001 |
Briefings Bioinform. | 2 |
| 2021 | A computational platform to identify origins of replication sites in eukaryotesabstractThe locations of the initiation of genomic DNA replication are defined as origins of replication sites (ORIs), which regulate the onset of DNA replication and play significant roles in the DNA replication process. The study of ORIs is essential for understanding the cell-division cycle and gene expression regulation. Accurate identification of ORIs will provide important clues for DNA replication research and drug development by developing computational methods. In this paper, the first integrated predictor named iORI-Euk was built to identify ORIs in multiple eukaryotes and multiple cell types. In the predictor, seven eukaryotic (Homo sapiens, Mus musculus, Drosophila melanogaster, Arabidopsis thaliana, Pichia pastoris, Schizosaccharomyces pombe and Kluyveromyces lactis) ORI data was collected from public database to construct benchmark datasets. Subsequently, three feature extraction strategies which are k-mer, binary encoding and combination of k-mer and binary were used to formulate DNA sequence samples. We also compared the different classification algorithms' performance. As a result, the best results were obtained by using support vector machine in 5-fold cross-validation test and independent dataset test. Based on the optimal model, an online web server called iORI-Euk (http://lin-group.cn/server/iORI-Euk/) was established for the novel ORI identification. Fu-Ying Dao, Hao Lv 0007, Hasan Zulfiqar, Hui Ding 0005, Hao Lin 0001 |
Briefings Bioinform. | 2 |
| 2021 | DeepYY1: a deep learning approach to identify YY1-mediated chromatin loopsabstractThe protein Yin Yang 1 (YY1) could form dimers that facilitate the interaction between active enhancers and promoter-proximal elements. YY1-mediated enhancer-promoter interaction is the general feature of mammalian gene control. Recently, some computational methods have been developed to characterize the interactions between DNA elements by elucidating important features of chromatin folding; however, no computational methods have been developed for identifying the YY1-mediated chromatin loops. In this study, we developed a deep learning algorithm named DeepYY1 based on word2vec to determine whether a pair of YY1 motifs would form a loop. The proposed models showed a high prediction performance (AUCs$\ge$0.93) on both training datasets and testing datasets in different cell types, demonstrating that DeepYY1 has an excellent performance in the identification of the YY1-mediated chromatin loops. Our study also suggested that sequences play an important role in the formation of YY1-mediated chromatin loops. Furthermore, we briefly discussed the distribution of the replication origin site in the loops. Finally, a user-friendly web server was established, and it can be freely accessed at http://lin-group.cn/server/DeepYY1. Fu-Ying Dao, Hao Lv 0007, Dan Zhang 0006, Zi-Mei Zhang, Hao Lin 0001 |
Briefings Bioinform. | 2 |
| 2021 | Deep-Kcr: accurate detection of lysine crotonylation sites using deep learning methodabstractAs a newly discovered protein posttranslational modification, histone lysine crotonylation (Kcr) involved in cellular regulation and human diseases. Various proteomics technologies have been developed to detect Kcr sites. However, experimental approaches for identifying Kcr sites are often time-consuming and labor-intensive, which is difficult to widely popularize in large-scale species. Computational approaches are cost-effective and can be used in a high-throughput manner to generate relatively precise identification. In this study, we develop a deep learning-based method termed as Deep-Kcr for Kcr sites prediction by combining sequence-based features, physicochemical property-based features and numerical space-derived information with information gain feature selection. We investigate the performances of convolutional neural network (CNN) and five commonly used classifiers (long short-term memory network, random forest, LogitBoost, naive Bayes and logistic regression) using 10-fold cross-validation and independent set test. Results show that CNN could always display the best performance with high computational efficiency on large dataset. We also compare the Deep-Kcr with other existing tools to demonstrate the excellent predictive power and robustness of our method. Based on the proposed model, a webserver called Deep-Kcr was established and is freely accessible at http://lin-group.cn/server/Deep-Kcr. Hao Lv 0007, Fu-Ying Dao, Zheng-Xing Guan, Yan-Wen Li, Hao Lin 0001 |
Briefings Bioinform. | 1 |
| 2021 | Application of artificial intelligence and machine learning for COVID-19 drug discovery and vaccine designabstractThe global pandemic of coronavirus disease 2019 (COVID-19), caused by severe acute respiratory syndrome coronavirus 2, has led to a dramatic loss of human life worldwide. Despite many efforts, the development of effective drugs and vaccines for this novel virus will take considerable time. Artificial intelligence (AI) and machine learning (ML) offer promising solutions that could accelerate the discovery and optimization of new antivirals. Motivated by this, in this paper, we present an extensive survey on the application of AI and ML for combating COVID-19 based on the rapidly emerging literature. Particularly, we point out the challenges and future directions associated with state-of-the-art solutions to effectively control the COVID-19 pandemic. We hope that this review provides researchers with new insights into the ways AI and ML fight and have fought the COVID-19 outbreak. Hao Lv 0007, Joshua William Berkenpas, Fu-Ying Dao, Hasan Zulfiqar, Hui Ding 0005, Yang Zhang 0125, Renzhi Cao |
Briefings Bioinform. | 1 |
| 2021 | iCarPS: a computational tool for identifying protein carbonylation sites by novel encoded featuresabstractMOTIVATION: Protein carbonylation is one of the most important oxidative stress-induced post-translational modifications, which is generally characterized as stability, irreversibility and relative early formation. It plays a significant role in orchestrating various biological processes and has been already demonstrated to be related to many diseases. However, the experimental technologies for carbonylation sites identification are not only costly and time consuming, but also unable of processing a large number of proteins at a time. Thus, rapidly and effectively identifying carbonylation sites by computational methods will provide key clues for the analysis of occurrence and development of diseases. RESULTS: In this study, we developed a predictor called iCarPS to identify carbonylation sites based on sequence information. A novel feature encoding scheme called residues conical coordinates combined with their physicochemical properties was proposed to formulate carbonylated protein and non-carbonylated protein samples. To remove potential redundant features and improve the prediction performance, a feature selection technique was used. The accuracy and robustness of iCarPS were proved by experiments on training and independent datasets. Comparison with other published methods demonstrated that the proposed method is powerful and could provide powerful performance for carbonylation sites identification. AVAILABILITY AND IMPLEMENTATION: Based on the proposed model, a user-friendly webserver and a software package were constructed, which can be freely accessed at http://lin-group.cn/server/iCarPS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dan Zhang 0006, Hao Lv 0007, Hao Lin 0001 |
Bioinform. | 5 |
| 2020 | Evaluation of different computational methods on 5-methylcytosine sites identificationabstract5-Methylcytosine (m5C) plays an extremely important role in the basic biochemical process. With the great increase of identified m5C sites in a wide variety of organisms, their epigenetic roles become largely unknown. Hence, accurate identification of m5C site is a key step in understanding its biological functions. Over the past several years, more attentions have been paid on the identification of m5C sites in multiple species. In this work, we firstly summarized the current progresses in computational prediction of m5C sites and then constructed a more powerful and reliable model for identifying m5C sites. To train the model, we collected experimentally confirmed m5C data from Homo sapiens, Mus musculus, Saccharomyces cerevisiae and Arabidopsis thaliana, and compared the performances of different feature extraction methods and classification algorithms for optimizing prediction model. Based on the optimal model, a novel predictor called iRNA-m5C was developed for the recognition of m5C sites. Finally, we critically evaluated the performance of iRNA-m5C and compared it with existing methods. The result showed that iRNA-m5C could produce the best prediction performance. We hope that this paper could provide a guide on the computational identification of m5C site and also anticipate that the proposed iRNA-m5C will become a powerful tool for large scale identification of m5C sites. Hao Lv 0007, Zi-Mei Zhang, Jiu-Xin Tan, Wei Chen 0064, Hao Lin 0001 |
Briefings Bioinform. | 1 |
| 2020 | A comparison and assessment of computational method for identifying recombination hotspots in Saccharomyces cerevisiaeabstractMeiotic recombination is one of the most important driving forces of biological evolution, which is initiated by double-strand DNA breaks. Recombination has important roles in genome diversity and evolution. This review firstly provides a comprehensive survey of the 15 computational methods developed for identifying recombination hotspots in Saccharomyces cerevisiae. These computational methods were discussed and compared in terms of underlying algorithms, extracted features, predictive capability and practical utility. Subsequently, a more objective benchmark data set was constructed to develop a new predictor iRSpot-Pse6NC2.0 (http://lin-group.cn/server/iRSpot-Pse6NC2.0). To further demonstrate the generalization ability of these methods, we compared iRSpot-Pse6NC2.0 with existing methods on the chromosome XVI of S. cerevisiae. The results of the independent data set test demonstrated that the new predictor is superior to existing tools in the identification of recombination hotspots. The iRSpot-Pse6NC2.0 will become an important tool for identifying recombination hotspot. Wuritu Yang, Fu-Ying Dao, Hao Lv 0007, Hui Ding 0005, Wei Chen 0064, Hao Lin 0001 |
Briefings Bioinform. | 4 |
| 2019 | i6mA-Pred: identifying DNA N6-methyladenine sites in the rice genomeabstractMOTIVATION: DNA N6-methyladenine (6mA) is associated with a wide range of biological processes. Since the distribution of 6mA site in the genome is non-random, accurate identification of 6mA sites is crucial for understanding its biological functions. Although experimental methods have been proposed for this regard, they are still cost-ineffective for detecting 6mA site in genome-wide scope. Therefore, it is desirable to develop computational methods to facilitate the identification of 6mA site. RESULTS: In this study, a computational method called i6mA-Pred was developed to identify 6mA sites in the rice genome, in which the optimal nucleotide chemical properties obtained by the using feature selection technique were used to encode the DNA sequences. It was observed that the i6mA-Pred yielded an accuracy of 83.13% in the jackknife test. Meanwhile, the performance of i6mA-Pred was also superior to other methods. AVAILABILITY AND IMPLEMENTATION: A user-friendly web-server, i6mA-Pred is freely accessible at http://lin-group.cn/server/i6mA-Pred. Wei Chen 0064, Hao Lv 0007, Fulei Nie, Hao Lin 0001 |
Bioinform. | 2 |
| 2019 | Identify origin of replication in Saccharomyces cerevisiae using two-step feature selection techniqueabstractMOTIVATION: DNA replication is a key step to maintain the continuity of genetic information between parental generation and offspring. The initiation site of DNA replication, also called origin of replication (ORI), plays an extremely important role in the basic biochemical process. Thus, rapidly and effectively identifying the location of ORI in genome will provide key clues for genome analysis. Although biochemical experiments could provide detailed information for ORI, it requires high experimental cost and long experimental period. As good complements to experimental techniques, computational methods could overcome these disadvantages. RESULTS: Thus, in this study, we developed a predictor called iORI-PseKNC2.0 to identify ORIs in the Saccharomyces cerevisiae genome based on sequence information. The PseKNC including 90 physicochemical properties was proposed to formulate ORI and non-ORI samples. In order to improve the accuracy, a two-step feature selection was proposed to exclude redundant and noise information. As a result, the overall success rate of 88.53% was achieved in the 5-fold cross-validation test by using support vector machine. AVAILABILITY AND IMPLEMENTATION: Based on the proposed model, a user-friendly webserver was established and can be freely accessed at http://lin-group.cn/server/iORI-PseKNC2.0. The webserver will provide more convenience to most of wet-experimental scholars. Fu-Ying Dao, Hao Lv 0007, Chao-Qin Feng, Hui Ding 0005, Wei Chen 0064, Hao Lin 0001 |
Bioinform. | 2 |