VLDB 2026 Research / reviewers in the wild / expert
Hao Sun 0001
dblp:82/2248-1
· DBLP profile ↗
13ranked-venue papers
1as first author
4since 2021 · last 2024
0000-0002-5547-9501ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | NanoDeep: a deep learning framework for nanopore adaptive sampling on microbial sequencingabstractNanopore sequencers can enrich or deplete the targeted DNA molecules in a library by reversing the voltage across individual nanopores. However, it requires substantial computational resources to achieve rapid operations in parallel at read-time sequencing. We present a deep learning framework, NanoDeep, to overcome these limitations by incorporating convolutional neural network and squeeze and excitation. We first showed that the raw squiggle derived from native DNA sequences determines the origin of microbial and human genomes. Then, we demonstrated that NanoDeep successfully classified bacterial reads from the pooled library with human sequence and showed enrichment for bacterial sequence compared with routine nanopore sequencing setting. Further, we showed that NanoDeep improves the sequencing efficiency and preserves the fidelity of bacterial genomes in the mock sample. In addition, NanoDeep performs well in the enrichment of metagenome sequences of gut samples, showing its potential applications in the enrichment of unknown microbiota. Our toolkit is available at https://github.com/lysovosyl/NanoDeep. Yusen Lin, Xing Zhao 0008, Xiaojuan Teng, Jingxia Lin, Bowen Shu, Hao Sun 0001, Yuhui Liao, Jiajian Zhou |
Briefings Bioinform. | 9 |
| 2024 | Occlusion enhanced pan-cancer classification via deep learningabstractQuantitative measurement of RNA expression levels through RNA-Seq is an ideal replacement for conventional cancer diagnosis via microscope examination. Currently, cancer-related RNA-Seq studies focus on two aspects: classifying the status and tissue of origin of a sample and discovering marker genes. Existing studies typically identify marker genes by statistically comparing healthy and cancer samples. However, this approach overlooks marker genes with low expression level differences and may be influenced by experimental results. This paper introduces "GENESO," a novel framework for pan-cancer classification and marker gene discovery using the occlusion method in conjunction with deep learning. we first trained a baseline deep LSTM neural network capable of distinguishing the origins and statuses of samples utilizing RNA-Seq data. Then, we propose a novel marker gene discovery method called "Symmetrical Occlusion (SO)". It collaborates with the baseline LSTM network, mimicking the "gain of function" and "loss of function" of genes to evaluate their importance in pan-cancer classification quantitatively. By identifying the genes of utmost importance, we then isolate them to train new neural networks, resulting in higher-performance LSTM models that utilize only a reduced set of highly relevant genes. The baseline neural network achieves an impressive validation accuracy of 96.59% in pan-cancer classification. With the help of SO, the accuracy of the second network reaches 98.30%, while using 67% fewer genes. Notably, our method excels in identifying marker genes that are not differentially expressed. Moreover, we assessed the feasibility of our method using single-cell RNA-Seq data, employing known marker genes as a validation test. Xing Zhao 0008, Zigui Chen, Huating Wang, Hao Sun 0001 |
BMC Bioinform. | 4 |
| 2021 | Large scale RNA-binding proteins/LncRNAs interaction analysis to uncover lncRNA nuclear localization mechanismsabstractLong non-coding RNAs (lncRNAs) are key regulators of major biological processes and their functional modes are dictated by their subcellular localization. Relative nuclear enrichment of lncRNAs compared to mRNAs is a prevalent phenomenon but the molecular mechanisms governing their nuclear retention in cells remain largely unknown. Here in this study, we harness the recently released eCLIP data for a large number of RNA-binding proteins (RBPs) in K562 and HepG2 cells and utilize multiple bioinformatics methods to comprehensively survey the roles of RBPs in lncRNA nuclear retention. We identify an array of splicing RBPs that bind to nuclear-enriched lincRNAs (large intergenic non-coding RNAs) thus may act as trans-factors regulating their nuclear retention. Further analyses reveal that these RBPs may bind with distinct core motifs, flanking sequence compositions, or secondary structures to drive lincRNA nuclear retention. Moreover, network analyses uncover potential co-regulatory RBP clusters and the physical interaction between HNRNPU and SAFB2 proteins in K562 cells is further experimentally verified. Altogether, our analyses reveal previously unknown factors and mechanisms that govern lincRNA nuclear localization in cells. Yile Huang, Yulong Qiao, Jiajian Zhou, Hao Sun 0001, Huating Wang |
Briefings Bioinform. | 7 |
| 2021 | NAMS webserver: coding potential assessment and functional annotation of plant transcriptsabstractRecent advances in transcriptomics have uncovered lots of novel transcripts in plants. To annotate such transcripts, dissecting their coding potential is a critical step. Computational approaches have been proven fruitful in this task; however, most current tools are designed/optimized for mammals and only a few of them have been tested on a limited number of plant species. In this work, we present NAMS webserver, which contains a novel coding potential classifier, NAMS, specifically optimized for plants. We have evaluated the performance of NAMS using a comprehensive dataset containing more than 3 million transcripts from various plant species, where NAMS demonstrates high accuracy and remarkable performance improvements over state-of-the-art software. Moreover, our webserver also furnishes functional annotations, aiming to provide users informative clues to the functions of their transcripts. Considering that most plant species are poorly characterized, our NAMS webserver could serve as a valuable resource to facilitate the transcriptomic studies. The webserver with testing dataset is freely available at http://sunlab.cpy.cuhk.edu.hk/NAMS/. Kun Sun 0003, Huating Wang, Hao Sun 0001 |
Briefings Bioinform. | 3 |
| 2020 | RNA binding proteins (RBPs) regulate lncRNA nuclear retentionabstractIt is well known that functional modes of lncRNAs are intimately associated with their subcellular localization1and relative nuclear enrichment of lncRNAs compared to mRNAs is a prevalent phenomenon2. As RBPs control the production, maturation, localization, translation, and degradation of cellular RNAs3, we reason that it is important to uncover partner RBPs that can bind and facilitate lncRNA nuclear localization. In this study, we thus harnessed the recently released large scale of eCLIP data4and subcellular RNA-seq data5available in K562 and HepG2 cell lines to characterize lncRNA-RBP interactome and uncovered potential factors and associated mechanisms determining lncRNA nuclear retention. Analyses of the subcellular RNA-seq data identified nuclear enriched lncRNAs (nuc-lncRNAs) and confirmed that lncRNAs and eRNAs (enhancer associated lncRNAs) are relatively nuclear enriched in both cells. By integrating the RBP binding profiles, we next generated RBP/nuc-lncRNA interaction map to identify RBPs associated with nuc-lncRNAs including HNRNPU, SAFB2, KHSRP and KHDRBS1 in K562 and HRNPNPC as well as HNRNPL in HepG2 cell lines. To further confirm the above findings, HNRNPU was knocked down and which led to nuclear retention of a panel of lncRNAs. Yile Huang, Yulong Qiao, Yingzhe Ding, Jiajian Zhou, Huating Wang, Hao Sun 0001 |
BIBM | 8 |
| 2019 | SKmDB: an integrated database of next generation sequencing information in skeletal muscleabstractMOTIVATION: Skeletal muscles have indispensable functions and also possess prominent regenerative ability. The rapid emergence of Next Generation Sequencing (NGS) data in recent years offers us an unprecedented perspective to understand gene regulatory networks governing skeletal muscle development and regeneration. However, the data from public NGS database are often in raw data format or processed with different procedures, causing obstacles to make full use of them. RESULTS: We provide SKmDB, an integrated database of NGS information in skeletal muscle. SKmDB not only includes all NGS datasets available in the human and mouse skeletal muscle tissues and cells, but also provide preliminary data analyses including gene/isoform expression levels, gene co-expression subnetworks, as well as assembly of putative lincRNAs, typical and super enhancers and transcription factor hotspots. Users can efficiently search, browse and visualize the information with the well-designed user interface and server side. SKmDB thus will offer wet lab biologists useful information to study gene regulatory mechanisms in the field of skeletal muscle development and regeneration. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at http://sunlab.cpy.cuhk.edu.hk/SKmDB. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiajian Zhou, Huating Wang, Hao Sun 0001 |
Bioinform. | 4 |
| 2018 | GeneCT: a generalizable cancerous status and tissue origin classifier for pan-cancer biopsiesabstractMotivation: Tissue biopsy is commonly used in cancer diagnosis and molecular studies. However, advanced skills are required for determining cancerous status of biopsies and tissue origin of tumor for cancerous ones. Correct classification is essential for downstream experiment design and result interpretation, especially in molecular cancer studies. Methods for accurate classification of cancerous status and tissue origin for pan-cancer biopsies are thus urgently needed. Results: We developed a deep learning-based classifier, named GeneCT, for predicting cancerous status and tissue origin of pan-cancer biopsies. GeneCT showed high performance on pan-cancer datasets from various sources and outperformed existing tools. We believe that GeneCT can potentially facilitate cancer diagnosis, tumor origin determination and molecular cancer studies. Availability and implementation: GeneCT is implemented in Perl/R and supported on GNU/Linux platforms. Source code, testing data and webserver are freely available at http://sunlab.cpy.cuhk.edu.hk/GeneCT/. Supplementary information: Supplementary data are available at Bioinformatics online. Kun Sun 0003, Huating Wang, Hao Sun 0001 |
Bioinform. | 4 |
| 2018 | lncFunTK: a toolkit for functional annotation of long noncoding RNAsabstractMotivation: Thousands of long noncoding RNAs (lncRNAs) were newly identified from high throughput RNA-seq data. Functional annotation and prioritization of these lncRNAs for further experimental validation as well as the functional investigation is the bottleneck step for many noncoding RNA studies. Results: Here we describe lncFunTK that can run either as standard application or webserver for this purpose. It integrates high throughput sequencing data (i.e. ChIP-seq, CLIP-seq and RNA-seq) to construct the regulatory network associated with lncRNAs. Through the network, it calculates the Functional Information Score (FIS) of each individual lncRNA for prioritizing and inferring its functions through Gene Ontology (GO) terms of neighboring genes. In addition, it also provides utility scripts to support the input data preprocessing and the parameter optimizing. We further demonstrate that lncFunTK can be widely used in various biological systems for lncRNA prioritization and functional annotation. Availability and implementation: The lncFunTK standalone version is an open source package and freely available at http://sunlab.cpy.cuhk.edu.hk/lncfuntk under the MIT license. A webserver implementation is also available at http://sunlab.cpy.cuhk.edu.hk/lncfuntk/runlncfuntk.html. Supplementary information: Supplementary data are available at Bioinformatics online. Jiajian Zhou, Yile Huang, Yingzhe Ding, Huating Wang, Hao Sun 0001 |
Bioinform. | 6 |
| 2017 | BSviewer: a genotype-preserving, nucleotide-level visualizer for bisulfite sequencing dataabstractMOTIVATION: The bisulfite sequencing technology has been widely used to study the DNA methylation profile in many species. However, most of the current visualization tools for bisulfite sequencing data only provide high-level views (i.e. overall methylation densities) while miss the methylation dynamics at nucleotide level. Meanwhile, they also focus on CpG sites while omit other information (such as genotypes on SNP sites) which could be helpful for interpreting the methylation pattern of the data. A bioinformatics tool that visualizes the methylation statuses at nucleotide level and preserves the most essential information of the sequencing data is thus valuable and needed. RESULTS: We have developed BSviewer, a lightweight nucleotide-level visualization tool for bisulfite sequencing data. Using an imprinting gene as an example, we show that BSviewer could be specifically helpful for interpreting the data with allele-specific DNA methylation pattern. AVAILABILITY AND IMPLEMENTATION: BSviewer is implemented in Perl and runs on most GNU/Linux platforms. Source code and testing dataset are freely available at http://sunlab.cpy.cuhk.edu.hk/BSviewer/. CONTACT: [email protected]. Kun Sun 0003, Fiona F. M. Lun, Peiyong Jiang, Hao Sun 0001 |
Bioinform. | 4 |
| 2012 | FetalQuant: deducing fractional fetal DNA concentration from massively parallel sequencing of DNA in maternal plasmaabstractMOTIVATION: The fractional fetal DNA concentration is one of the critical parameters for non-invasive prenatal diagnosis based on the analysis of DNA in maternal plasma. Massively parallel sequencing (MPS) of DNA in maternal plasma has been demonstrated to be a powerful tool for the non-invasive prenatal diagnosis of fetal chromosomal aneuploidies. With the rapid advance of MPS technologies, the sequencing cost per base is dramatically reducing, especially when using targeted MPS. Even though several approaches have been developed for deducing the fractional fetal DNA concentration, none of them can be used to deduce the fractional fetal DNA concentration directly from the sequencing data without prior genotype information. RESULT: In this study, we implement a statistical mixture model, named FetalQuant, which utilizes the maximum likelihood to estimate the fractional fetal DNA concentration directly from targeted MPS of DNA in maternal plasma. This method allows the improved deduction of the fractional fetal DNA concentration, obviating the need of genotype information without loss of accuracy. Furthermore, by using Bayes' rule, this method can distinguish the informative single-nucleotide polymorphism loci where the mother is homozygous and the fetus is heterozygous. We believe that FetalQuant can help expand the spectrum of diagnostic applications using MPS on DNA in maternal plasma. AVAILABILITY: Software and simulation data are available at http://sourceforge.net/projects/fetalquant/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Peiyong Jiang, K. C. Allen Chan, Gary J. W. Liao, Yama W. L. Zheng, Tak Y. Leung, Rossa W. K. Chiu, Yuk Ming Dennis Lo, Hao Sun 0001 |
Bioinform. | 8 |
| 2005 | OMGProm: a database of orthologous mammalian gene promotersabstractAbstract Summary: Sequence comparisons between human and rodents are increasingly being used for the identification of gene regulatory regions. The effectiveness of such an approach largely depends on the quality and availability of promoter sequences. We developed OMGProm by integrating three data sources: (1) experimentally supported full-length cDNA, promoter and first exon sequences; (2) homology information from HomoloGene and (3) the human and mouse genomic sequences. The current version of OMGProm contains 8550 promoter pairs of 6373 orthologous human and mouse genes, where supporting experimental evidence for transcription start site annotation exists in at least one species. Availability: OMGProm can be accessed from http://bioinformatics.med.ohio-state.edu/OMGProm Contact: [email protected] Supplementary information: Additional information on methods and implementation is available at http://bioinformatics.med.ohio-state.edu/OMGProm/si.jsp. Saranyan K. Palaniswamy, Victor X. Jin, Hao Sun 0001, Ramana V. Davuluri |
Bioinform. | 3 |
| 2004 | Java-based application framework for visualization of gene regulatory region annotationsabstractMOTIVATION: The genome sequences of several organisms are either complete, or being sequenced. Each genome needs to be integrated with various types of annotations, e.g. locations of genes, promoters and other functional elements such as transcriptional regulatory elements. A robust application framework will be useful for developing web-based applications to visualize various genome annotations. RESULTS: We developed genome data visualization toolkit (GDVTK) as an application framework that consists of a set of data structures and core classes, using Java technology. GDVTK is a sound framework for developing web-based applications to present the gene regulatory region annotations in visual form. The current version of GDVTK consists of eight packages and 38 Java classes that are portable, reusable and extensible for plugging in new data sources and models. We implemented GDVTK for visualization of promoter annotations in Mammalian Promoter Database (MPromDb), a web-based gene-regulatory information server. AVAILABILITY: GDVTK is available under GNU general public license. Source code and software documentation can be found at the URL http://bioinformatics.med.ohio-state.edu/GDVTK. Hao Sun 0001, Ramana V. Davuluri |
Bioinform. | 1 |
| 2003 | AGRIS: Arabidopsis Gene Regulatory Information Server, an information resource of Arabidopsis cis-regulatory elements and transcription factorsabstractBACKGROUND: The gene regulatory information is hardwired in the promoter regions formed by cis-regulatory elements that bind specific transcription factors (TFs). Hence, establishing the architecture of plant promoters is fundamental to understanding gene expression. The determination of the regulatory circuits controlled by each TF and the identification of the cis-regulatory sequences for all genes have been identified as two of the goals of the Multinational Coordinated Arabidopsis thaliana Functional Genomics Project by the Multinational Arabidopsis Steering Committee (June 2002). RESULTS: AGRIS is an information resource of Arabidopsis promoter sequences, transcription factors and their target genes. AGRIS currently contains two databases, AtTFDB (Arabidopsis thaliana transcription factor database) and AtcisDB (Arabidopsis thaliana cis-regulatory database). AtTFDB contains information on approximately 1,400 transcription factors identified through motif searches and grouped into 34 families. AtTFDB links the sequence of the transcription factors with available mutants and, when known, with the possible genes they may regulate. AtcisDB consists of the 5' regulatory sequences of all 29,388 annotated genes with a description of the corresponding cis-regulatory elements. Users can search the databases for (i) promoter sequences, (ii) a transcription factor, (iii) a direct target genes for a specific transcription factor, or (vi) a regulatory network that consists of transcription factors and their target genes. CONCLUSION: AGRIS provides the necessary software tools on Arabidopsis transcription factors and their putative binding sites on all genes to initiate the identification of transcriptional regulatory networks in the model dicotyledoneous plant Arabidopsis thaliana. AGRIS can be accessed from http://arabidopsis.med.ohio-state.edu. Ramana V. Davuluri, Hao Sun 0001, Saranyan K. Palaniswamy, Nicole Matthews, Mike Kurtz, Erich Grotewold |
BMC Bioinform. | 2 |