EDBT 2026 Demo / reviewers in the wild / expert
Haiyan Hu 0004
dblp:57/476-4
· DBLP profile ↗
22ranked-venue papers
1as first author
10since 2021 · last 2024
0000-0002-4580-5975ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 1 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A survey of experimental and computational identification of small proteinsabstractSmall proteins (SPs) are typically characterized as eukaryotic proteins shorter than 100 amino acids and prokaryotic proteins shorter than 50 amino acids. Historically, they were disregarded because of the arbitrary size thresholds to define proteins. However, recent research has revealed the existence of many SPs and their crucial roles. Despite this, the identification of SPs and the elucidation of their functions are still in their infancy. To pave the way for future SP studies, we briefly introduce the limitations and advancements in experimental techniques for SP identification. We then provide an overview of available computational tools for SP identification, their constraints, and their evaluation. Additionally, we highlight existing resources for SP research. This survey aims to initiate further exploration into SPs and encourage the development of more sophisticated computational tools for SP identification in prokaryotes and microbiomes. Joshua Beals, Haiyan Hu 0004, Xiaoman Shawn Li |
Briefings Bioinform. | 2 |
| 2024 | A deep learning method to integrate extracelluar miRNA with mRNA for cancer studiesabstractMOTIVATION: Extracellular miRNAs (exmiRs) and intracellular mRNAs both can serve as promising biomarkers and therapeutic targets for various diseases. However, exmiR expression data is often noisy, and obtaining intracellular mRNA expression data usually involves intrusive procedures. To gain valuable insights into disease mechanisms, it is thus essential to improve the quality of exmiR expression data and develop noninvasive methods for assessing intracellular mRNA expression. RESULTS: We developed CrossPred, a deep-learning multi-encoder model for the cross-prediction of exmiRs and mRNAs. Utilizing contrastive learning, we created a shared embedding space to integrate exmiRs and mRNAs. This shared embedding was then used to predict intracellular mRNA expression from noisy exmiR data and to predict exmiR expression from intracellular mRNA data. We evaluated CrossPred on three types of cancers and assessed its effectiveness in predicting the expression levels of exmiRs and mRNAs. CrossPred outperformed the baseline encoder-decoder model, exmiR or mRNA-based models, and variational autoencoder models. Moreover, the integration of exmiR and mRNA data uncovered important exmiRs and mRNAs associated with cancer. Our study offers new insights into the bidirectional relationship between mRNAs and exmiRs. AVAILABILITY AND IMPLEMENTATION: The datasets and tool are available at https://doi.org/10.5281/zenodo.13891508. Tasbiraha Athaya, Xiaoman Shawn Li, Haiyan Hu 0004 |
Bioinform. | 3 |
| 2023 | Multimodal deep learning approaches for single-cell multi-omics data integrationabstractIntegrating single-cell multi-omics data is a challenging task that has led to new insights into complex cellular systems. Various computational methods have been proposed to effectively integrate these rapidly accumulating datasets, including deep learning. However, despite the proven success of deep learning in integrating multi-omics data and its better performance over classical computational methods, there has been no systematic study of its application to single-cell multi-omics data integration. To fill this gap, we conducted a literature review to explore the use of multimodal deep learning techniques in single-cell multi-omics data integration, taking into account recent studies from multiple perspectives. Specifically, we first summarized different modalities found in single-cell multi-omics data. We then reviewed current deep learning techniques for processing multimodal data and categorized deep learning-based integration methods for single-cell multi-omics data according to data modality, deep learning architecture, fusion strategy, key tasks and downstream analysis. Finally, we provided insights into using these deep learning models to integrate multi-omics data and better understand single-cell biological mechanisms. Tasbiraha Athaya, Rony Chowdhury Ripan, Xiaoman Shawn Li, Haiyan Hu 0004 |
Briefings Bioinform. | 4 |
| 2022 | Computational analyses of bacterial strains from shotgun readsabstractShotgun sequencing is routinely employed to study bacteria in microbial communities. With the vast amount of shotgun sequencing reads generated in a metagenomic project, it is crucial to determine the microbial composition at the strain level. This study investigated 20 computational tools that attempt to infer bacterial strain genomes from shotgun reads. For the first time, we discussed the methodology behind these tools. We also systematically evaluated six novel-strain-targeting tools on the same datasets and found that BHap, mixtureS and StrainFinder performed better than other tools. Because the performance of the best tools is still suboptimal, we discussed future directions that may address the limitations. Minerva Fatimae Ventolero, Saidi Wang, Haiyan Hu 0004, Xiaoman Shawn Li |
Briefings Bioinform. | 3 |
| 2022 | Letter to the editor: evaluating computational tools for lncRNA identification on independent datasetsabstractThe authors of the BASiNET tool claim that the survey paper 'A systematic evaluation of computational tools for lncRNA identification' incorrectly evaluates the BASiNET tool. Here, we point out that the survey paper correctly evaluates the BASiNET tool and why the evaluation should not be carried out as BASiNET authors suggest. Hansi Zheng, Xiaoman Shawn Li, Haiyan Hu 0004 |
Briefings Bioinform. | 3 |
| 2021 | Interpretation of deep learning in genomics and epigenomicsabstractMachine learning methods have been widely applied to big data analysis in genomics and epigenomics research. Although accuracy and efficiency are common goals in many modeling tasks, model interpretability is especially important to these studies towards understanding the underlying molecular and cellular mechanisms. Deep neural networks (DNNs) have recently gained popularity in various types of genomic and epigenomic studies due to their capabilities in utilizing large-scale high-throughput bioinformatics data and achieving high accuracy in predictions and classifications. However, DNNs are often challenged by their potential to explain the predictions due to their black-box nature. In this review, we present current development in the model interpretation of DNNs, focusing on their applications in genomics and epigenomics. We first describe state-of-the-art DNN interpretation methods in representative machine learning fields. We then summarize the DNN interpretation methods in recent studies on genomics and epigenomics, focusing on current data- and computing-intensive topics such as sequence motif identification, genetic variations, gene expression, chromatin interactions and non-coding RNAs. We also present the biological discoveries that resulted from these interpretation methods. We finally discuss the advantages and limitations of current interpretation approaches in the context of genomic and epigenomic studies. Contact:[email protected], [email protected]. Amlan Talukder, Clayton Barham, Xiaoman Shawn Li, Haiyan Hu 0004 |
Briefings Bioinform. | 4 |
| 2021 | Computational annotation of miRNA transcription start sitesabstractMOTIVATION: MicroRNAs (miRNAs) are small noncoding RNAs that play important roles in gene regulation and phenotype development. The identification of miRNA transcription start sites (TSSs) is critical to understand the functional roles of miRNA genes and their transcriptional regulation. Unlike protein-coding genes, miRNA TSSs are not directly detectable from conventional RNA-Seq experiments due to miRNA-specific process of biogenesis. In the past decade, large-scale genome-wide TSS-Seq and transcription activation marker profiling data have become available, based on which, many computational methods have been developed. These methods have greatly advanced genome-wide miRNA TSS annotation. RESULTS: In this study, we summarized recent computational methods and their results on miRNA TSS annotation. We collected and performed a comparative analysis of miRNA TSS annotations from 14 representative studies. We further compiled a robust set of miRNA TSSs (RSmirT) that are supported by multiple studies. Integrative genomic and epigenomic data analysis on RSmirT revealed the genomic and epigenomic features of miRNA TSSs as well as their relations to protein-coding and long non-coding genes. CONTACT: [email protected], [email protected]. Saidi Wang, Amlan Talukder, Mingyu Cha, Xiaoman Shawn Li, Haiyan Hu 0004 |
Briefings Bioinform. | 5 |
| 2021 | Corrigendum to: Computational annotation of miRNA transcription start sitesabstractThe first version of this text did not label Saidi Wang and Amlan Talukder as equally contributing authors or Xiaoman Li and Haiyan Hu as co-corresponding authors, and also included an incorrect link to the supplementary material. This has now been corrected. Saidi Wang, Amlan Talukder, Mingyu Cha, Xiaoman Shawn Li, Haiyan Hu 0004 |
Briefings Bioinform. | 5 |
| 2021 | A systematic evaluation of the computational tools for lncRNA identificationabstractThe computational identification of long non-coding RNAs (lncRNAs) is important to study lncRNAs and their functions. Despite the existence of many computation tools for lncRNA identification, to our knowledge, there is no systematic evaluation of these tools on common datasets and no consensus regarding their performance and the importance of the features used. To fill this gap, in this study, we assessed the performance of 17 tools on several common datasets. We also investigated the importance of the features used by the tools. We found that the deep learning-based tools have the best performance in terms of identifying lncRNAs, and the peptide features do not contribute much to the tool accuracy. Moreover, when the transcripts in a cell type were considered, the performance of all tools significantly dropped, and the deep learning-based tools were no longer as good as other tools. Our study will serve as an excellent starting point for selecting tools and features for lncRNA identification. Hansi Zheng, Amlan Talukder, Xiaoman Shawn Li, Haiyan Hu 0004 |
Briefings Bioinform. | 4 |
| 2021 | mixtureS: a novel tool for bacterial strain genome reconstruction from readsabstractMOTIVATION: It is essential to study bacterial strains in environmental samples. Existing methods and tools often depend on known strains or known variations, cannot work on individual samples, not reliable, or not easy to use, etc. It is thus important to develop more user-friendly tools that can identify bacterial strains more accurately. RESULTS: We developed a new tool called mixtureS that can de novo identify bacterial strains from shotgun reads of a clonal or metagenomic sample, without prior knowledge about the strains and their variations. Tested on 243 simulated datasets and 195 experimental datasets, mixtureS reliably identified the strains, their numbers and their abundance. Compared with three tools, mixtureS showed better performance in almost all simulated datasets and the vast majority of experimental datasets. AVAILABILITY AND IMPLEMENTATION: The source code and tool mixtureS is available at http://www.cs.ucf.edu/˜xiaoman/mixtureS/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haiyan Hu 0004, Xiaoman Shawn Li |
Bioinform. | 2 |
| 2020 | Deep Learning to Identify Transcription Start Sites from CAGE DataabstractGene transcription start site (TSS) identification is important to understanding transcriptional gene regulation. Cap Analysis Gene Expression (CAGE) experiments have recently become common practice for direct measurement of TSSs. Currently, CAGE data available in public databases created unprecedented opportunities to study gene transcriptional initiation mechanisms under various cellular conditions. However, due to potential transcriptional noises inherent in CAGE data, in-silico methods are required to identify bonafide TSSs from noises further. Here we present a computational approach dlCAGE, an end-to-end deep neural network to identify TSSs from CAGE data. dlCAGE incorporate de-novo DNA regulatory motif features discovered by DeepBind model architecture, as well as existing sequence and structural features. Testing results of dlCAGE in several cell lines in comparison with current state-of-the-art approaches showed its superior performance and promise in TSS identification from CAGE experiments. Hansi Zheng, Xiaoman Shawn Li, Haiyan Hu 0004 |
BIBM | 3 |
| 2020 | Position-wise binding preference is important for miRNA target site predictionabstractMOTIVATION: It is a fundamental task to identify microRNAs (miRNAs) targets and accurately locate their target sites. Genome-scale experiments for miRNA target site detection are still costly. The prediction accuracies of existing computational algorithms and tools are often not up to the expectation due to a large number of false positives. One major obstacle to achieve a higher accuracy is the lack of knowledge of the target binding features of miRNAs. The published high-throughput experimental data provide an opportunity to analyze position-wise preference of miRNAs in terms of target binding, which can be an important feature in miRNA target prediction algorithms. RESULTS: We developed a Markov model to characterize position-wise pairing patterns of miRNA-target interactions. We further integrated this model as a scoring method and developed a dynamic programming (DP) algorithm, MDPS (Markov model-scored Dynamic Programming algorithm for miRNA target site Selection) that can screen putative target sites of miRNA-target binding. The MDPS algorithm thus can take into account both the dependency of neighboring pairing positions and the global pairing information. Based on the trained Markov models from both miRNA-specific and general datasets, we discovered that the position-wise binding information specific to a given miRNA would benefit its target prediction. We also found that miRNAs maintain region-wise similarity in their target binding patterns. Combining MDPS with existing methods significantly improves their precision while only slightly reduces their recall. Therefore, position-wise pairing patterns have the promise to improve target prediction if incorporated into existing software tools. AVAILABILITY AND IMPLEMENTATION: The source code and tool to calculate MDPS score is available at http://hulab.ucf.edu/research/projects/MDPS/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Amlan Talukder, Xiaoman Shawn Li, Haiyan Hu 0004 |
Bioinform. | 3 |
| 2019 | BHap: a novel approach for bacterial haplotype reconstructionabstractMOTIVATION: The bacterial haplotype reconstruction is critical for selecting proper treatments for diseases caused by unknown haplotypes. Existing methods and tools do not work well on this task, because they are usually developed for viral instead of bacterial populations. RESULTS: In this study, we developed BHap, a novel algorithm based on fuzzy flow networks, for reconstructing bacterial haplotypes from next generation sequencing data. Tested on simulated and experimental datasets, we showed that BHap was capable of reconstructing haplotypes of bacterial populations with an average F1 score of 0.87, an average precision of 0.87 and an average recall of 0.88. We also demonstrated that BHap had a low susceptibility to sequencing errors, was capable of reconstructing haplotypes with low coverage and could handle a wide range of mutation rates. Compared with existing approaches, BHap outperformed them in terms of higher F1 scores, better precision, better recall and more accurate estimation of the number of haplotypes. AVAILABILITY AND IMPLEMENTATION: The BHap tool is available at http://www.cs.ucf.edu/∼xiaoman/BHap/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Samaneh Saadat, Haiyan Hu 0004, Xiaoman Shawn Li |
Bioinform. | 3 |
| 2019 | EPIP: a novel approach for condition-specific enhancer-promoter interaction predictionabstractMOTIVATION: The identification of enhancer-promoter interactions (EPIs), especially condition-specific ones, is important for the study of gene transcriptional regulation. Existing experimental approaches for EPI identification are still expensive, and available computational methods either do not consider or have low performance in predicting condition-specific EPIs. RESULTS: We developed a novel computational method called EPIP to reliably predict EPIs, especially condition-specific ones. EPIP is capable of predicting interactions in samples with limited data as well as in samples with abundant data. Tested on more than eight cell lines, EPIP reliably identifies EPIs, with an average area under the receiver operating characteristic curve of 0.95 and an average area under the precision-recall curve of 0.73. Tested on condition-specific EPIPs, EPIP correctly identified 99.26% of them. Compared with two recently developed methods, EPIP outperforms them with a better accuracy. AVAILABILITY AND IMPLEMENTATION: The EPIP tool is freely available at http://www.cs.ucf.edu/˜xiaoman/EPIP/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Amlan Talukder, Samaneh Saadat, Xiaoman Shawn Li, Haiyan Hu 0004 |
Bioinform. | 4 |
| 2018 | CCmiR: a computational approach for competitive and cooperative microRNA binding predictionabstractMOTIVATION: The identification of microRNA (miRNA) target sites is important. In the past decade, dozens of computational methods have been developed to predict miRNA target sites. Despite their existence, rarely does a method consider the well-known competition and cooperation among miRNAs when attempts to discover target sites. To fill this gap, we developed a new approach called CCmiR, which takes the cooperation and competition of multiple miRNAs into account in a statistical model to predict their target sites. RESULTS: Tested on four different datasets, CCmiR predicted miRNA target sites with a high recall and a reasonable precision, and identified known and new cooperative and competitive miRNAs supported by literature. Compared with three state-of-the-art computational methods, CCmiR had a higher recall and a higher precision. AVAILABILITY AND IMPLEMENTATION: CCmiR is freely available at http://hulab.ucf.edu/research/projects/miRNA/CCmiR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaoman Shawn Li, Haiyan Hu 0004 |
Bioinform. | 3 |
| 2017 | Discover the semantic structure of human reference epigenome by differential latent dirichlet allocationabstractUnderstanding epigenetic changes across various conditions is a fundamental problem to epigenome annotation. With more high-throughput epigenomic data available, computational methods have been developed to quantify various types of epigenetic modification signals, to compare epigenetic marks between different conditions and to understand the functional consequences of epigenetic changes. However, currently few studies on epigenomes aim to provide a global view of epigenetic changes through large-scale high-throughput data integration. We apply a probabilistic graphical model called differential latent Dirichlet allocation (DLDA) to discover latent epigenetic modification modules (EMMs) from 56 reference epigenomes. We demonstrate the identified EMMs can characterize the global semantic structure of the reference epigenomes. The resulted EMMs show their condition-relevance to the corresponding reference epigenomes. The genes involved in these EMMs show epigenome-relevant functionality. Study of the involved epigenetic modification marks involved in these EMMs reveals the relative activity levels of epigenetic marks in different epigenomes. Clustering gene-epigenetic modification pairs leads to the discovery of more functional epigenetic modification groups. Yiyu Zheng, Xiaoman Shawn Li, Haiyan Hu 0004 |
BIBM | 3 |
| 2016 | TarPmiR: a new approach for microRNA target site predictionabstractMOTIVATION: The identification of microRNA (miRNA) target sites is fundamentally important for studying gene regulation. There are dozens of computational methods available for miRNA target site prediction. Despite their existence, we still cannot reliably identify miRNA target sites, partially due to our limited understanding of the characteristics of miRNA target sites. The recently published CLASH (crosslinking ligation and sequencing of hybrids) data provide an unprecedented opportunity to study the characteristics of miRNA target sites and improve miRNA target site prediction methods. RESULTS: Applying four different machine learning approaches to the CLASH data, we identified seven new features of miRNA target sites. Combining these new features with those commonly used by existing miRNA target prediction algorithms, we developed an approach called TarPmiR for miRNA target site prediction. Testing on two human and one mouse non-CLASH datasets, we showed that TarPmiR predicted more than 74.2% of true miRNA target sites in each dataset. Compared with three existing approaches, we demonstrated that TarPmiR is superior to these existing approaches in terms of better recall and better precision. AVAILABILITY AND IMPLEMENTATION: The TarPmiR software is freely available at http://hulab.ucf.edu/research/projects/miRNA/TarPmiR/ CONTACTS: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaoman Shawn Li, Haiyan Hu 0004 |
Bioinform. | 3 |
| 2015 | MicroRNA modules prefer to bind weak and unconventional target sitesabstractMOTIVATION: MicroRNAs (miRNAs) play critical roles in gene regulation. Although it is well known that multiple miRNAs may work as miRNA modules to synergistically regulate common target mRNAs, the understanding of miRNA modules is still in its infancy. RESULTS: We employed the recently generated high throughput experimental data to study miRNA modules. We predicted 181 miRNA modules and 306 potential miRNA modules. We observed that the target sites of these predicted modules were in general weaker compared with those not bound by miRNA modules. We also discovered that miRNAs in predicted modules preferred to bind unconventional target sites rather than canonical sites. Surprisingly, contrary to a previous study, we found that most adjacent miRNA target sites from the same miRNA modules were not within the range of 10-130 nucleotides. Interestingly, the distance of target sites bound by miRNAs in the same modules was shorter when miRNA modules bound unconventional instead of canonical sites. Our study shed new light on miRNA binding and miRNA target sites, which will likely advance our understanding of miRNA regulation. AVAILABILITY AND IMPLEMENTATION: The software miRModule can be freely downloaded at http://hulab.ucf.edu/research/projects/miRNA/miRModule. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected] or [email protected]. Xiaoman Shawn Li, Haiyan Hu 0004 |
Bioinform. | 3 |
| 2015 | MBBC: an efficient approach for metagenomic binning based on clusteringabstractBACKGROUND: Binning environmental shotgun reads is one of the most fundamental tasks in metagenomic studies, in which mixed reads from different species or operational taxonomical units (OTUs) are separated into different groups. While dozens of binning methods are available, there is still room for improvement. RESULTS: We developed a novel taxonomy-independent approach called MBBC (Metagenomic Binning Based on Clustering) to cluster environmental shotgun reads, by considering k-mer frequency in reads and Markov properties of the inferred OTUs. Tested on twelve simulated datasets, MBBC reliably estimated the species number, the genome size, and the relative abundance of each species, independent of whether there are errors in reads. Tested on multiple experimental datasets, MBBC outperformed two state-of-the-art taxonomy-independent methods, in terms of the accuracy of the estimated species number, genome sizes, and percentages of correctly assigned reads, among other metrics. CONCLUSIONS: We have developed a novel method for binning metagenomic reads based on clustering. This method is demonstrated to reliably predict species numbers, genome sizes, relative species abundances, and k-mer coverage in simple datasets. Our method also has a high accuracy in read binning. The MBBC software is freely available at http://eecs.ucf.edu/~xiaoman/MBBC/MBBC.html . Haiyan Hu 0004, Xiaoman Shawn Li |
BMC Bioinform. | 2 |
| 2010 | Mining patterns in disease classification forests
Haiyan Hu 0004 |
J. Biomed. Informatics | 1 |
| 2007 | Tree Gibbs Sampler: identifying conserved motifs without aligning orthologous sequencesabstractSUMMARY: Tree Gibbs Sampler is a software for identifying motifs by simultaneously using the motif overrepresentation property and the motif evolutionary conservation property. It identifies motifs without depending on pre-aligned orthologous sequences, which makes it useful for the extraction of regulatory elements in multiple genomes of both closely related and distant species. AVAILABILITY: The Tree Gibbs Sampler software is freely downloadable at https://compbio.iupui.edu/xiaomanli/LiSoftware/retrieve.php?ID=tgs Xiaohui Cai, Haiyan Hu 0004, Xiaoman Shawn Li |
Bioinform. | 2 |
| 2006 | Integrative Array Analyzer: a software package for analysis of cross-platform and cross-species microarray dataabstractThe rapid accumulation of microarray data translates into an urgent need for tools to perform integrative microarray analysis. Integrative Array Analyzer is a comprehensive analysis and visualization software toolkit, which aims to facilitate the reuse of the large amount of cross-platform and cross-species microarray data. It is composed of the data preprocess module, the co-expression analysis module, the differential expression analysis module, the functional and transcriptional annotation module and the graph visualization module. Kiran Kamath, Kangyu Zhang, Sudip Pulapura, Avinash Achar, Juan Nunez-Iglesias, Yu Huang 0003, Xifeng Yan, Jiawei Han 0001, Haiyan Hu 0004, Min Xu 0009, Jianjun Hu, Xianghong Jasmine Zhou |
Bioinform. | 10 |