EDBT 2026 Demo / reviewers in the wild / expert
Yue Li 0017
dblp:61/500-17
· DBLP profile ↗
24ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TrajGPT: Irregular Time-Series Representation Learning of Health TrajectoryabstractIn the healthcare domain, time-series data are often irregularly sampled with varying intervals through outpatient visits, posing challenges for existing models designed for equally spaced sequential data. To address this, we propose Trajectory Generative Pre-trained Transformer (TrajGPT) for representation learning on irregularly-sampled healthcare time series. TrajGPT introduces a novel Selective Recurrent Attention (SRA) module that leverages a data-dependent decay to adaptively filter irrelevant past information. As a discretized ordinary differential equation (ODE) framework, TrajGPT captures underlying continuous dynamics and enables a time-specific inference for forecasting arbitrary target timesteps without auto-regressive prediction. Experimental results based on the longitudinal EHR data PopHR from Montreal health system and eICU from PhysioNet showcase TrajGPT's superior zero-shot performance in disease forecasting, drug usage prediction, and sepsis detection. The inferred trajectories of diabetic and cardiac patients reveal meaningful comorbidity conditions, underscoring TrajGPT as a useful tool for forecasting patient health evolution. Qincheng Lu, David L. Buckeridge, Yue Li 0017 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | A Unified Solution to Diverse Heterogeneities in One-Shot Federated LearningabstractOne-Shot Federated Learning (OSFL) restricts communication between the server and clients to a single round, significantly reducing communication costs and minimizing privacy leakage risks compared to traditional Federated Learning (FL), which requires multiple rounds of communication. However, existing OSFL frameworks remain vulnerable to distributional heterogeneity, as they primarily focus on model heterogeneity while neglecting data heterogeneity. To bridge this gap, we propose FedHydra, a unified, data-free, OSFL framework designed to effectively address both model and data heterogeneity. Unlike existing OSFL approaches, FedHydra introduces a novel two-stage learning mechanism. Specifically, it incorporates model stratification and heterogeneity-aware stratified aggregation to mitigate the challenges posed by both model and data heterogeneity. By this design, the data and model heterogeneity issues are simultaneously monitored from different aspects during learning. Consequently, FedHydra can effectively mitigate both issues by minimizing their inherent conflicts. We compared FedHydra with five SOTA baselines on four benchmark datasets. Experimental results show that our method outperforms the previous OSFL methods in both homogeneous and heterogeneous settings. The code is available at https://github.com/Jun-B0518/FedHydra. Yiliao Song, Di Wu 0050, Atul Sajjanhar, Yong Xiang 0001, Wei Zhou 0044, Xiaohui Tao 0001, Yan Li 0002, Yue Li 0017 |
KDD (2) | 9 |
| 2025 | ConvNTC: convolutional neural tensor completion for detecting "A-A-B" type biological tripletsabstractSystematically investigating interactions among molecules of the same type across different contexts is crucial for unraveling disease mechanisms and developing potential therapeutic strategies. The "A-A-B" triplet paradigm provides a principled approach to model such context-specific interactions, and leveraging third-order tensor to capture such type ternary relationships is an efficient strategy. However, effectively modeling both multilinear and nonlinear characteristics to accurately identify such triplets using tensor-based methods remains a challenge. In this paper, we propose a novel Convolutional Neural Tensor Completion (ConvNTC) framework that collaboratively learns the multilinear and nonlinear representations to model triplet-based network interactions. ConvNTC consists of a multilinear module and a nonlinear module. The former is a tensor decomposition approach that integrates multiple constraints to learn the tensor factor embeddings. The latter contains three components: an embedding generator to produce position-specific index embeddings for each tensor entry in addition to the factor embeddings, a convolutional encoder to perform nonlinear feature mapping while preserving the tensor's rank-one property, and a Kolmogorov-Arnold Network (KAN) based predictor to effectively capture high-dimensional relationships aligned with the intrinsic structure of real-world data. We evaluate ConvNTC on two types triplet datasets of the "A-A-B" type: miRNA-miRNA-disease and drug-drug-cell. Comprehensive experiments against 11 state-of-the-art methods demonstrate the superiority of ConvNTC in terms of triplet prediction. ConvNTC reveals promising prognostic values of the miRNA-miRNA interactions on breast cancer and detects synergistic drug combinations in cancer cell lines. Yue Li 0017, Jiawei Luo 0001 |
Briefings Bioinform. | 3 |
| 2025 | SpaTM: topic models for inferring spatially informed transcriptional programsabstractSpatial transcriptomics enables the contextualization of gene expression with spatial organization, advancing our understanding of development, disease, and tissue architecture. However, existing analysis pipelines require multiple tools to explore spatial domains, and few methods can jointly analyse spatial data from annotation-free and annotation-guided perspectives with high interpretability. We therefore propose the Spatial Topic Model (SpaTM), a topic-modelling framework capable of annotation-guided and annotation-free analysis of spatial transcriptomes. SpaTM can learn gene programs that represent histology-based annotations while also inferring spatial domains with an annotation-free approach if manual annotations are limited or noisy. In benchmarking experiments, SpaTM achieves competitive performance at spatial label prediction and clustering when compared with existing state-of-the-art methods. We demonstrate SpaTM's interpretability by using topic mixtures to capture transcriptional programs in dorsolateral prefrontal cortex and ductal carcinoma samples and show how its intuitive framework facilitates the integration of spatial transcriptomics tasks. Finally, we showcase how SpaTM can extend the analysis of large-scale snRNA-seq atlases in human brains with Major Depressive Disorder. Overall, SpaTM provides a unified and interpretable analysis framework for spatial transcriptomics, enabling competitive performance in multiple tasks while inferring biologically informed gene programs. Adrien Osakwe, Wenqi Dong, Qihuang Zhang, Robert Sladek, Yue Li 0017 |
Briefings Bioinform. | 5 |
| 2025 | Sparse polygenic risk score inference with the spike-and-slab LASSOabstractMOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs. Junyi Song, Shadi Zabad, Archer Yang, Simon Gravel, Yue Li 0017 |
Bioinform. | 5 |
| 2024 | MiRGraph: A hybrid deep learning approach to identify microRNA-target interactions by integrating heterogeneous regulatory network and genomic sequencesabstractMicroRNAs (miRNAs) mediates gene expression regulation by targeting specific messenger RNAs (mRNAs) in the cytoplasm. They can function as both tumor suppressors and oncogenes depending on the specific miRNA and its target genes. Detecting miRNA-target interactions (MTIs) is critical for unraveling the complex mechanisms of gene regulation and promising towards RNA therapy for cancer. There is currently a lack of MTIs prediction methods that simultaneously perform feature learning from heterogeneous gene regulatory network (GRN) and genomic sequences. To improve the prediction performance of MTIs, we present a novel transformer-based multi-view feature learning method – MiRGraph, which consists of two main modules for learning the sequence-based and GRN-based feature embedding. For the former, we utilize the mature miRNA sequences and the complete 3'UTR sequence of the target mRNAs to encode sequence features using a hybrid transformer and convolutional neural network (CNN) (TransCNN) architecture. For the latter, we utilize a heterogeneous graph transformer (HGT) module to extract the relational and structural information from the GRN consisting of miRNA-miRNA, gene-gene and miRNA-target interactions. The TransCNN and HGT modules can be learned end-to-end to predict experimentally validated MTIs from MiRTarBase. MiRGraph outperforms existing methods in not only recapitulating the true MTIs but also in predicting strength of the MTIs based on the in-vitro measurements of miRNA transfections. In a case study on breast cancer, we identified plausible target genes of an oncomir. Ying Liu 0027, Jiawei Luo 0001, Yue Li 0017 |
BIBM | 4 |
| 2024 | Cell ontology guided transcriptome foundation modelabstractTranscriptome foundation models (TFMs) hold great promises of deciphering the transcriptomic language that dictate diverse cell functions by self-supervised learning on large-scale single-cell gene expression data, and ultimately unraveling the complex mechanisms of human diseases. However, current TFMs treat cells as independent samples and ignore the taxonomic relationships between cell types, which are available in cell ontology graphs. We argue that effectively leveraging this ontology information during the TFM pre-training can improve learning biologically meaningful gene co-expression patterns while preserving TFM as a general purpose foundation model for downstream zero-shot and fine-tuning tasks. To this end, we present **s**ingle **c**ell, **Cell**-**o**ntology guided TFM (scCello). We introduce cell-type coherence loss and ontology alignment loss, which are minimized along with the masked gene expression prediction loss during the pre-training. The novel loss component guide scCello to learn the cell-type-specific representation and the structural relation between cell types from the cell ontology graph, respectively. We pre-trained scCello on 22 million cells from CellxGene database leveraging their cell-type labels mapped to the cell ontology graph from Open Biological and Biomedical Ontology Foundry. Our TFM demonstrates competitive generalization and transferability performance over the existing TFMs on biologically important tasks including identifying novel cell types of unseen cells, prediction of cell-type-specific marker genes, and cancer drug responses. Source code and model
weights are available at https://github.com/DeepGraphLearning/scCello. Xinyu Yuan, Zhihao Zhan, Zuobai Zhang, Manqi Zhou, Jianan Zhao 0002, Yue Li 0017, Jian Tang 0005 |
NeurIPS | 7 |
| 2024 | GFETM: Genome Foundation-Based Embedded Topic Model for scATAC-seq Modeling
Yimin Fan, Yu Li 0006, Yue Li 0017 |
RECOMB | 4 |
| 2024 | SharePro: an accurate and efficient genetic colocalization method accounting for multiple causal signalsabstractMOTIVATION: Colocalization analysis is commonly used to assess whether two or more traits share the same genetic signals identified in genome-wide association studies (GWAS), and is important for prioritizing targets for functional follow-up of GWAS results. Existing colocalization methods can have suboptimal performance when there are multiple causal variants in one genomic locus. RESULTS: We propose SharePro to extend the COLOC framework for colocalization analysis. SharePro integrates linkage disequilibrium (LD) modeling and colocalization assessment by grouping correlated variants into effect groups. With an efficient variational inference algorithm, posterior colocalization probabilities can be accurately estimated. In simulation studies, SharePro demonstrated increased power with a well-controlled false positive rate at a low computational cost. Compared to existing methods, SharePro provided stronger and more consistent colocalization evidence for known lipid-lowering drug target proteins and their corresponding lipid traits. Through an additional challenging case of the colocalization analysis of the circulating abundance of R-spondin 3 GWAS and estimated bone mineral density GWAS, we demonstrated the utility of SharePro in identifying biologically plausible colocalized signals. AVAILABILITY AND IMPLEMENTATION: SharePro for colocalization analysis is written in Python and openly available at https://github.com/zhwm/SharePro_coloc. Wenmin Zhang, Tianyuan Lu, Robert Sladek, Yue Li 0017, Hamed S. Najafabadi, Josée Dupuis |
Bioinform. | 4 |
| 2024 | MixEHR-SurG: A joint proportional hazard and guided topic model for inferring mortality-associated topics from electronic health recordsabstractSurvival models can help medical practitioners to evaluate the prognostic importance of clinical variables to patient outcomes such as mortality or hospital readmission and subsequently design personalized treatment regimes. Electronic Health Records (EHRs) hold the promise for large-scale survival analysis based on systematically recorded clinical features for each patient. However, existing survival models either do not scale to high dimensional and multi-modal EHR data or are difficult to interpret. In this study, we present a supervised topic model called MixEHR-SurG to simultaneously integrate heterogeneous EHR data and model survival hazard. Our contributions are three-folds: (1) integrating EHR topic inference with Cox proportional hazards likelihood; (2) integrating patient-specific topic hyperparameters using the PheCode concepts such that each topic can be identified with exactly one PheCode-associated phenotype; (3) multi-modal survival topic inference. This leads to a highly interpretable survival topic model that can infer PheCode-specific phenotype topics associated with patient mortality. We evaluated MixEHR-SurG using a simulated dataset and two real-world EHR datasets: the Quebec Congenital Heart Disease (CHD) data consisting of 8211 subjects with 75,187 outpatient claim records of 1767 unique ICD codes; the MIMIC-III consisting of 1458 subjects with multi-modal EHR records. Compared to the baselines, MixEHR-SurG achieved a superior dynamic AUROC for mortality prediction, with a mean AUROC score of 0.89 in the simulation dataset and a mean AUROC of 0.645 on the CHD dataset. Qualitatively, MixEHR-SurG associates severe cardiac conditions with high mortality risk among the CHD patients after the first heart failure hospitalization and critical brain injuries with increased mortality among the MIMIC-III patients after their ICU discharge. Together, the integration of the Cox proportional hazards model and EHR topic inference in MixEHR-SurG not only leads to competitive mortality prediction but also meaningful phenotype topics for in-depth survival analysis. The software is available at GitHub: https://github.com/li-lab-mcgill/MixEHR-SurG. Yi Yang 0046, Ariane Marelli, Yue Li 0017 |
J. Biomed. Informatics | 4 |
| 2022 | Automatic Phenotyping by a Seed-guided Topic ModelabstractElectronic health records (EHRs) provide rich clinical information and the opportunities to extract epidemiological patterns to understand and predict patient disease risks with suitable machine learning methods such as topic models. However, existing topic models do not generate identifiable topics each predicting a unique phenotype. One promising direction is to use known phenotype concepts to guide topic inference. We present a seed-guided Bayesian topic model called MixEHR-Seed with 3 contributions: (1) for each phenotype, we infer a dual-form of topic distribution: a seed-topic distribution over a small set of key EHR codes and a regular topic distribution over the entire EHR vocabulary; (2) we model age-dependent disease progression as Markovian dynamic topic priors; (3) we infer seed-guided multi-modal topics over distinct EHR data types. For inference, we developed a variational inference algorithm. Using MixEHR-Seed, we inferred 1569 PheCode-guided phenotype topics from an EHR database in Quebec, Canada covering 1.3 million patients for up to 20-year follow-up with 122 million records for 8539 and 1126 unique diagnostic and drug codes, respectively. We observed (1) accurate phenotype prediction by the guided topics, (2) clinically relevant PheCode-guided disease topics, (3) meaningful age-dependent disease prevalence. Source code is available at GitHub: https://github.com/li-lab-mcgill/MixEHR-Seed. Yuanyi Hu, Aman Verma, David L. Buckeridge, Yue Li 0017 |
KDD | 5 |
| 2022 | MixEHR-Guided: A guided multi-modal topic modeling approach for large-scale automatic phenotyping using the electronic health recordabstractElectronic Health Records (EHRs) contain rich clinical data collected at the point of the care, and their increasing adoption offers exciting opportunities for clinical informatics, disease risk prediction, and personalized treatment recommendation. However, effective use of EHR data for research and clinical decision support is often hampered by a lack of reliable disease labels. To compile gold-standard labels, researchers often rely on clinical experts to develop rule-based phenotyping algorithms from billing codes and other surrogate features. This process is tedious and error-prone due to recall and observer biases in how codes and measures are selected, and some phenotypes are incompletely captured by a handful of surrogate features. To address this challenge, we present a novel automatic phenotyping model called MixEHR-Guided (MixEHR-G), a multimodal hierarchical Bayesian topic model that efficiently models the EHR generative process by identifying latent phenotype structure in the data. Unlike existing topic modeling algorithms wherein the inferred topics are not identifiable, MixEHR-G uses prior information from informative surrogate features to align topics with known phenotypes. We applied MixEHR-G to an openly-available EHR dataset of 38,597 intensive care patients (MIMIC-III) in Boston, USA and to administrative claims data for a population-based cohort (PopHR) of 1.3 million people in Quebec, Canada. Qualitatively, we demonstrate that MixEHR-G learns interpretable phenotypes and yields meaningful insights about phenotype similarities, comorbidities, and epidemiological associations. Quantitatively, MixEHR-G outperforms existing unsupervised phenotyping methods on a phenotype label annotation task, and it can accurately estimate relative phenotype prevalence functions without gold-standard phenotype information. Altogether, MixEHR-G is an important step towards building an interpretable and automated phenotyping system using EHR data. Yuri Ahuja, Yuesong Zou, Aman Verma, David L. Buckeridge, Yue Li 0017 |
J. Biomed. Informatics | 5 |
| 2021 | Deep feature extraction of single-cell transcriptomes by generative adversarial networkabstractMOTIVATION: Single-cell RNA-sequencing (scRNA-seq) offers the opportunity to dissect heterogeneous cellular compositions and interrogate the cell-type-specific gene expression patterns across diverse conditions. However, batch effects such as laboratory conditions and individual-variability hinder their usage in cross-condition designs. RESULTS: Here, we present a single-cell Generative Adversarial Network (scGAN) to simultaneously acquire patterns from raw data while minimizing the confounding effect driven by technical artifacts or other factors inherent to the data. Specifically, scGAN models the data likelihood of the raw scRNA-seq counts by projecting each cell onto a latent embedding. Meanwhile, scGAN attempts to minimize the correlation between the latent embeddings and the batch labels across all cells. We demonstrate scGAN on three public scRNA-seq datasets and show that our method confers superior performance over the state-of-the-art methods in forming clusters of known cell types and identifying known psychiatric genes that are associated with major depressive disorder. AVAILABILITYAND IMPLEMENTATION: The scGAN code and the information for the public scRNA-seq datasets are available at https://github.com/li-lab-mcgill/singlecell-deepfeature. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mojtaba Bahrami, Malosree Maitra, Corina Nagy, Gustavo Turecki, Hamid R. Rabiee 0001, Yue Li 0017 |
Bioinform. | 6 |
| 2020 | Supervised learning on phylogenetically distributed dataabstractMOTIVATION: The ability to develop robust machine-learning (ML) models is considered imperative to the adoption of ML techniques in biology and medicine fields. This challenge is particularly acute when data available for training is not independent and identically distributed (iid), in which case trained models are vulnerable to out-of-distribution generalization problems. Of particular interest are problems where data correspond to observations made on phylogenetically related samples (e.g. antibiotic resistance data). RESULTS: We introduce DendroNet, a new approach to train neural networks in the context of evolutionary data. DendroNet explicitly accounts for the relatedness of the training/testing data, while allowing the model to evolve along the branches of the phylogenetic tree, hence accommodating potential changes in the rules that relate genotypes to phenotypes. Using simulated data, we demonstrate that DendroNet produces models that can be significantly better than non-phylogenetically aware approaches. DendroNet also outperforms other approaches at two biological tasks of significant practical importance: antiobiotic resistance prediction in bacteria and trophic level prediction in fungi. AVAILABILITY AND IMPLEMENTATION: https://github.com/BlanchetteLab/DendroNet. Elliot Layne, Erika N. Dort, Richard C. Hamelin, Yue Li 0017, Mathieu Blanchette |
Bioinform. | 4 |
| 2017 | Evolving Transcription Factor Binding Site Models From Protein Binding Microarray DataabstractProtein binding microarray (PBM) is a high-throughput platform that can measure the DNA binding preference of a protein in a comprehensive and unbiased manner. In this paper, we describe the PBM motif model building problem. We apply several evolutionary computation methods and compare their performance with the interior point method, demonstrating their performance advantages. In addition, given the PBM domain knowledge, we propose and describe a novel method called kmerGA which makes domain-specific assumptions to exploit PBM data properties to build more accurate models than the other models built. The effectiveness and robustness of kmerGA is supported by comprehensive performance benchmarking on more than 200 datasets, time complexity analysis, convergence analysis, parameter analysis, and case studies. To demonstrate its utility further, kmerGA is applied to two real world applications: 1) PBM rotation testing and 2) ChIP-Seq peak sequence prediction. The results support the biological relevance of the models learned by kmerGA, and thus its real world applicability. Ka-Chun Wong, Chengbin Peng 0001, Yue Li 0017 |
IEEE Trans. Cybern. | 3 |
| 2016 | Identification of coupling DNA motif pairs on long-range chromatin interactions in human K562 cellsabstractMOTIVATION: The protein-DNA interactions between transcription factors (TFs) and transcription factor binding sites (TFBSs, also known as DNA motifs) are critical activities in gene transcription. The identification of the DNA motifs is a vital task for downstream analysis. Unfortunately, the long-range coupling information between different DNA motifs is still lacking. To fill the void, as the first-of-its-kind study, we have identified the coupling DNA motif pairs on long-range chromatin interactions in human. RESULTS: The coupling DNA motif pairs exhibit substantially higher DNase accessibility than the background sequences. Half of the DNA motifs involved are matched to the existing motif databases, although nearly all of them are enriched with at least one gene ontology term. Their motif instances are also found statistically enriched on the promoter and enhancer regions. Especially, we introduce a novel measurement called motif pairing multiplicity which is defined as the number of motifs that are paired with a given motif on chromatin interactions. Interestingly, we observe that motif pairing multiplicity is linked to several characteristics such as regulatory region type, motif sequence degeneracy, DNase accessibility and pairing genomic distance. Taken into account together, we believe the coupling DNA motif pairs identified in this study can shed lights on the gene transcription mechanism under long-range chromatin interactions. AVAILABILITY AND IMPLEMENTATION: The identified motif pair data is compressed and available in the supplementary materials associated with this manuscript. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001 |
Bioinform. | 2 |
| 2016 | A Novel Method to Detect Functional microRNA Regulatory Modules by Bicliques MergingabstractUNLABELLED: MicroRNAs (miRNAs) are post-transcriptional regulators that repress the expression of their targets. They are known to work cooperatively with genes and play important roles in numerous cellular processes. Identification of miRNA regulatory modules (MRMs) would aid deciphering the combinatorial effects derived from the many-to-many regulatory relationships in complex cellular systems. Here, we develop an effective method called BiCliques Merging (BCM) to predict MRMs based on bicliques merging. By integrating the miRNA/mRNA expression profiles from The Cancer Genome Atlas (TCGA) with the computational target predictions, we construct a weighted miRNA regulatory network for module discovery. The maximal bicliques detected in the network are statistically evaluated and filtered accordingly. We then employed a greedy-based strategy to iteratively merge the remaining bicliques according to their overlaps together with edge weights and the gene-gene interactions. Comparing with existing methods on two cancer datasets from TCGA, we showed that the modules identified by our method are more densely connected and functionally enriched. Moreover, our predicted modules are more enriched for miRNA families and the miRNA-mRNA pairs within the modules are more negatively correlated. Finally, several potential prognostic modules are revealed by Kaplan-Meier survival analysis and breast cancer subtype analysis. AVAILABILITY: BCM is implemented in Java and available for download in the supplementary materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/ TCBB.2015.2462370. Cheng Liang 0001, Yue Li 0017, Jiawei Luo 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2016 | A Comparison Study for DNA Motif Modeling on Protein Binding MicroarrayabstractTranscription factor binding sites (TFBSs) are relatively short (5-15 bp) and degenerate. Identifying them is a computationally challenging task. In particular, protein binding microarray (PBM) is a high-throughput platform that can measure the DNA binding preference of a protein in a comprehensive and unbiased manner; for instance, a typical PBM experiment can measure binding signal intensities of a protein to all possible DNA k-mers (k = 8∼10). Since proteins can often bind to DNA with different binding intensities, one of the major challenges is to build TFBS (also known as DNA motif) models which can fully capture the quantitative binding affinity data. To learn DNA motif models from the non-convex objective function landscape, several optimization methods are compared and applied to the PBM motif model building problem. In particular, representative methods from different optimization paradigms have been chosen for modeling performance comparison on hundreds of PBM datasets. The results suggest that the multimodal optimization methods are very effective for capturing the binding preference information from PBM data. In particular, we observe a general performance improvement if choosing di-nucleotide modeling over mono-nucleotide modeling. In addition, the models learned by the best-performing method are applied to two independent applications: PBM probe rotation testing and ChIP-Seq peak sequence prediction, demonstrating its biological applicability. Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001, Hau-San Wong |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | A novel motif-discovery algorithm to identify co-regulatory motifs in large transcription factor and microRNA co-regulatory networks in humanabstractMOTIVATION: Interplays between transcription factors (TFs) and microRNAs (miRNAs) in gene regulation are implicated in various physiological processes. It is thus important to identify biologically meaningful network motifs involving both types of regulators to understand the key co-regulatory mechanisms underlying the cellular identity and function. However, existing motif finders do not scale well for large networks and are not designed specifically for co-regulatory networks. RESULTS: In this study, we propose a novel algorithm CoMoFinder to accurately and efficiently identify composite network motifs in genome-scale co-regulatory networks. We define composite network motifs as network patterns involving at least one TF, one miRNA and one target gene that are statistically significant than expected. Using two published disease-related co-regulatory networks, we show that CoMoFinder outperforms existing methods in both accuracy and robustness. We then applied CoMoFinder to human TF-miRNA co-regulatory network derived from The Encyclopedia of DNA Elements project and identified 44 recurring composite network motifs of size 4. The functional analysis revealed that genes involved in the 44 motifs are enriched for significantly higher number of biological processes or pathways comparing with non-motifs. We further analyzed the identified composite bi-fan motif and showed that gene pairs involved in this motif structure tend to physically interact and are functionally more similar to each other than expected. AVAILABILITY AND IMPLEMENTATION: CoMoFinder is implemented in Java and available for download at http://www.cs.utoronto.ca/∼yueli/como.html. Cheng Liang 0001, Yue Li 0017, Jiawei Luo 0001, Zhaolei Zhang |
Bioinform. | 2 |
| 2015 | SignalSpider: probabilistic pattern discovery on multiple normalized ChIP-Seq signal profilesabstractMOTIVATION: Chromatin immunoprecipitation (ChIP) followed by high-throughput sequencing (ChIP-Seq) measures the genome-wide occupancy of transcription factors in vivo. Different combinations of DNA-binding protein occupancies may result in a gene being expressed in different tissues or at different developmental stages. To fully understand the functions of genes, it is essential to develop probabilistic models on multiple ChIP-Seq profiles to decipher the combinatorial regulatory mechanisms by multiple transcription factors. RESULTS: In this work, we describe a probabilistic model (SignalSpider) to decipher the combinatorial binding events of multiple transcription factors. Comparing with similar existing methods, we found SignalSpider performs better in clustering promoter and enhancer regions. Notably, SignalSpider can learn higher-order combinatorial patterns from multiple ChIP-Seq profiles. We have applied SignalSpider on the normalized ChIP-Seq profiles from the ENCODE consortium and learned model instances. We observed different higher-order enrichment and depletion patterns across sets of proteins. Those clustering patterns are supported by Gene Ontology (GO) enrichment, evolutionary conservation and chromatin interaction enrichment, offering biological insights for further focused studies. We also proposed a specific enrichment map visualization method to reveal the genome-wide transcription factor combinatorial patterns from the models built, which extend our existing fine-scale knowledge on gene regulation to a genome-wide level. AVAILABILITY AND IMPLEMENTATION: The matrix-algebra-optimized executables and source codes are available at the authors' websites: http://www.cs.toronto.edu/∼wkc/SignalSpider. Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001, Zhaolei Zhang |
Bioinform. | 2 |
| 2015 | Probabilistic Inference on Multiple Normalized Signal Profiles from Next Generation Sequencing: Transcription Factor Binding SitesabstractWith the prevalence of chromatin immunoprecipitation (ChIP) with sequencing (ChIP-Seq) technology, massive ChIP-Seq data has been accumulated. The ChIP-Seq technology measures the genome-wide occupancy of DNA-binding proteins in vivo. It is well-known that different DNA-binding protein occupancies may result in a gene being regulated in different conditions (e.g. different cell types). To fully understand a gene's function, it is essential to develop probabilistic models on multiple ChIP-Seq profiles for deciphering the gene transcription causalities. In this work, we propose and describe two probabilistic models. Assuming the conditional independence of different DNA-binding proteins' occupancies, the first method (SignalRanker) is developed as an intuitive method for ChIP-Seq genome-wide signal profile inference. Unfortunately, such an assumption may not always hold in some gene regulation cases. Thus, we propose and describe another method (FullSignalRanker) which does not make the conditional independence assumption. The proposed methods are compared with other existing methods on ENCODE ChIP-Seq datasets, demonstrating its regression and classification ability. The results suggest that FullSignalRanker is the best-performing method for recovering the signal ranks on the promoter and enhancer regions. In addition, FullSignalRanker is also the best-performing method for peak sequence classification. We envision that SignalRanker and FullSignalRanker will become important in the era of next generation sequencing. FullSignalRanker program is available on the following website: http://www.cs.toronto.edu/~wkc/FullSignalRanker/. Ka-Chun Wong, Chengbin Peng 0001, Yue Li 0017 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2014 | A probabilistic approach to explore human miRNA targetome by integrating miRNA-overexpression data and sequence informationabstractMOTIVATION: Systematic identification of microRNA (miRNA) targets remains a challenge. The miRNA overexpression coupled with genome-wide expression profiling is a promising new approach and calls for a new method that integrates expression and sequence information. RESULTS: We developed a probabilistic scoring method called targetScore. TargetScore infers miRNA targets as the transformed fold-changes weighted by the Bayesian posteriors given observed target features. To this end, we compiled 84 datasets from Gene Expression Omnibus corresponding to 77 human tissue or cells and 113 distinct transfected miRNAs. Comparing with other methods, targetScore achieves significantly higher accuracy in identifying known targets in most tests. Moreover, the confidence targets from targetScore exhibit comparable protein downregulation and are more significantly enriched for Gene Ontology terms. Using targetScore, we explored oncomir-oncogenes network and predicted several potential cancer-related miRNA-messenger RNA interactions. AVAILABILITY AND IMPLEMENTATION: TargetScore is available at Bioconductor: http://www.bioconductor.org/packages/devel/bioc/html/TargetScore.html. Yue Li 0017, Anna Goldenberg, Ka-Chun Wong, Zhaolei Zhang |
Bioinform. | 1 |
| 2014 | Mirsynergy: detecting synergistic miRNA regulatory modules by overlapping neighbourhood expansionabstractMOTIVATION: Identification of microRNA regulatory modules (MiRMs) will aid deciphering aberrant transcriptional regulatory network in cancer but is computationally challenging. Existing methods are stochastic or require a fixed number of regulatory modules. RESULTS: We propose Mirsynergy, an efficient deterministic overlapping clustering algorithm adapted from a recently developed framework. Mirsynergy operates in two stages: it first forms MiRMs based on co-occurring microRNA (miRNA) targets and then expands each MiRM by greedily including (excluding) mRNAs into (from) the MiRM to maximize the synergy score, which is a function of miRNA-mRNA and gene-gene interactions. Using expression data for ovarian, breast and thyroid cancer from The Cancer Genome Atlas, we compared Mirsynergy with internal controls and existing methods. Mirsynergy-MiRMs exhibit significantly higher functional enrichment and more coherent miRNA-mRNA expression anti-correlation. Based on Kaplan-Meier survival analysis, we proposed several prognostically promising MiRMs and envisioned their utility in cancer research. AVAILABILITY AND IMPLEMENTATION: Mirsynergy is implemented/available as an R/Bioconductor package at www.cs.utoronto.ca/∼yueli/Mirsynergy.html. Yue Li 0017, Cheng Liang 0001, Ka-Chun Wong, Jiawei Luo 0001, Zhaolei Zhang |
Bioinform. | 1 |
| 2014 | Regression Analysis of Combined Gene Expression Regulation in Acute Myeloid LeukemiaabstractGene expression is a combinatorial function of genetic/epigenetic factors such as copy number variation (CNV), DNA methylation (DM), transcription factors (TF) occupancy, and microRNA (miRNA) post-transcriptional regulation. At the maturity of microarray/sequencing technologies, large amounts of data measuring the genome-wide signals of those factors became available from Encyclopedia of DNA Elements (ENCODE) and The Cancer Genome Atlas (TCGA). However, there is a lack of an integrative model to take full advantage of these rich yet heterogeneous data. To this end, we developed RACER (Regression Analysis of Combined Expression Regulation), which fits the mRNA expression as response using as explanatory variables, the TF data from ENCODE, and CNV, DM, miRNA expression signals from TCGA. Briefly, RACER first infers the sample-specific regulatory activities by TFs and miRNAs, which are then used as inputs to infer specific TF/miRNA-gene interactions. Such a two-stage regression framework circumvents a common difficulty in integrating ENCODE data measured in generic cell-line with the sample-specific TCGA measurements. As a case study, we integrated Acute Myeloid Leukemia (AML) data from TCGA and the related TF binding data measured in K562 from ENCODE. As a proof-of-concept, we first verified our model formalism by 10-fold cross-validation on predicting gene expression. We next evaluated RACER on recovering known regulatory interactions, and demonstrated its superior statistical power over existing methods in detecting known miRNA/TF targets. Additionally, we developed a feature selection procedure, which identified 18 regulators, whose activities clustered consistently with cytogenetic risk groups. One of the selected regulators is miR-548p, whose inferred targets were significantly enriched for leukemia-related pathway, implicating its novel role in AML pathogenesis. Moreover, survival analysis using the inferred activities identified C-Fos as a potential AML prognostic marker. Together, we provided a novel framework that successfully integrated the TCGA and ENCODE data in revealing AML-specific regulatory program at global level. Yue Li 0017, Minggao Liang, Zhaolei Zhang |
PLoS Comput. Biol. | 1 |