EDBT 2026 Demo / reviewers in the wild / expert
Manolis Kellis
dblp:75/2690
· DBLP profile ↗
16ranked-venue papers
0as first author
3since 2021 · last 2024
0000-0001-7113-9630ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
11 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
2 papers |
Trustworthy machine learning · 61% Generative modeling · 30% Probabilistic and Bayesian machine learning · 9% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 54% Graph algorithms and graph theory · 46% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
phylogenetics |
1.0 | 5 | 2018 | RANGER-DTL 2.0: rigorous reconstruction of gene-family evolution by duplication, transfer and loss · Bioinform. 2018 Improved gene tree error correction in the presence of horizontal gene transfer · Bioinform. 2015 Pareto-optimal phylogenetic tree reconciliation · Bioinform. 2014 |
Bioinformatics and computational biology › phylogenetics
gene tree reconciliation |
0.8 | 4 | 2018 | RANGER-DTL 2.0: rigorous reconstruction of gene-family evolution by duplication, transfer and loss · Bioinform. 2018 Pareto-optimal phylogenetic tree reconciliation · Bioinform. 2014 Reconciliation Revisited: Handling Multiple Optima When Reconciling with Duplication, Transfer, and Loss · RECOMB 2013 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | A versatile informative diffusion model for single-cell ATAC-seq data generation and analysis · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
large language model trustworthiness |
0.8 | 1 | 2024 | Position: TrustLLM: Trustworthiness in Large Language Models · ICML 2024 |
Bioinformatics and computational biology › single-cell analysis › single-cell epigenomics
single-cell ATAC-seq analysis |
0.8 | 1 | 2024 | A versatile informative diffusion model for single-cell ATAC-seq data generation and analysis · NeurIPS 2024 |
Bioinformatics and computational biology › phylogenetics › gene tree reconciliation
duplication-transfer-loss reconciliation |
0.6 | 3 | 2018 | RANGER-DTL 2.0: rigorous reconstruction of gene-family evolution by duplication, transfer and loss · Bioinform. 2018 Reconciliation Revisited: Handling Multiple Optima When Reconciling with Duplication, Transfer, and Loss · RECOMB 2013 Efficient algorithms for the reconciliation problem with gene duplication, horizontal transfer and loss · Bioinform. 2012 |
Bioinformatics and computational biology › cancer genomics
cancer driver gene identification |
0.4 | 1 | 2019 | ncdDetect2: improved models of the site-specific mutation rate in cancer and driver detection with robust significance evaluation · Bioinform. 2019 |
Bioinformatics and computational biology
cancer genomics |
0.4 | 1 | 2019 | ncdDetect2: improved models of the site-specific mutation rate in cancer and driver detection with robust significance evaluation · Bioinform. 2019 |
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics › RNA structure prediction
RNA secondary structure prediction |
0.2 | 1 | 2016 | SwiSpot: modeling riboswitches by spotting out switching sequences · Bioinform. 2016 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.2 | 1 | 2024 | A versatile informative diffusion model for single-cell ATAC-seq data generation and analysis · NeurIPS 2024 |
Bioinformatics and computational biology › phylogenetics › phylogenetic inference
gene tree inference |
0.2 | 1 | 2015 | Improved gene tree error correction in the presence of horizontal gene transfer · Bioinform. 2015 |
Bioinformatics and computational biology › metagenomics
horizontal gene transfer detection |
0.2 | 1 | 2015 | Improved gene tree error correction in the presence of horizontal gene transfer · Bioinform. 2015 |
Mathematical optimization
combinatorial optimization |
0.2 | 1 | 2013 | Reconciliation Revisited: Handling Multiple Optima When Reconciling with Duplication, Transfer, and Loss · RECOMB 2013 |
Bioinformatics and computational biology › molecular evolution
gene family evolution |
0.1 | 1 | 2012 | Efficient algorithms for the reconciliation problem with gene duplication, horizontal transfer and loss · Bioinform. 2012 |
Bioinformatics and computational biology › epigenomics › chromatin analysis
chromatin state analysis |
0.1 | 1 | 2011 | Discovery and Characterization of Chromatin States for Systematic Annotation of the Human Genome · RECOMB 2011 |
Bioinformatics and computational biology
comparative genomics |
0.1 | 1 | 2011 | PhyloCSF: a comparative genomics method to distinguish protein coding and non-coding regions · Bioinform. 2011 |
Bioinformatics and computational biology
epigenomics |
0.1 | 1 | 2011 | Discovery and Characterization of Chromatin States for Systematic Annotation of the Human Genome · RECOMB 2011 |
Bioinformatics and computational biology
genome annotation |
0.1 | 1 | 2011 | Discovery and Characterization of Chromatin States for Systematic Annotation of the Human Genome · RECOMB 2011 |
Bioinformatics and computational biology › genome annotation
protein coding region prediction |
0.1 | 1 | 2011 | PhyloCSF: a comparative genomics method to distinguish protein coding and non-coding regions · Bioinform. 2011 |
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA structure |
0.1 | 1 | 2016 | SwiSpot: modeling riboswitches by spotting out switching sequences · Bioinform. 2016 |
Bioinformatics and computational biology › network bioinformatics
biological network analysis |
0.1 | 1 | 2007 | Network Motif Discovery Using Subgraph Enumeration and Symmetry-Breaking · RECOMB 2007 |
Bioinformatics and computational biology › biological network
network biology |
0.1 | 1 | 2007 | Network Motif Discovery Using Subgraph Enumeration and Symmetry-Breaking · RECOMB 2007 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis › network topology analysis
network motif discovery |
0.1 | 1 | 2007 | Network Motif Discovery Using Subgraph Enumeration and Symmetry-Breaking · RECOMB 2007 |
Graph algorithms and graph theory
graph algorithms |
0.1 | 1 | 2007 | Network Motif Discovery Using Subgraph Enumeration and Symmetry-Breaking · RECOMB 2007 |
Graph algorithms and graph theory › graph algorithms
subgraph enumeration |
0.1 | 1 | 2007 | Network Motif Discovery Using Subgraph Enumeration and Symmetry-Breaking · RECOMB 2007 |
Bioinformatics and computational biology › transcriptomics › non-coding RNA analysis
long non-coding RNA identification |
0.0 | 1 | 2011 | PhyloCSF: a comparative genomics method to distinguish protein coding and non-coding regions · Bioinform. 2011 |
Bioinformatics and computational biology
transcriptomics |
0.0 | 1 | 2011 | PhyloCSF: a comparative genomics method to distinguish protein coding and non-coding regions · Bioinform. 2011 |
Methods — techniques the papers use, named apart from their topics
mutual information regularization · 1.5gaussian mixture model · 1.5diffusion model · 1.5evaluation study · 0.8benchmark construction · 0.8site-specific mutation rate model · 0.4bayesian posterior ranking · 0.4maximum likelihood estimation · 0.3homology-based search · 0.2configuration-based pairing evaluation · 0.2statistical hypothesis testing · 0.2shimodaira-hasegawa test · 0.2reconciliation algorithms · 0.2symmetry breaking · 0.1subgraph enumeration · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Position: TrustLLM: Trustworthiness in Large Language ModelsabstractLarge language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs, and discussion of open challenges and future directions. Specifically, we first propose a set of principles for trustworthy LLMs that span eight different dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. We then present a study evaluating 16 mainstream LLMs in TrustLLM, consisting of over 30 datasets. Our findings firstly show that in general trustworthiness and capability (i.e., functional effectiveness) are positively related. Secondly, our observations reveal that proprietary LLMs generally outperform most open-source counterparts in terms of trustworthiness, raising concerns about the potential risks of widely accessible open-source LLMs. However, a few open-source LLMs come very close to proprietary ones, suggesting that open-source models can achieve high levels of trustworthiness without additional mechanisms like moderator, offering valuable insights for developers in this field. Thirdly, it is important to note that some LLMs may be overly calibrated towards exhibiting trustworthiness, to the extent that they compromise their utility by mistakenly treating benign prompts as harmful and consequently not responding. Besides these observations, we’ve uncovered key insights into the multifaceted trustworthiness in LLMs. We emphasize the importance of ensuring transparency not only in the models themselves but also in the technologies that underpin trustworthiness. We advocate that the establishment of an AI alliance between industry, academia, the open-source community to foster collaboration is imperative to advance the trustworthiness of LLMs. Yue Huang 0001, Lichao Sun 0001, Haoran Wang 0005, Siyuan Wu 0001, Qihui Zhang, Chujie Gao, Wenhan Lyu, Yixuan Zhang 0001, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu 0002, Yijue Wang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Heng Ji 0001, Hongyi Wang 0001, Huan Zhang 0001, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang 0001, Mohit Bansal, James Zou 0001, Jian Pei 0001, Jianfeng Gao 0001, Jiawei Han 0001, Jieyu Zhao 0001, Jiliang Tang, Jindong Wang 0001, Joaquin Vanschoren, John C. Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang 0001, Lifang He 0001, Lifu Huang, Michael Backes 0001, Neil Zhenqiang Gong, Philip S. Yu, Quanquan Gu, Ran Xu 0001, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen 0001, Tianming Liu 0001, Tianyi Zhou 0001, William Yang Wang, Xiang Li 0001, Xiangliang Zhang 0001, Xiao Wang 0012, Xing Xie 0001, Xuyu Wang, Yan Liu 0002, Yanfang Ye 0001, Yinzhi Cao, Yong Chen 0016, Yue Zhao 0016 |
ICML | 29 |
| 2024 | A versatile informative diffusion model for single-cell ATAC-seq data generation and analysisabstractThe rapid advancement of single-cell ATAC sequencing (scATAC-seq) technologies holds great promise for investigating the heterogeneity of epigenetic landscapes at the cellular level. The amplification process in scATAC-seq experiments often introduces noise due to dropout events, which results in extreme sparsity that hinders accurate analysis. Consequently, there is a significant demand for the generation of high-quality scATAC-seq data in silico. Furthermore, current methodologies are typically task-specific, lacking a versatile framework capable of handling multiple tasks within a single model. In this work, we propose ATAC-Diff, a versatile framework, which is based on a diffusion model conditioned on the latent auxiliary variables to adapt for various tasks. ATAC-Diff is the first diffusion model for the scATAC-seq data generation and analysis, composed of auxiliary modules encoding the latent high-level variables to enable the model to learn the semantic information to sample high-quality data. Gaussian Mixture Model (GMM) as the latent prior and auxiliary decoder, the yield variables reserve the refined genomic information beneficial for downstream analyses. Another innovation is the incorporation of mutual information between observed and hidden variables as a regularization term to prevent the model from decoupling from latent variables. Through extensive experiments, we demonstrate that ATAC-Diff achieves high performance in both generation and analysis tasks, outperforming state-of-the-art models. Zunpeng Liu, Ka-Chun Wong, Manolis Kellis |
NeurIPS | 6 |
| 2024 | Single-cell RNA sequencing data imputation using bi-level feature propagationabstractSingle-cell RNA sequencing (scRNA-seq) enables the exploration of cellular heterogeneity by analyzing gene expression profiles in complex tissues. However, scRNA-seq data often suffer from technical noise, dropout events and sparsity, hindering downstream analyses. Although existing works attempt to mitigate these issues by utilizing graph structures for data denoising, they involve the risk of propagating noise and fall short of fully leveraging the inherent data relationships, relying mainly on one of cell-cell or gene-gene associations and graphs constructed by initial noisy data. To this end, this study presents single-cell bilevel feature propagation (scBFP), two-step graph-based feature propagation method. It initially imputes zero values using non-zero values, ensuring that the imputation process does not affect the non-zero values due to dropout. Subsequently, it denoises the entire dataset by leveraging gene-gene and cell-cell relationships in the respective steps. Extensive experimental results on scRNA-seq data demonstrate the effectiveness of scBFP in various downstream tasks, uncovering valuable biological insights. Junseok Lee 0002, Sukwon Yun, Yeongmin Kim, Tianlong Chen 0001, Manolis Kellis, Chanyoung Park 0001 |
Briefings Bioinform. | 5 |
| 2019 | ncdDetect2: improved models of the site-specific mutation rate in cancer and driver detection with robust significance evaluationabstractMotivation: Understanding the mutational processes that act during cancer development is a key topic of cancer biology. Nevertheless, much remains to be learned, as a complex interplay of processes with dependencies on a range of genomic features creates highly heterogeneous cancer genomes. Accurate driver detection relies on unbiased models of the mutation rate that also capture rate variation from uncharacterized sources. Results: Here, we analyse patterns of observed-to-expected mutation counts across 505 whole cancer genomes, and find that genomic features missing from our mutation-rate model likely operate on a megabase length scale. We extend our site-specific model of the mutation rate to include the additional variance from these sources, which leads to robust significance evaluation of candidate cancer drivers. We thus present ncdDetect v.2, with greatly improved cancer driver detection specificity. Finally, we show that ranking candidates by their posterior mean value of their effect sizes offers an equivalent and more computationally efficient alternative to ranking by their P-values. Availability and implementation: ncdDetect v.2 is implemented as an R-package and is freely available at http://github.com/TobiasMadsen/ncdDetect2. Supplementary information: Supplementary data are available at Bioinformatics online. Malene Juul, Tobias Madsen, Qianyun Guo, Johanna Bertl, Asger Hobolth, Manolis Kellis, Jakob Skou Pedersen |
Bioinform. | 6 |
| 2018 | RANGER-DTL 2.0: rigorous reconstruction of gene-family evolution by duplication, transfer and lossabstractSummary: RANGER-DTL 2.0 is a software program for inferring gene family evolution using Duplication-Transfer-Loss reconciliation. This new software is highly scalable and easy to use, and offers many new features not currently available in any other reconciliation program. RANGER-DTL 2.0 has a particular focus on reconciliation accuracy and can account for many sources of reconciliation uncertainty including uncertain gene tree rooting, gene tree topological uncertainty, multiple optimal reconciliations and alternative event cost assignments. RANGER-DTL 2.0 is open-source and written in C++ and Python. Availability and implementation: Pre-compiled executables, source code (open-source under GNU GPL) and a detailed manual are freely available from http://compbio.engr.uconn.edu/software/RANGER-DTL/. Supplementary information: Supplementary data are available at Bioinformatics online. Mukul S. Bansal, Manolis Kellis, Misagh Kordi, Soumya Kundu |
Bioinform. | 2 |
| 2016 | SwiSpot: modeling riboswitches by spotting out switching sequencesabstractMOTIVATION: Riboswitches are cis-regulatory elements in mRNA, mostly found in Bacteria, which exhibit two main secondary structure conformations. Although one of them prevents the gene from being expressed, the other conformation allows its expression, and this switching process is typically driven by the presence of a specific ligand. Although there are a handful of known riboswitches, our knowledge in this field has been greatly limited due to our inability to identify their alternate structures from their sequences. Indeed, current methods are not able to predict the presence of the two functionally distinct conformations just from the knowledge of the plain RNA nucleotide sequence. Whether this would be possible, for which cases, and what prediction accuracy can be achieved, are currently open questions. RESULTS: Here we show that the two alternate secondary structures of riboswitches can be accurately predicted once the 'switching sequence' of the riboswitch has been properly identified. The proposed SwiSpot approach is capable of identifying the switching sequence inside a putative, complete riboswitch sequence, on the basis of pairing behaviors, which are evaluated on proper sets of configurations. Moreover, it is able to model the switching behavior of riboswitches whose generated ensemble covers both alternate configurations. Beyond structural predictions, the approach can also be paired to homology-based riboswitch searches. AVAILABILITY AND IMPLEMENTATION: SwiSpot software, along with the reference dataset files, is available at: http://www.iet.unipi.it/a.bechini/swispot/Supplementary information: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Marco Barsacchi, Eva Maria Novoa, Manolis Kellis, Alessio Bechini |
Bioinform. | 3 |
| 2015 | Improved gene tree error correction in the presence of horizontal gene transferabstractMOTIVATION: The accurate inference of gene trees is a necessary step in many evolutionary studies. Although the problem of accurate gene tree inference has received considerable attention, most existing methods are only applicable to gene families unaffected by horizontal gene transfer. As a result, the accurate inference of gene trees affected by horizontal gene transfer remains a largely unaddressed problem. RESULTS: In this study, we introduce a new and highly effective method for gene tree error correction in the presence of horizontal gene transfer. Our method efficiently models horizontal gene transfers, gene duplications and losses, and uses a statistical hypothesis testing framework [Shimodaira-Hasegawa (SH) test] to balance sequence likelihood with topological information from a known species tree. Using a thorough simulation study, we show that existing phylogenetic methods yield inaccurate gene trees when applied to horizontally transferred gene families and that our method dramatically improves gene tree accuracy. We apply our method to a dataset of 11 cyanobacterial species and demonstrate the large impact of gene tree accuracy on downstream evolutionary analyses. AVAILABILITY AND IMPLEMENTATION: An implementation of our method is available at http://compbio.mit.edu/treefix-dtl/ CONTACT: : [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mukul S. Bansal, Yi-Chieh Wu, Eric J. Alm, Manolis Kellis |
Bioinform. | 4 |
| 2014 | Pareto-optimal phylogenetic tree reconciliationabstractMOTIVATION: Phylogenetic tree reconciliation is a widely used method for reconstructing the evolutionary histories of gene families and species, hosts and parasites and other dependent pairs of entities. Reconciliation is typically performed using maximum parsimony, in which each evolutionary event type is assigned a cost and the objective is to find a reconciliation of minimum total cost. It is generally understood that reconciliations are sensitive to event costs, but little is understood about the relationship between event costs and solutions. Moreover, choosing appropriate event costs is a notoriously difficult problem. RESULTS: We address this problem by giving an efficient algorithm for computing Pareto-optimal sets of reconciliations, thus providing the first systematic method for understanding the relationship between event costs and reconciliations. This, in turn, results in new techniques for computing event support values and, for cophylogenetic analyses, performing robust statistical tests. We provide new software tools and demonstrate their use on a number of datasets from evolutionary genomic and cophylogenetic studies. AVAILABILITY AND IMPLEMENTATION: Our Python tools are freely available at www.cs.hmc.edu/∼hadas/xscape. . Ran Libeskind-Hadas, Yi-Chieh Wu, Mukul S. Bansal, Manolis Kellis |
Bioinform. | 4 |
| 2013 | Reconciliation Revisited: Handling Multiple Optima When Reconciling with Duplication, Transfer, and Loss
Mukul S. Bansal, Eric J. Alm, Manolis Kellis |
RECOMB | 3 |
| 2013 | RFECS: A Random-Forest Based Algorithm for Enhancer Identification from Chromatin StateabstractTranscriptional enhancers play critical roles in regulation of gene expression, but their identification in the eukaryotic genome has been challenging. Recently, it was shown that enhancers in the mammalian genome are associated with characteristic histone modification patterns, which have been increasingly exploited for enhancer identification. However, only a limited number of cell types or chromatin marks have previously been investigated for this purpose, leaving the question unanswered whether there exists an optimal set of histone modifications for enhancer prediction in different cell types. Here, we address this issue by exploring genome-wide profiles of 24 histone modifications in two distinct human cell types, embryonic stem cells and lung fibroblasts. We developed a Random-Forest based algorithm, RFECS (Random Forest based Enhancer identification from Chromatin States) to integrate histone modification profiles for identification of enhancers, and used it to identify enhancers in a number of cell-types. We show that RFECS not only leads to more accurate and precise prediction of enhancers than previous methods, but also helps identify the most informative and robust set of three chromatin marks for enhancer prediction. Nisha Rajagopal, Uli Wagner 0002, Wei Wang 0051, John A. Stamatoyannopoulos, Jason Ernst, Manolis Kellis |
PLoS Comput. Biol. | 8 |
| 2012 | Efficient algorithms for the reconciliation problem with gene duplication, horizontal transfer and lossabstractMOTIVATION: Gene family evolution is driven by evolutionary events such as speciation, gene duplication, horizontal gene transfer and gene loss, and inferring these events in the evolutionary history of a given gene family is a fundamental problem in comparative and evolutionary genomics with numerous important applications. Solving this problem requires the use of a reconciliation framework, where the input consists of a gene family phylogeny and the corresponding species phylogeny, and the goal is to reconcile the two by postulating speciation, gene duplication, horizontal gene transfer and gene loss events. This reconciliation problem is referred to as duplication-transfer-loss (DTL) reconciliation and has been extensively studied in the literature. Yet, even the fastest existing algorithms for DTL reconciliation are too slow for reconciling large gene families and for use in more sophisticated applications such as gene tree or species tree reconstruction. RESULTS: We present two new algorithms for the DTL reconciliation problem that are dramatically faster than existing algorithms, both asymptotically and in practice. We also extend the standard DTL reconciliation model by considering distance-dependent transfer costs, which allow for more accurate reconciliation and give an efficient algorithm for DTL reconciliation under this extended model. We implemented our new algorithms and demonstrated up to 100 000-fold speed-up over existing methods, using both simulated and biological datasets. This dramatic improvement makes it possible to use DTL reconciliation for performing rigorous evolutionary analyses of large gene families and enables its use in advanced reconciliation-based gene and species tree reconstruction methods. AVAILABILITY: Our programs can be freely downloaded from http://compbio.mit.edu/ranger-dtl/. Mukul S. Bansal, Eric J. Alm, Manolis Kellis |
Bioinform. | 3 |
| 2011 | Discovery and Characterization of Chromatin States for Systematic Annotation of the Human Genome
Jason Ernst, Manolis Kellis |
RECOMB | 2 |
| 2011 | PhyloCSF: a comparative genomics method to distinguish protein coding and non-coding regionsabstractMOTIVATION: As high-throughput transcriptome sequencing provides evidence for novel transcripts in many species, there is a renewed need for accurate methods to classify small genomic regions as protein coding or non-coding. We present PhyloCSF, a novel comparative genomics method that analyzes a multispecies nucleotide sequence alignment to determine whether it is likely to represent a conserved protein-coding region, based on a formal statistical comparison of phylogenetic codon models. RESULTS: We show that PhyloCSF's classification performance in 12-species Drosophila genome alignments exceeds all other methods we compared in a previous study. We anticipate that this method will be widely applicable as the transcriptomes of many additional species, tissues and subcellular compartments are sequenced, particularly in the context of ENCODE and modENCODE, and as interest grows in long non-coding RNAs, often initially recognized by their lack of protein coding potential rather than conserved RNA secondary structures. AVAILABILITY AND IMPLEMENTATION: The Objective Caml source code and executables for GNU/Linux and Mac OS X are freely available at http://compbio.mit.edu/PhyloCSF CONTACT: [email protected]; [email protected]. Michael F. Lin, Irwin Jungreis, Manolis Kellis |
Bioinform. | 3 |
| 2010 | Motif discovery in physiological datasets: A methodology for inferring predictive elementsabstractIn this article, we propose a methodology for identifying predictive physiological patterns in the absence of prior knowledge. We use the principle of conservation to identify activity that consistently precedes an outcome in patients, and describe a two-stage process that allows us to efficiently search for such patterns in large datasets. This involves first transforming continuous physiological signals from patients into symbolic sequences, and then searching for patterns in these reduced representations that are strongly associated with an outcome.Our strategy of identifying conserved activity that is unlikely to have occurred purely by chance in symbolic data is analogous to the discovery of regulatory motifs in genomic datasets. We build upon existing work in this area, generalizing the notion of a regulatory motif and enhancing current techniques to operate robustly on non-genomic data. We also address two significant considerations associated with motif discovery in general: computational efficiency and robustness in the presence of degeneracy and noise. To deal with these issues, we introduce the concept of active regions and new subset-based techniques such as a two-layer Gibbs sampling algorithm. These extensions allow for a framework for information inference, where precursors are identified as approximately conserved activity of arbitrary complexity preceding multiple occurrences of an event.We evaluated our solution on a population of patients who experienced sudden cardiac death and attempted to discover electrocardiographic activity that may be associated with the endpoint of death. To assess the predictive patterns discovered, we compared likelihood scores for motifs in the sudden death population against control populations of normal individuals and those with non-fatal supraventricular arrhythmias. Our results suggest that predictive motif discovery may be able to identify clinically relevant information even in the absence of significant prior knowledge. Zeeshan Syed, Collin M. Stultz, Manolis Kellis, Piotr Indyk, John V. Guttag |
ACM Trans. Knowl. Discov. Data | 3 |
| 2008 | Performance and Scalability of Discriminative Metrics for Comparative Gene Identification in 12 Drosophila GenomesabstractComparative genomics of multiple related species is a powerful methodology for the discovery of functional genomic elements, and its power should increase with the number of species compared. Here, we use 12 Drosophila genomes to study the power of comparative genomics metrics to distinguish between protein-coding and non-coding regions. First, we study the relative power of different comparative metrics and their relationship to single-species metrics. We find that even relatively simple multi-species metrics robustly outperform advanced single-species metrics, especially for shorter exons (< or =240 nt), which are common in animal genomes. Moreover, the two capture largely independent features of protein-coding genes, with different sensitivity/specificity trade-offs, such that their combinations lead to even greater discriminatory power. In addition, we study how discovery power scales with the number and phylogenetic distance of the genomes compared. We find that species at a broad range of distances are comparably effective informants for pairwise comparative gene identification, but that these are surpassed by multi-species comparisons at similar evolutionary divergence. In particular, while pairwise discovery power plateaued at larger distances and never outperformed the most advanced single-species metrics, multi-species comparisons continued to benefit even from the most distant species with no apparent saturation. Last, we find that genes in functional categories typically considered fast-evolving can nonetheless be recovered at very high rates using comparative methods. Our results have implications for comparative genomics analyses in any species, including the human. Michael F. Lin, Ameya N. Deoras, Matthew D. Rasmussen, Manolis Kellis |
PLoS Comput. Biol. | 4 |
| 2007 | Network Motif Discovery Using Subgraph Enumeration and Symmetry-Breaking
Joshua A. Grochow, Manolis Kellis |
RECOMB | 2 |