EDBT 2026 Demo / reviewers in the wild / expert
Eleazar Eskin
dblp:e/EEskin
· DBLP profile ↗
69ranked-venue papers
8as first author
2since 2021 · last 2024
0000-0003-1149-4758ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 55 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 1 first-authorSecurity and privacy · 4Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
42 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
5 papers |
Probabilistic and Bayesian machine learning · 67% Knowledge representation and reasoning · 24% Information extraction and text analysis · 5% |
Topics — the 30 heaviest of 66, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
genomics |
2.6 | 12 | 2020 | Identifying Causal Variants by Fine Mapping Across Multiple Studies · RECOMB 2020 A Unifying Framework for Summary Statistic Imputation · RECOMB 2018 Applying meta-analysis to genotype-tissue expression data from multiple tissues to identify eQTLs and increase the number of eGenes · Bioinform. 2017 |
Bioinformatics and computational biology › genomics
genome-wide association study |
1.5 | 7 | 2020 | Finding associated variants in genome-wide association studies on multiple traits · Bioinform. 2018 Improved methods for multi-trait fine mapping of pleiotropic risk loci · Bioinform. 2017 Identification of causal genes for complex traits · Bioinform. 2015 |
Bioinformatics and computational biology
statistical genetics |
1.1 | 5 | 2017 | Increasing the power of meta-analysis of genome-wide association studies to detect heterogeneous effects · Bioinform. 2017 Improved methods for multi-trait fine mapping of pleiotropic risk loci · Bioinform. 2017 Identification of causal genes for complex traits · Bioinform. 2015 |
Bioinformatics and computational biology › statistical genetics › fine-mapping
causal variant identification |
0.9 | 3 | 2020 | Identifying Causal Variants by Fine Mapping Across Multiple Studies · RECOMB 2020 Improving Imputation Accuracy by Inferring Causal Variants in Genetic Studies · RECOMB 2017 Identification of causal genes for complex traits · Bioinform. 2015 |
Bioinformatics and computational biology › statistical genetics
fine-mapping |
0.9 | 3 | 2020 | Identifying Causal Variants by Fine Mapping Across Multiple Studies · RECOMB 2020 Improved methods for multi-trait fine mapping of pleiotropic risk loci · Bioinform. 2017 Identification of causal genes for complex traits · Bioinform. 2015 |
Bioinformatics and computational biology › genomics › genotyping
genotype imputation |
0.9 | 4 | 2018 | A Unifying Framework for Summary Statistic Imputation · RECOMB 2018 Improving Imputation Accuracy by Inferring Causal Variants in Genetic Studies · RECOMB 2017 A Spatial-Aware Haplotype Copying Model with Applications to Genotype Imputation · RECOMB 2014 |
Bioinformatics and computational biology › functional genomics
eQTL mapping |
0.5 | 2 | 2017 | Applying meta-analysis to genotype-tissue expression data from multiple tissues to identify eQTLs and increase the number of eGenes · Bioinform. 2017 Using genomic annotations increases statistical power to detect eGenes · Bioinform. 2016 |
Bioinformatics and computational biology › statistical genetics
genetic association study |
0.5 | 3 | 2017 | Improving Imputation Accuracy by Inferring Causal Variants in Genetic Studies · RECOMB 2017 Increasing Power of Groupwise Association Test with Likelihood Ratio Test · RECOMB 2011 Increasing Power in Association Studies by Using Linkage Disequilibrium Structure and Molecular Function as Prior Information · RECOMB 2008 |
Bioinformatics and computational biology › genomics
haplotype inference |
0.4 | 4 | 2013 | Leveraging reads that span multiple single nucleotide polymorphisms for haplotype inference from sequencing data · Bioinform. 2013 Optimal algorithms for haplotype assembly from whole-genome sequence data · Bioinform. 2010 Haplotype reconstruction from genotype data using Imperfect Phylogeny · Bioinform. 2004 |
Bioinformatics and computational biology › genomics › genome-wide association study
summary statistic imputation |
0.3 | 1 | 2018 | A Unifying Framework for Summary Statistic Imputation · RECOMB 2018 |
Bioinformatics and computational biology › single-cell analysis
cell type composition estimation |
0.3 | 1 | 2017 | A Bayesian Framework for Estimating Cell Type Composition from DNA Methylation Without the Need for Methylation Reference · RECOMB 2017 |
Bioinformatics and computational biology
epigenetics |
0.3 | 1 | 2017 | A Bayesian Framework for Estimating Cell Type Composition from DNA Methylation Without the Need for Methylation Reference · RECOMB 2017 |
Bioinformatics and computational biology › statistical genetics › genetic association study
genome-wide association study meta-analysis |
0.3 | 1 | 2017 | Increasing the power of meta-analysis of genome-wide association studies to detect heterogeneous effects · Bioinform. 2017 |
Bioinformatics and computational biology
sequence analysis |
0.3 | 2 | 2016 | Long Single-Molecule Reads Can Resolve the Complexity of the Influenza Virus Composed of Rare, Closely Related Mutant Variants · RECOMB 2016 Finding composite regulatory patterns in DNA sequences · ISMB 2002 |
Bioinformatics and computational biology › genomics › structural variation
copy number variation |
0.3 | 2 | 2012 | CNVeM: Copy Number Variation Detection Using Uncertainty of Read Mapping · RECOMB 2012 Efficient algorithms for tandem copy number variation reconstruction in repeat-rich regions · Bioinform. 2011 |
Bioinformatics and computational biology › genomics
viral genomics |
0.2 | 1 | 2016 | Long Single-Molecule Reads Can Resolve the Complexity of the Influenza Virus Composed of Rare, Closely Related Mutant Variants · RECOMB 2016 |
Bioinformatics and computational biology
population genetics |
0.2 | 2 | 2014 | A Spatial-Aware Haplotype Copying Model with Applications to Genotype Imputation · RECOMB 2014 Haplotype reconstruction from genotype data using Imperfect Phylogeny · Bioinform. 2004 |
Bioinformatics and computational biology › population genetics
population stratification correction |
0.2 | 1 | 2015 | Efficient and Accurate Multiple-Phenotypes Regression Method for High Dimensional Data Considering Population Structure · RECOMB 2015 |
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics |
0.2 | 1 | 2015 | Efficient and Accurate Multiple-Phenotypes Regression Method for High Dimensional Data Considering Population Structure · RECOMB 2015 |
Bioinformatics and computational biology
gene expression analysis |
0.2 | 3 | 2011 | Detecting the Presence and Absence of Causal Relationships between Expression of Yeast Genes with Very Few Samples · RECOMB 2009 Discovering tightly regulated and differentially expressed gene sets in whole genome expression data · Bioinform. 2007 Mixed-model coexpression: calculating gene coexpression while accounting for expression heterogeneity · Bioinform. 2011 |
Bioinformatics and computational biology › statistical genetics › gene-gene interaction
gene-gene interaction detection |
0.2 | 1 | 2014 | Gene-Gene Interactions Detection Using a Two-Stage Model · RECOMB 2014 |
Bioinformatics and computational biology › statistical genetics › rare variant analysis
rare variant detection |
0.2 | 1 | 2014 | Accurate viral population assembly from ultra-deep sequencing data · Bioinform. 2014 |
Bioinformatics and computational biology › sequence analysis
sequencing error correction |
0.2 | 1 | 2014 | Accurate viral population assembly from ultra-deep sequencing data · Bioinform. 2014 |
Privacy and data protection › health data privacy
privacy-preserving genomic analysis |
0.2 | 1 | 2014 | Privacy preserving protocol for detecting genetic relatives using rare variants · Bioinform. 2014 |
Cryptographic protocols and secure computation
secure multiparty computation |
0.2 | 1 | 2014 | Privacy preserving protocol for detecting genetic relatives using rare variants · Bioinform. 2014 |
Bioinformatics and computational biology › statistical genetics › haplotype analysis
haplotype phasing |
0.2 | 2 | 2012 | Hap-seq: An Optimal Algorithm for Haplotype Phasing with Imputation Using Sequencing Data · RECOMB 2012 Optimal algorithms for haplotype assembly from whole-genome sequence data · Bioinform. 2010 |
Bioinformatics and computational biology › statistical genetics › genotype data analysis
genotype data integration |
0.2 | 1 | 2013 | eALPS: Estimating Abundance Levels in Pooled Sequencing Using Available Genotyping Data · RECOMB 2013 |
Bioinformatics and computational biology › population genetics
pedigree reconstruction |
0.2 | 1 | 2013 | IPED: Inheritance Path Based Pedigree Reconstruction Algorithm Using Genotype Data · RECOMB 2013 |
Bioinformatics and computational biology › genomics › high-throughput sequencing
pooled sequencing |
0.2 | 1 | 2013 | eALPS: Estimating Abundance Levels in Pooled Sequencing Using Available Genotyping Data · RECOMB 2013 |
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number variation detection |
0.1 | 1 | 2012 | CNVeM: Copy Number Variation Detection Using Uncertainty of Read Mapping · RECOMB 2012 |
Methods — techniques the papers use, named apart from their topics
meta-analysis · 1.1statistical association testing · 0.6importance sampling · 0.5bayesian fine mapping · 0.4statistical genetics · 0.4mixed model · 0.3statistical modeling · 0.3statistical imputation · 0.3bayesian model inference · 0.3bayesian inference · 0.3secure computation · 0.2rare variant analysis · 0.2genotype data analysis · 0.2sequencing data analysis · 0.1dynamic programming · 0.1probability distribution learning · 0.1machine learning · 0.0linguistic features · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | VISTA: an integrated framework for structural variant discoveryabstractStructural variation (SV) refers to insertions, deletions, inversions, and duplications in human genomes. SVs are present in approximately 1.5% of the human genome. Still, this small subset of genetic variation has been implicated in the pathogenesis of psoriasis, Crohn's disease and other autoimmune disorders, autism spectrum and other neurodevelopmental disorders, and schizophrenia. Since identifying structural variants is an important problem in genetics, several specialized computational techniques have been developed to detect structural variants directly from sequencing data. With advances in whole-genome sequencing (WGS) technologies, a plethora of SV detection methods have been developed. However, dissecting SVs from WGS data remains a challenge, with the majority of SV detection methods prone to a high false-positive rate, and no existing method able to precisely detect a full range of SVs present in a sample. Previous studies have shown that none of the existing SV callers can maintain high accuracy across various SV lengths and genomic coverages. Here, we report an integrated structural variant calling framework, Variant Identification and Structural Variant Analysis (VISTA), that leverages the results of individual callers using a novel and robust filtering and merging algorithm. In contrast to existing consensus-based tools which ignore the length and coverage, VISTA overcomes this limitation by executing various combinations of top-performing callers based on variant length and genomic coverage to generate SV events with high accuracy. We evaluated the performance of VISTA on comprehensive gold-standard datasets across varying organisms and coverage. We benchmarked VISTA using the Genome-in-a-Bottle gold standard SV set, haplotype-resolved de novo assemblies from the Human Pangenome Reference Consortium, along with an in-house polymerase chain reaction (PCR)-validated mouse gold standard set. VISTA maintained the highest F1 score among top consensus-based tools measured using a comprehensive gold standard across both mouse and human genomes. VISTA also has an optimized mode, where the calls can be optimized for precision or recall. VISTA-optimized can attain 100% precision and the highest sensitivity among other variant callers. In conclusion, VISTA represents a significant advancement in structural variant calling, offering a robust and accurate framework that outperforms existing consensus-based tools and sets a new standard for SV detection in genomic research. Varuni Sarwal, Seungmo Lee, Jianzhi Yang, Sriram Sankararaman, Mark Chaisson, Eleazar Eskin, Serghei Mangul |
Briefings Bioinform. | 6 |
| 2022 | A comprehensive benchmarking of WGS-based deletion structural variant callersabstractAdvances in whole-genome sequencing (WGS) promise to enable the accurate and comprehensive structural variant (SV) discovery. Dissecting SVs from WGS data presents a substantial number of challenges and a plethora of SV detection methods have been developed. Currently, evidence that investigators can use to select appropriate SV detection tools is lacking. In this article, we have evaluated the performance of SV detection tools on mouse and human WGS data using a comprehensive polymerase chain reaction-confirmed gold standard set of SVs and the genome-in-a-bottle variant set, respectively. In contrast to the previous benchmarking studies, our gold standard dataset included a complete set of SVs allowing us to report both precision and sensitivity rates of the SV detection methods. Our study investigates the ability of the methods to detect deletions, thus providing an optimistic estimate of SV detection performance as the SV detection methods that fail to detect deletions are likely to miss more complex SVs. We found that SV detection tools varied widely in their performance, with several methods providing a good balance between sensitivity and precision. Additionally, we have determined the SV callers best suited for low- and ultralow-pass sequencing data as well as for different deletion length categories. Varuni Sarwal, Sebastian Niehus, Ram Ayyala, Aditya Sarkar, Sei Chang, Angela Lu, Neha Rajkumar, Nicholas Darci-Maher, Russell Littman, Karishma Chhugani, Arda Söylev, Zoia Comarova, Emily E. Wesel, Jacqueline Castellanos, Rahul Chikka, Margaret G. Distler, Eleazar Eskin, Jonathan Flint, Serghei Mangul |
Briefings Bioinform. | 18 |
| 2020 | Identifying Causal Variants by Fine Mapping Across Multiple Studies
Nathan Lapierre, Kodi Taraszka, Helen Huang, Rosemary He, Farhad Hormozdiari, Eleazar Eskin |
RECOMB | 6 |
| 2018 | A Unifying Framework for Summary Statistic Imputation
Eleazar Eskin, Sriram Sankararaman |
RECOMB | 2 |
| 2018 | Finding associated variants in genome-wide association studies on multiple traitsabstractMotivation: Many variants identified by genome-wide association studies (GWAS) have been found to affect multiple traits, either directly or through shared pathways. There is currently a wealth of GWAS data collected in numerous phenotypes, and analyzing multiple traits at once can increase power to detect shared variant effects. However, traditional meta-analysis methods are not suitable for combining studies on different traits. When applied to dissimilar studies, these meta-analysis methods can be underpowered compared to univariate analysis. The degree to which traits share variant effects is often not known, and the vast majority of GWAS meta-analysis only consider one trait at a time. Results: Here, we present a flexible method for finding associated variants from GWAS summary statistics for multiple traits. Our method estimates the degree of shared effects between traits from the data. Using simulations, we show that our method properly controls the false positive rate and increases power when an effect is present in a subset of traits. We then apply our method to the North Finland Birth Cohort and UK Biobank datasets using a variety of metabolic traits and discover novel loci. Availability and implementation: Our source code is available at https://github.com/lgai/CONFIT. Supplementary information: Supplementary data are available at Bioinformatics online. Lisa Gai, Eleazar Eskin |
Bioinform. | 2 |
| 2017 | A Bayesian Framework for Estimating Cell Type Composition from DNA Methylation Without the Need for Methylation Reference
Elior Rahmani, Regev Schweiger, Liat Shenhav, Eleazar Eskin, Eran Halperin |
RECOMB | 4 |
| 2017 | Improving Imputation Accuracy by Inferring Causal Variants in Genetic Studies
Farhad Hormozdiari, Jong Wha J. Joo, Eleazar Eskin |
RECOMB | 4 |
| 2017 | Applying meta-analysis to genotype-tissue expression data from multiple tissues to identify eQTLs and increase the number of eGenesabstractMOTIVATION: There is recent interest in using gene expression data to contextualize findings from traditional genome-wide association studies (GWAS). Conditioned on a tissue, expression quantitative trait loci (eQTLs) are genetic variants associated with gene expression, and eGenes are genes whose expression levels are associated with genetic variants. eQTLs and eGenes provide great supporting evidence for GWAS hits and important insights into the regulatory pathways involved in many diseases. When a significant variant or a candidate gene identified by GWAS is also an eQTL or eGene, there is strong evidence to further study this variant or gene. Multi-tissue gene expression datasets like the Gene Tissue Expression (GTEx) data are used to find eQTLs and eGenes. Unfortunately, these datasets often have small sample sizes in some tissues. For this reason, there have been many meta-analysis methods designed to combine gene expression data across many tissues to increase power for finding eQTLs and eGenes. However, these existing techniques are not scalable to datasets containing many tissues, like the GTEx data. Furthermore, these methods ignore a biological insight that the same variant may be associated with the same gene across similar tissues. RESULTS: We introduce a meta-analysis model that addresses these problems in existing methods. We focus on the problem of finding eGenes in gene expression data from many tissues, and show that our model is better than other types of meta-analyses. AVAILABILITY AND IMPLEMENTATION: Source code is at https://github.com/datduong/RECOV . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dat Duong, Lisa Gai, Sagi Snir, Eun Yong Kang, Buhm Han, Jae Hoon Sul, Eleazar Eskin |
Bioinform. | 7 |
| 2017 | Improved methods for multi-trait fine mapping of pleiotropic risk lociabstractMOTIVATION: Genome-wide association studies (GWAS) have identified thousands of regions in the genome that contain genetic variants that increase risk for complex traits and diseases. However, the variants uncovered in GWAS are typically not biologically causal, but rather, correlated to the true causal variant through linkage disequilibrium (LD). To discern the true causal variant(s), a variety of statistical fine-mapping methods have been proposed to prioritize variants for functional validation. RESULTS: In this work we introduce a new approach, fastPAINTOR, that leverages evidence across correlated traits, as well as functional annotation data, to improve fine-mapping accuracy at pleiotropic risk loci. To improve computational efficiency, we describe an new importance sampling scheme to perform model inference. First, we demonstrate in simulations that by leveraging functional annotation data, fastPAINTOR increases fine-mapping resolution relative to existing methods. Next, we show that jointly modeling pleiotropic risk regions improves fine-mapping resolution compared to standard single trait and pleiotropic fine mapping strategies. We report a reduction in the number of SNPs required for follow-up in order to capture 90% of the causal variants from 23 SNPs per locus using a single trait to 12 SNPs when fine-mapping two traits simultaneously. Finally, we analyze summary association data from a large-scale GWAS of lipids and show that these improvements are largely sustained in real data. AVAILABILITY AND IMPLEMENTATION: The fastPAINTOR framework is implemented in the PAINTOR v3.0 package which is publicly available to the research community http://bogdan.bioinformatics.ucla.edu/software/paintor CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online. Gleb Kichaev, Megan Roytman, Ruth Johnson, Eleazar Eskin, Sara Lindström, Peter Kraft, Bogdan Pasaniuc |
Bioinform. | 4 |
| 2017 | Increasing the power of meta-analysis of genome-wide association studies to detect heterogeneous effectsabstractMOTIVATION: Meta-analysis is essential to combine the results of genome-wide association studies (GWASs). Recent large-scale meta-analyses have combined studies of different ethnicities, environments and even studies of different related phenotypes. These differences between studies can manifest as effect size heterogeneity. We previously developed a modified random effects model (RE2) that can achieve higher power to detect heterogeneous effects than the commonly used fixed effects model (FE). However, RE2 cannot perform meta-analysis of correlated statistics, which are found in recent research designs, and the identified variants often overlap with those found by FE. RESULTS: Here, we propose RE2C, which increases the power of RE2 in two ways. First, we generalized the likelihood model to account for correlations of statistics to achieve optimal power, using an optimization technique based on spectral decomposition for efficient parameter estimation. Second, we designed a novel statistic to focus on the heterogeneous effects that FE cannot detect, thereby, increasing the power to identify new associations. We developed an efficient and accurate p -value approximation procedure using analytical decomposition of the statistic. In simulations, RE2C achieved a dramatic increase in power compared with the decoupling approach (71% vs. 21%) when the statistics were correlated. Even when the statistics are uncorrelated, RE2C achieves a modest increase in power. Applications to real genetic data supported the utility of RE2C. RE2C is highly efficient and can meta-analyze one hundred GWASs in one day. AVAILABILITY AND IMPLEMENTATION: The software is freely available at http://software.buhmhan.com/RE2C . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Eleazar Eskin, Buhm Han |
Bioinform. | 2 |
| 2017 | IPED2: Inheritance Path Based Pedigree Reconstruction Algorithm for Complicated PedigreesabstractReconstruction of family trees, or pedigree reconstruction, for a group of individuals is a fundamental problem in genetics. The problem is known to be NP-hard even for datasets known to only contain siblings. Some recent methods have been developed to accurately and efficiently reconstruct pedigrees. These methods, however, still consider relatively simple pedigrees, for example, they are not able to handle half-sibling situations where a pair of individuals only share one parent. In this work, we propose an efficient method, IPED2, based on our previous work, which specifically targets reconstruction of complicated pedigrees that include half-siblings. We note that the presence of half-siblings makes the reconstruction problem significantly more challenging which is why previous methods exclude the possibility of half-siblings. We proposed a novel model as well as an efficient graph algorithm and experiments show that our algorithm achieves relatively accurate reconstruction. To our knowledge, this is the first method that is able to handle pedigree reconstruction from genotype data when half-sibling exists in any generation of the pedigree. Dan He 0001, Zhanyong Wang, Laxmi Parida, Eleazar Eskin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2016 | HapIso: An Accurate Method for the Haplotype-Specific Isoforms Reconstruction from Long Single-Molecule Reads
Serghei Mangul, Harry (Taegyun) Yang, Farhad Hormozdiari, Elizabeth Tseng, Alex Zelikovsky, Eleazar Eskin |
ISBRA | 6 |
| 2016 | Long Single-Molecule Reads Can Resolve the Complexity of the Influenza Virus Composed of Rare, Closely Related Mutant Variants
Alexander Artyomenko, Nicholas C. Wu, Serghei Mangul, Eleazar Eskin, Ren Sun, Alex Zelikovsky |
RECOMB | 4 |
| 2016 | Using genomic annotations increases statistical power to detect eGenesabstractMOTIVATION: Expression quantitative trait loci (eQTLs) are genetic variants that affect gene expression. In eQTL studies, one important task is to find eGenes or genes whose expressions are associated with at least one eQTL. The standard statistical method to determine whether a gene is an eGene requires association testing at all nearby variants and the permutation test to correct for multiple testing. The standard method however does not consider genomic annotation of the variants. In practice, variants near gene transcription start sites (TSSs) or certain histone modifications are likely to regulate gene expression. In this article, we introduce a novel eGene detection method that considers this empirical evidence and thereby increases the statistical power. RESULTS: We applied our method to the liver Genotype-Tissue Expression (GTEx) data using distance from TSSs, DNase hypersensitivity sites, and six histone modifications as the genomic annotations for the variants. Each of these annotations helped us detected more candidate eGenes. Distance from TSS appears to be the most important annotation; specifically, using this annotation, our method discovered 50% more candidate eGenes than the standard permutation method. CONTACT: [email protected] or [email protected]. Dat Duong, Jennifer Zou, Farhad Hormozdiari, Jae Hoon Sul, Jason Ernst, Buhm Han, Eleazar Eskin |
Bioinform. | 7 |
| 2015 | Efficient and Accurate Multiple-Phenotypes Regression Method for High Dimensional Data Considering Population Structure
Jong Wha J. Joo, Eun Yong Kang, Elin Org, Nicholas A. Furlotte, Brian Parks, Aldons J. Lusis, Eleazar Eskin |
RECOMB | 7 |
| 2015 | Identification of causal genes for complex traitsabstractMOTIVATION: Although genome-wide association studies (GWAS) have identified thousands of variants associated with common diseases and complex traits, only a handful of these variants are validated to be causal. We consider 'causal variants' as variants which are responsible for the association signal at a locus. As opposed to association studies that benefit from linkage disequilibrium (LD), the main challenge in identifying causal variants at associated loci lies in distinguishing among the many closely correlated variants due to LD. This is particularly important for model organisms such as inbred mice, where LD extends much further than in human populations, resulting in large stretches of the genome with significantly associated variants. Furthermore, these model organisms are highly structured and require correction for population structure to remove potential spurious associations. RESULTS: In this work, we propose CAVIAR-Gene (CAusal Variants Identification in Associated Regions), a novel method that is able to operate across large LD regions of the genome while also correcting for population structure. A key feature of our approach is that it provides as output a minimally sized set of genes that captures the genes which harbor causal variants with probability ρ. Through extensive simulations, we demonstrate that our method not only speeds up computation, but also have an average of 10% higher recall rate compared with the existing approaches. We validate our method using a real mouse high-density lipoprotein data (HDL) and show that CAVIAR-Gene is able to identify Apoa2 (a gene known to harbor causal variants for HDL), while reducing the number of genes that need to be tested for functionality by a factor of 2. AVAILABILITY AND IMPLEMENTATION: Software is freely available for download at genetics.cs.ucla.edu/caviar. Farhad Hormozdiari, Gleb Kichaev, Wen-Yun Yang, Bogdan Pasaniuc, Eleazar Eskin |
Bioinform. | 5 |
| 2014 | Gene-Gene Interactions Detection Using a Two-Stage Model
Zhanyong Wang, Jae Hoon Sul, Sagi Snir, José Antonio Lozano 0001, Eleazar Eskin |
RECOMB | 5 |
| 2014 | A Spatial-Aware Haplotype Copying Model with Applications to Genotype Imputation
Wen-Yun Yang, Farhad Hormozdiari, Eleazar Eskin, Bogdan Pasaniuc |
RECOMB | 3 |
| 2014 | Fast pairwise IBD association testing in genome-wide association studiesabstractMOTIVATION: Recently, investigators have proposed state-of-the-art Identity-by-descent (IBD) mapping methods to detect IBD segments between purportedly unrelated individuals. The IBD information can then be used for association testing in genetic association studies. One approach for this IBD association testing strategy is to test for excessive IBD between pairs of cases ('pairwise method'). However, this approach is inefficient because it requires a large number of permutations. Moreover, a limited number of permutations define a lower bound for P-values, which makes fine-mapping of associated regions difficult because, in practice, a much larger genomic region is implicated than the region that is actually associated. RESULTS: In this article, we introduce a new pairwise method 'Fast-Pairwise'. Fast-Pairwise uses importance sampling to improve efficiency and enable approximation of extremely small P-values. Fast-Pairwise method takes only days to complete a genome-wide scan. In the application to the WTCCC type 1 diabetes data, Fast-Pairwise successfully fine-maps a known human leukocyte antigen gene that is known to cause the disease. AVAILABILITY: Fast-Pairwise is publicly available at: http://genetics.cs.ucla.edu/graphibd. Buhm Han, Eun Yong Kang, Soumya Raychaudhuri, Paul I. W. de Bakker, Eleazar Eskin |
Bioinform. | 5 |
| 2014 | Privacy preserving protocol for detecting genetic relatives using rare variantsabstractMOTIVATION: High-throughput sequencing technologies have impacted many areas of genetic research. One such area is the identification of relatives from genetic data. The standard approach for the identification of genetic relatives collects the genomic data of all individuals and stores it in a database. Then, each pair of individuals is compared to detect the set of genetic relatives, and the matched individuals are informed. The main drawback of this approach is the requirement of sharing your genetic data with a trusted third party to perform the relatedness test. RESULTS: In this work, we propose a secure protocol to detect the genetic relatives from sequencing data while not exposing any information about their genomes. We assume that individuals have access to their genome sequences but do not want to share their genomes with anyone else. Unlike previous approaches, our approach uses both common and rare variants which provide the ability to detect much more distant relationships securely. We use a simulated data generated from the 1000 genomes data and illustrate that we can easily detect up to fifth degree cousins which was not possible using the existing methods. We also show in the 1000 genomes data with cryptic relationships that our method can detect these individuals. AVAILABILITY: The software is freely available for download at http://genetics.cs.ucla.edu/crypto/. Farhad Hormozdiari, Jong Wha J. Joo, Akshay Wadia, Feng Guan, Rafail Ostrovsky, Amit Sahai, Eleazar Eskin |
Bioinform. | 7 |
| 2014 | Accurate viral population assembly from ultra-deep sequencing dataabstractMOTIVATION: Next-generation sequencing technologies sequence viruses with ultra-deep coverage, thus promising to revolutionize our understanding of the underlying diversity of viral populations. While the sequencing coverage is high enough that even rare viral variants are sequenced, the presence of sequencing errors makes it difficult to distinguish between rare variants and sequencing errors. RESULTS: In this article, we present a method to overcome the limitations of sequencing technologies and assemble a diverse viral population that allows for the detection of previously undiscovered rare variants. The proposed method consists of a high-fidelity sequencing protocol and an accurate viral population assembly method, referred to as Viral Genome Assembler (VGA). The proposed protocol is able to eliminate sequencing errors by using individual barcodes attached to the sequencing fragments. Highly accurate data in combination with deep coverage allow VGA to assemble rare variants. VGA uses an expectation-maximization algorithm to estimate abundances of the assembled viral variants in the population. RESULTS on both synthetic and real datasets show that our method is able to accurately assemble an HIV viral population and detect rare variants previously undetectable due to sequencing errors. VGA outperforms state-of-the-art methods for genome-wide viral assembly. Furthermore, our method is the first viral assembly method that scales to millions of sequencing reads. AVAILABILITY: Our tool VGA is freely available at http://genetics.cs.ucla.edu/vga/ Serghei Mangul, Nicholas C. Wu, Nicholas Mancuso, Alex Zelikovsky, Ren Sun, Eleazar Eskin |
Bioinform. | 6 |
| 2013 | IPEDX: An exact algorithm for pedigree reconstruction using genotype dataabstractThe problem of inference of family trees, or pedigree reconstruction, for a group of individuals has attracted lots of attentions recently. Various methods have been proposed to automate the process of pedigree reconstruction given the genotypes or haplotypes of a set of individuals. The state-of-the-art method IPED is able to reconstruct large pedigrees with reasonable accuracy. However, the algorithm is shown to be an approximate algorithm. In this work, we proposed an exact method IPEDX, where two dynamic programming algorithms are developed to compute inheritance paths between ancestors and descendants as well as exact paths between extant individuals, respectively. Then IPEDX reconstructs the pedigrees utilizing the outputs of the two algorithms. Experiments show that as an exact algorithm, IPEDX generally achieves better results than IPED does. It does require longer computation time but is still very efficient for pedigrees which have a large number of generations. Dan He 0001, Eleazar Eskin |
BIBM | 2 |
| 2013 | eALPS: Estimating Abundance Levels in Pooled Sequencing Using Available Genotyping Data
Itamar Eskin, Farhad Hormozdiari, Lucía Conde, Jacques Riby, Chris Skibola, Eleazar Eskin, Eran Halperin |
RECOMB | 6 |
| 2013 | IPED: Inheritance Path Based Pedigree Reconstruction Algorithm Using Genotype Data
Dan He 0001, Zhanyong Wang, Buhm Han, Laxmi Parida, Eleazar Eskin |
RECOMB | 5 |
| 2013 | Efficiently Identifying Significant Associations in Genome-Wide Association Studies
Emrah Kostem, Eleazar Eskin |
RECOMB | 2 |
| 2013 | Leveraging reads that span multiple single nucleotide polymorphisms for haplotype inference from sequencing dataabstractMOTIVATION: Haplotypes, defined as the sequence of alleles on one chromosome, are crucial for many genetic analyses. As experimental determination of haplotypes is extremely expensive, haplotypes are traditionally inferred using computational approaches from genotype data, i.e. the mixture of the genetic information from both haplotypes. Best performing approaches for haplotype inference rely on Hidden Markov Models, with the underlying assumption that the haplotypes of a given individual can be represented as a mosaic of segments from other haplotypes in the same population. Such algorithms use this model to predict the most likely haplotypes that explain the observed genotype data conditional on reference panel of haplotypes. With rapid advances in short read sequencing technologies, sequencing is quickly establishing as a powerful approach for collecting genetic variation information. As opposed to traditional genotyping-array technologies that independently call genotypes at polymorphic sites, short read sequencing often collects haplotypic information; a read spanning more than one polymorphic locus (multi-single nucleotide polymorphic read) contains information on the haplotype from which the read originates. However, this information is generally ignored in existing approaches for haplotype phasing and genotype-calling from short read data. RESULTS: In this article, we propose a novel framework for haplotype inference from short read sequencing that leverages multi-single nucleotide polymorphic reads together with a reference panel of haplotypes. The basis of our approach is a new probabilistic model that finds the most likely haplotype segments from the reference panel to explain the short read sequencing data for a given individual. We devised an efficient sampling method within a probabilistic model to achieve superior performance than existing methods. Using simulated sequencing reads from real individual genotypes in the HapMap data and the 1000 Genomes projects, we show that our method is highly accurate and computationally efficient. Our haplotype predictions improve accuracy over the basic haplotype copying model by ∼20% with comparable computational time, and over another recently proposed approach Hap-SeqX by ∼10% with significantly reduced computational time and memory usage. AVAILABILITY: Publicly available software is available at http://genetics.cs.ucla.edu/harsh CONTACT: [email protected] or [email protected]. Wen-Yun Yang, Farhad Hormozdiari, Zhanyong Wang, Dan He 0001, Bogdan Pasaniuc, Eleazar Eskin |
Bioinform. | 6 |
| 2012 | Hap-seq: An Optimal Algorithm for Haplotype Phasing with Imputation Using Sequencing Data
Dan He 0001, Buhm Han, Eleazar Eskin |
RECOMB | 3 |
| 2012 | CNVeM: Copy Number Variation Detection Using Uncertainty of Read Mapping
Zhanyong Wang, Farhad Hormozdiari, Wen-Yun Yang, Eran Halperin, Eleazar Eskin |
RECOMB | 5 |
| 2012 | Incorporating prior information into association studiesabstractUNLABELLED: Recent technological developments in measuring genetic variation have ushered in an era of genome-wide association studies which have discovered many genes involved in human disease. Current methods to perform association studies collect genetic information and compare the frequency of variants in individuals with and without the disease. Standard approaches do not take into account any information on whether or not a given variant is likely to have an effect on the disease. We propose a novel method for computing an association statistic which takes into account prior information. Our method improves both power and resolution by 8% and 27%, respectively, over traditional methods for performing association studies when applied to simulations using the HapMap data. Advantages of our method are that it is as simple to apply to association studies as standard methods, the results of the method are interpretable as the method reports p-values, and the method is optimal in its use of prior information in regards to statistical power. AVAILABILITY: The method presented herein is available at http://masa.cs.ucla.edu. Gregory Darnell, Dat Duong, Buhm Han, Eleazar Eskin |
Bioinform. | 4 |
| 2011 | Increasing Power of Groupwise Association Test with Likelihood Ratio Test
Jae Hoon Sul, Buhm Han, Eleazar Eskin |
RECOMB | 3 |
| 2011 | Mixed-model coexpression: calculating gene coexpression while accounting for expression heterogeneityabstractMOTIVATION: The analysis of gene coexpression is at the core of many types of genetic analysis. The coexpression between two genes can be calculated by using a traditional Pearson's correlation coefficient. However, unobserved confounding effects may cause inflation of the Pearson's correlation so that uncorrelated genes appear correlated. Many general methods have been suggested, which aim to remove the effects of confounding from gene expression data. However, the residual confounding which is not accounted for by these generic correction procedures has the potential to induce correlation between genes. Therefore, a method that specifically aims to calculate gene coexpression between gene expression arrays, while accounting for confounding effects, is desirable. RESULTS: In this article, we present a statistical model for calculating gene coexpression called mixed model coexpression (MMC), which models coexpression within a mixed model framework. Confounding effects are expected to be encoded in the matrix representing the correlation between arrays, the inter-sample correlation matrix. By conditioning on the information in the inter-sample correlation matrix, MMC is able to produce gene coexpressions that are not influenced by global confounding effects and thus significantly reduce the number of spurious coexpressions observed. We applied MMC to both human and yeast datasets and show it is better able to effectively prioritize strong coexpressions when compared to a traditional Pearson's correlation and a Pearson's correlation applied to data corrected with surrogate variable analysis (SVA). AVAILABILITY: The method is implemented in the R programming language and may be found at http://genetics.cs.ucla.edu/mmc. CONTACT: [email protected]; [email protected]. Nicholas A. Furlotte, Hyun Min Kang, Chun Ye, Eleazar Eskin |
Bioinform. | 4 |
| 2011 | Efficient algorithms for tandem copy number variation reconstruction in repeat-rich regionsabstractMOTIVATION: Structural variations and in particular copy number variations (CNVs) have dramatic effects of disease and traits. Technologies for identifying CNVs have been an active area of research for over 10 years. The current generation of high-throughput sequencing techniques presents new opportunities for identification of CNVs. Methods that utilize these technologies map sequencing reads to a reference genome and look for signatures which might indicate the presence of a CNV. These methods work well when CNVs lie within unique genomic regions. However, the problem of CNV identification and reconstruction becomes much more challenging when CNVs are in repeat-rich regions, due to the multiple mapping positions of the reads. RESULTS: In this study, we propose an efficient algorithm to handle these multi-mapping reads such that the CNVs can be reconstructed with high accuracy even for repeat-rich regions. To our knowledge, this is the first attempt to both identify and reconstruct CNVs in repeat-rich regions. Our experiments show that our method is not only computationally efficient but also accurate. Dan He 0001, Farhad Hormozdiari, Nicholas A. Furlotte, Eleazar Eskin |
Bioinform. | 4 |
| 2011 | Genotyping common and rare variation using overlapping pool sequencingabstractBACKGROUND: Peptide identification from tandem mass spectrometry (MS/MS) data is one of the most important problems in computational proteomics. This technique relies heavily on the accurate assessment of the quality of peptide-spectrum matches (PSMs). However, current MS technology and PSM scoring algorithm are far from perfect, leading to the generation of incorrect peptide-spectrum pairs. Thus, it is critical to develop new post-processing techniques that can distinguish true identifications from false identifications effectively. RESULTS: In this paper, we present a consistency-based PSM re-ranking method to improve the initial identification results. This method uses one additional assumption that two peptides belonging to the same protein should be correlated to each other. We formulate an optimization problem that embraces two objectives through regularization: the smoothing consistency among scores of correlated peptides and the fitting consistency between new scores and initial scores. This optimization problem can be solved analytically. The experimental study on several real MS/MS data sets shows that this re-ranking method improves the identification performance. CONCLUSIONS: The score regularization method can be used as a general post-processing step for improving peptide identifications. Source codes and data sets are available at: http://bioinformatics.ust.hk/SRPI.rar. Dan He 0001, Noah Zaitlen, Bogdan Pasaniuc, Eleazar Eskin, Eran Halperin |
BMC Bioinform. | 4 |
| 2011 | Assembly of non-unique insertion content using next-generation sequencingabstractRecent studies in genomics have highlighted the significance of sequence insertions in determining individual variation. Efforts to discover the content of these sequence insertions have been limited to short insertions and long unique insertions. Much of the inserted sequence in the typical human genome, however, is a mixture of repeated and unique sequence. Current methods are designed to assemble only unique sequence insertions, using reads that do not map to the reference. These methods are not able to assemble repeated sequence insertions, as the reads will map to the reference in a different locus.In this paper, we present a computational method for discovering the content of sequence insertions that are unique, repeated, or a combination of the two. Our method analyzes the read mappings and depth of coverage of paired-end reads to identify reads that originated from inserted sequence. We demonstrate the process of assembling these reads to characterize the insertion content. Our method is based on the idea of segment extension, which progressively extends segments of known content using paired-end reads. We apply our method in simulation to discover the content of inserted sequences in a modified mouse chromosome and show that our method produces reliable results at 40x coverage. Nathaniel Parrish, Farhad Hormozdiari, Eleazar Eskin |
BMC Bioinform. | 3 |
| 2010 | Respecting Markov Equivalence in Computing Posterior Probabilities of Causal Graphical FeaturesabstractThere have been many efforts to identify causal graphical features such as directed edges between random variables from observational data. Recently, Tian et al. proposed a new dynamic programming algorithm which computes marginalized posterior probabilities of directed edge features over all the possible structures in O(n3n) time when the number of parents per node is bounded by a constant, where n is the number of variables of interest. However the main drawback of this approach is that deciding a single appropriate threshold for the existence of the directed edge feature is difficult due to the scale difference of the posterior probabilities between the directed edges forming v-structures and the directed edges not forming v-structures. We claim that computing posterior probabilities of both adjacencies and v-structures is necessary and more effective for discovering causal graphical features, since it allows us to find a single appropriate decision threshold for the existence of the feature that we are testing. For efficient computation, we provide a novel dynamic programming algorithm which computes the posterior probabilities of all of n(n – 1)/2 adjacency and n(n–1 choose 2) v-structure features in O(n3 * 3n) time. Eun Yong Kang, Ilya Shpitser, Eleazar Eskin |
AAAI | 3 |
| 2010 | Effective Algorithms for Fusion Gene Detection
Dan He 0001, Eleazar Eskin |
WABI | 2 |
| 2010 | Multi-marker tagging single nucleotide polymorphism selection using estimation of distribution algorithms
Roberto Santana 0001, Alexander Mendiburu, Noah Zaitlen, Eleazar Eskin, José Antonio Lozano 0001 |
Artif. Intell. Medicine | 4 |
| 2010 | Optimal algorithms for haplotype assembly from whole-genome sequence dataabstractMOTIVATION: Haplotype inference is an important step for many types of analyses of genetic variation in the human genome. Traditional approaches for obtaining haplotypes involve collecting genotype information from a population of individuals and then applying a haplotype inference algorithm. The development of high-throughput sequencing technologies allows for an alternative strategy to obtain haplotypes by combining sequence fragments. The problem of 'haplotype assembly' is the problem of assembling the two haplotypes for a chromosome given the collection of such fragments, or reads, and their locations in the haplotypes, which are pre-determined by mapping the reads to a reference genome. Errors in reads significantly increase the difficulty of the problem and it has been shown that the problem is NP-hard even for reads of length 2. Existing greedy and stochastic algorithms are not guaranteed to find the optimal solutions for the haplotype assembly problem. RESULTS: In this article, we proposed a dynamic programming algorithm that is able to assemble the haplotypes optimally with time complexity O(m x 2(k) x n), where m is the number of reads, k is the length of the longest read and n is the total number of SNPs in the haplotypes. We also reduce the haplotype assembly problem into the maximum satisfiability problem that can often be solved optimally even when k is large. Taking advantage of the efficiency of our algorithm, we perform simulation experiments demonstrating that the assembly of haplotypes using reads of length typical of the current sequencing technologies is not practical. However, we demonstrate that the combination of this approach and the traditional haplotype phasing approaches allow us to practically construct haplotypes containing both common and rare variants. Dan He 0001, Arthur Choi, Knot Pipatsrisawat, Adnan Darwiche, Eleazar Eskin |
Bioinform. | 5 |
| 2010 | Detection and reconstruction of tandemly organized de novo copy number variationsabstractBACKGROUND: The characterization of structural variations (SV) such as insertions, deletions and copy number variations is a critical step in the process of understanding the full genetic architecture of organisms. Copy number variations (CNV) have attracted much recent attention due to their effects on gene expression and disease status. RESULTS: In this paper, we present a method that utilizes next-generation sequencing technologies (NGS), in order to both detect and reconstruct CNVs. We focus on a special type of CNV, namely tandemly organized de novo CNVs, which have been shown to occur with high frequency in the mouse genome. CONCLUSIONS: We apply our method to CNV regions randomly inserted into the reference mouse genome and show that our method achieves good performance for both detection and reconstruction of tandemly organized de novo CNVs. Dan He 0001, Nicholas A. Furlotte, Eleazar Eskin |
BMC Bioinform. | 3 |
| 2009 | Detecting the Presence and Absence of Causal Relationships between Expression of Yeast Genes with Very Few Samples
Eun Yong Kang, Ilya Shpitser, Chun Ye, Eleazar Eskin |
RECOMB | 4 |
| 2009 | An Adaptive and Memory Efficient Algorithm for Genotype Imputation
Hyun Min Kang, Noah Zaitlen, Buhm Han, Eleazar Eskin |
RECOMB | 4 |
| 2009 | Using Network Component Analysis to Dissect Regulatory Networks Mediated by Transcription Factors in YeastabstractUnderstanding the relationship between genetic variation and gene expression is a central question in genetics. With the availability of data from high-throughput technologies such as ChIP-Chip, expression, and genotyping arrays, we can begin to not only identify associations but to understand how genetic variations perturb the underlying transcription regulatory networks to induce differential gene expression. In this study, we describe a simple model of transcription regulation where the expression of a gene is completely characterized by two properties: the concentrations and promoter affinities of active transcription factors. We devise a method that extends Network Component Analysis (NCA) to determine how genetic variations in the form of single nucleotide polymorphisms (SNPs) perturb these two properties. Applying our method to a segregating population of Saccharomyces cerevisiae, we found statistically significant examples of trans-acting SNPs located in regulatory hotspots that perturb transcription factor concentrations and affinities for target promoters to cause global differential expression and cis-acting genetic variations that perturb the promoter affinities of transcription factors on a single gene to cause local differential expression. Although many genetic variations linked to gene expressions have been identified, it is not clear how they perturb the underlying regulatory networks that govern gene expression. Our work begins to fill this void by showing that many genetic variations affect the concentrations of active transcription factors in a cell and their affinities for target promoters. Understanding the effects of these perturbations can help us to paint a more complete picture of the complex landscape of transcription regulation. The software package implementing the algorithms discussed in this work is available as a MATLAB package upon request. Chun Ye, Simon J. Galbraith, James C. Liao, Eleazar Eskin |
PLoS Comput. Biol. | 4 |
| 2008 | Increasing Power in Association Studies by Using Linkage Disequilibrium Structure and Molecular Function as Prior Information
Eleazar Eskin |
RECOMB | 1 |
| 2008 | Efficient Genome Wide Tagging by Reduction to SAT
Arthur Choi, Noah Zaitlen, Buhm Han, Knot Pipatsrisawat, Adnan Darwiche, Eleazar Eskin |
WABI | 6 |
| 2007 | Identification of Deletion Polymorphisms from Haplotypes
Erik Corona, Benjamin J. Raphael, Eleazar Eskin |
RECOMB | 3 |
| 2007 | Reconstructing the Phylogeny of Mobile Elements
Sean O'Rourke, Noah Zaitlen, Nebojsa Jojic, Eleazar Eskin |
RECOMB | 4 |
| 2007 | Discovering tightly regulated and differentially expressed gene sets in whole genome expression dataabstractMOTIVATION: Recently, a new type of expression data is being collected which aims to measure the effect of genetic variation on gene expression in pathways. In these datasets, expression profiles are constructed for multiple strains of the same model organism under the same condition. The goal of analyses of these data is to find differences in regulatory patterns due to genetic variation between strains, often without a phenotype of interest in mind. We present a new method based on notions of tight regulation and differential expression to look for sets of genes which appear to be significantly affected by genetic variation. RESULTS: When we use categorical phenotype information, as in the Alzheimer's and diabetes datasets, our method finds many of the same gene sets as gene set enrichment analysis. In addition, our notion of correlated gene sets allows us to focus our efforts on biological processes subjected to tight regulation. In murine hematopoietic stem cells, we are able to discover significant gene sets independent of a phenotype of interest. Some of these gene sets are associated with several blood-related phenotypes. AVAILABILITY: The programs are available by request from the authors. Chun Ye, Eleazar Eskin |
Bioinform. | 2 |
| 2006 | 10 Years of the International Conference on Research in Computational Molecular Biology (RECOMB)
Sarah J. Aerni, Eleazar Eskin |
RECOMB | 2 |
| 2006 | Discrete profile comparison using information bottleneckabstractSequence homologs are an important source of information about proteins. Amino acid profiles, representing the position-specific mutation probabilities found in profiles, are a richer encoding of biological sequences than the individual sequences themselves. However, profile comparisons are an order of magnitude slower than sequence comparisons, making profiles impractical for large datasets. Also, because they are such a rich representation, profiles are difficult to visualize. To address these problems, we describe a method to map probabilistic profiles to a discrete alphabet while preserving most of the information in the profiles. We find an informationally optimal discretization using the Information Bottleneck approach (IB). We observe that an 80-character IB alphabet captures nearly 90% of the amino acid occurrence information found in profiles, compared to the consensus sequence's 78%. Distant homolog search with IB sequences is 88% as sensitive as with profiles compared to 61% with consensus sequences (AUC scores 0.73, 0.83, and 0.51, respectively), but like simple sequence comparison, is 30 times faster. Discrete IB encoding can therefore expand the range of sequence problems to which profile information can be applied to include batch queries over large databases like SwissProt, which were previously computationally infeasible. Sean O'Rourke, Gal Chechik, Robin Friedman, Eleazar Eskin |
BMC Bioinform. | 4 |
| 2005 | The Homology Kernel: A Biologically Motivated Sequence Embedding into Euclidean Space
Eleazar Eskin, Sagi Snir |
CIBCB | 1 |
| 2005 | A comparative evaluation of two algorithms for Windows Registry Anomaly DetectionabstractWe present a component anomaly detector for a host-based intrusion detection system (IDS) for Microsoft Windows. The core of the detector is a learning-based anomaly detection algorithm that detects attacks on a host machine by looking for anomalous Salvatore J. Stolfo, Frank Apap, Eleazar Eskin, Katherine A. Heller, Shlomo Hershkop, Andrew Honig, Krysta M. Svore |
J. Comput. Secur. | 3 |
| 2005 | Searching Genomes for Noncoding RNA Using FastRabstractThe discovery of novel noncoding RNAs has been among the most exciting recent developments in biology. It has been hypothesized that there is, in fact, an abundance of functional noncoding RNAs (ncRNAs) with various catalytic and regulatory functions. However, the inherent signal for ncRNA is weaker than the signal for protein coding genes, making these harder to identify. We consider the following problem: Given an RNA sequence with a known secondary structure, efficiently detect all structural homologs in a genomic database by computing the sequence and structure similarity to the query. Our approach, based on structural filters that eliminate a large portion of the database while retaining the true homologs, allows us to search a typical bacterial genome in minutes on a standard PC. The results are two orders of magnitude better than the currently available software for the problem. We applied FastR to the discovery of novel riboswitches, which are a class of RNA domains found in the untranslated regions. They are of interest because they regulate metabolite synthesis by directly binding metabolites. We searched all available eubacterial and archaeal genomes for riboswitches from purine, lysine, thiamin, and riboflavin subfamilies. Our results point to a number of novel candidates for each of these subfamilies and include genomes that were not known to contain riboswitches. Shaojie Zhang 0001, Brian Haas, Eleazar Eskin, Vineet Bafna |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2004 | Discrete profile alignment via constrained information bottleneckabstractAmino acid profiles, which capture position-specific mutation prob- abilities, are a richer encoding of biological sequences than the in- dividual sequences themselves. However, profile comparisons are much more computationally expensive than discrete symbol com- parisons, making profiles impractical for many large datasets. Fur- thermore, because they are such a rich representation, profiles can be difficult to visualize. To overcome these problems, we propose a discretization for profiles using an expanded alphabet representing not just individual amino acids, but common profiles. By using an extension of information bottleneck (IB) incorporating constraints and priors on the class distributions, we find an informationally optimal alphabet. This discretization yields a concise, informative textual representation for profile sequences. Also alignments be- tween these sequences, while nearly as accurate as the full profile- profile alignments, can be computed almost as quickly as those between individual or consensus sequences. A full pairwise align- ment of SwissProt would take years using profiles, but less than 3 days using a discrete IB encoding, illustrating how discrete en- coding can expand the range of sequence problems to which profile information can be applied. 1 Introduction One of the most powerful techniques in protein analysis is the comparison of a target amino acid sequence with phylogenetically related or homologous proteins. Such comparisons give insight into which portions of the protein are important by revealing the parts that were conserved through natural selection. While mutations in non-functional regions may be harmless, mutations in functional regions are often lethal. For this reason, functional regions of a protein tend to be conserved between organisms while non-functional regions diverge. Department of Computer Science and Engineering, University of California San Diego Department of Computer Science, Stanford University Many of the state-of-the-art protein analysis techniques incorporate homologous sequences by representing a set of homologous sequences as a probabilistic profile, a sequence of the marginal distributions of amino acids at each position in the sequence. For example, Yona et al.[10] uses profiles to align distant homologues from the SCOP database[3]; the resulting alignments are similar to results from structural alignments, and tend to reflect both secondary and tertiary protein structure. The PHD algorithm[5] uses profiles purely for structure prediction. PSIBLAST[6] uses them to refine database searches. Although profiles provide a lot of information about the sequence, the use of pro- files comes at a steep price. While extremely efficient string algorithms exist for aligning protein sequences (Smith-Waterman[8]) and performing database queries (BLAST[6]), these algorithms operate on strings and are not immediately applica- ble to profile alignment or profile database queries. While profile-based methods can be substantially more accurate than sequence-based ones, they can require at least an order of magnitude more computation time, since substitution penalties must be calculated by computing distances between probability distributions. This makes profiles impractical for use with large bioinformatics databases like SwissProt, which recently passed 150,000 sequences. Another drawback of profile as compared to string representations is that it is much more difficult to visually interpret a sequence of 20 dimensional vectors than a sequence of letters. Discretizing the profiles addresses both of these problems. First, once a profile is rep- resented using a discrete alphabet, alignment and database search can be performed using the efficient string algorithms developed for sequences. For example, when aligning sequences of 1000 elements, runtime decreases from 20 seconds for profiles to 2 for discrete sequences. Second, by representing each class as a letter, discretized profiles can be presented in plain text like the original or consensus sequences, while conveying more information about the underlying profiles. This makes them more accurate than consensus sequences, and more dense than sequence logos (see figure 1). To make this representation intuitive, we want the discretization not only to minimize information loss, but also to reflect biologically meaningful categories by forming a superset of the standard 20-character amino acid alphabet. For example, we use "A" and "a" for strongly- and weakly-conserved Alanine. This formulation demands two types of constraints: similarities of the centroids to predefined values, and specific structural similarities between strongly- and weakly-conserved variants. We show below how these constraints can be added to the original IB formalism. In this paper, we present a new discrete representation of proteins that takes into account information from homologues. The main idea behind our approach is to compress the space of probabilistic profiles in a data-dependent manner by clustering the actual profiles and representing them by a small alphabet of distributions. Since this discretization removes some of the information carried by the full profiles, we cluster the distribution in a way that is directly targeted at minimizing the information loss. This is achieved using a variant of Information Bottleneck (IB)[9], a distributional clustering approach for informationally optimal discretization. We apply our algorithm to a subset of MEROPS[4], a database of peptidases or- ganized structurally by family and clan, and analyze the results in terms of both information loss and alignment quality. We show that multivariate IB in particular preserves much of the information in the original profiles using a small number of classes. Furthermore, optimal alignments for profile sequences encoded with these classes are much closer to the original profile-profile alignments than are alignments between the seed proteins. IB discretization is therefore an attractive way to gain some of the additional sensitivity of profiles with less computational cost. 0.0 0.0 0.0 0.09 0.34 0.23 0.12 0.0 0.0 0.0 0.0 0.0 0.0 0.04 0.01 0.01 0.03 0.0 0.0 0.0 0.0 0.0 1.0 0.01 0.05 0.14 0.09 0.0 1.0 0.0 0.0 0.0 0.0 0.38 0.04 0.00 0.04 0.0 0.0 0.0 0.0 0.0 0.0 0.06 0.00 0.08 0.04 0.0 0.0 1.0 0.0 0.0 0.0 0.00 0.06 0.01 0.03 1.0 0.0 0.0 0.0 0.0 0.0 0.02 0.00 0.04 0.00 0.0 0.0 0.0 N 0.0 0.0 0.0 0.00 0.00 0.03 0.00 0.0 0.0 0.0 ND GDF 0.0 0.0 0.0 0.04 0.01 0.01 0.00 0.0 0.0 0.0 S EAAS S V PT S A A T D F F D N G S L K D Q E T R N H F Y V 0.0 0.0 0.0 0.01 0.01 0.00 0.09 0.0 0.0 0.0 Q Y E A A A A A A 0.0 0.0 0.0 0.00 0.00 0.03 0.00 0.0 0.0 0.0 (b) 0.5 1.0 0.0 0.05 0.05 0.01 0.01 0.0 0.0 0.0 0.0 0.0 0.0 0.02 0.00 0.23 0.00 0.0 0.0 0.0 P00790 Seq.: ---EAPT--- 0.0 0.0 0.0 0.04 0.05 0.00 0.00 0.0 0.0 0.0 Consensus Seq.: NNDEAASGDF 0.0 0.0 0.0 0.04 0.01 0.00 0.00 0.0 0.0 0.0 0.5 0.0 0.0 0.16 0.10 0.06 0.29 0.0 0.0 0.0 IB Seq.: NNDeaptGDF 0.0 0.0 0.0 0.02 0.10 0.05 0.20 0.0 0.0 0.0 0.0 0.0 0.0 0.00 0.14 0.03 0.04 0.0 0.0 0.0 (c) 0.0 0.0 0.0 0.00 0.00 0.00 0.00 0.0 0.0 0.0 0.0 0.0 0.0 0.01 0.00 0.04 0.04 0.0 0.0 0.0 (a) Figure 1: (a) Profile, (b) sequence logo[2], and (c) textual representations for part of an alignment of Pepsin A precursor P00790, showing IB's concision compared to profiles and logos, and its precision compared to single sequences. 2 Information Bottleneck Information Bottleneck [9] is an information theoretic approach for distributional clustering. Given a joint distribution p(X, Y ) of two random variables X and Y , the goal is to obtain a compressed representation C of X, while preserving the informa- tion about Y . The two goals of compression and information preservation are quan- tified by the same measure of mutual information I(X; Y ) = p(x, y) log p(x,y) x,y p(x)p(y) and the problem is therefore defined as the constrained optimization problem minp(c|x):I(C;Y )>K I(C; X) where K is a constraint on the level of information preserved about Y , and the problem should also obey the constraints p(y|c) = p(y|x)p(x|c) and p(y) = p(y|x)p(x). This constrained optimization can be x x reformulated using Lagrange multipliers, and turned into a tradeoff optimization function with Lagrange multiplier : min L def = I(C; X) - I(C; Y ) (1) p(c|x) As an unsupervised learning technique, IB aims to characterize the set of solutions for the complete spectrum of constraint values K. This set of solutions is identical to the set of solutions of the tradeoff optimization problem obtained for the spectrum of values. When X is discrete, its natural compression is fuzzy clustering. In this case, the problem is not convex and cannot be guaranteed to contain a single global minimum. Fortunately, its solutions can be characterized analytically by a set of self consistent equations. These self consistent equations can then be used in an iterative algorithm that is guaranteed to converge to a local minimum. While the optimal solutions of the IB functional are in general soft clusters, in practice, hard cluster solutions are sometimes more easily interpreted. A series of algorithms was developed for hard IB, including an algorithm that can be viewed as a one-step look-ahead sequential version of K-Means [7]. To apply IB to the problem of profiles discretization discussed here, X is a given set of probabilistic profiles obtained from a set of aligned sequences and Y is the set of 20 amino acids. 2.1 Constraints on centroids' semantics The application studied in this paper differs from standard IB applications in that we are interested in obtaining a representation that is both efficient and biologi- cally meaningful. This requires that we add two kinds of constraints on clusters' distributions, discussed below. First, some clusters' meanings are naturally determined by limiting them to corre- spond to the common 20-letter alphabet used to describe amino acids. From the point of view of distributions over amino acids, each of these symbols is used today as the delta function distribution which is fully concentrated on a single amino acid. For the goal of finding an efficient representation, we require the centroids to be close to these delta distributions. More generally, we require the centroids to be close to some predefined values ^ ci, thus adding constraints to the IB target function of the form DKL[p(y|^ ci)||p(y|ci)] < Ki for each constrained centroid. While solving the constrained optimization problem is difficult, the corresponding tradeoff opti- mization problem can be made very similar to standard IB. With the additional constraints, the IB trade-off optimization problem becomes min L I(C; X) - I(C; Y ) + (ci)DKL[p(y|^ ci)||p(y|ci)] . (2) p(c|x) ciC We now use the following identity p(x, c)DKL[p(y|x)||p(y|c)] x,c = p(x) p(y|x) log p(y|x) - p(c) log p(y|c) p(y|x)p(x|c) x y c y x = -H(Y |X) + H(Y |C) = I(X; Y ) - I(Y ; C) to rewrite the IB functional of Eq. (1) as L = I(C; X) + p(x, c)DKL[p(y|x)||p(y|c)] - I(X; Y ) cC xX When (ci) 1 we can similarly rewrite Eq. (2) as L = I(C; X) + p(x) p(ci|x)DKL[p(y|x)||p(y|ci)] (3) xX ciC + (ci)DKL[p(y|^ ci)||p(y|ci)] - I(X; Y ) ciC = I(C; X) + p(x ) p(ci|x )DKL[p(y|x )||p(y|ci)] - I(X; Y ) x X ciC The optimization problem therefore becomes equivalent to the original IB problem, but with a modified set of samples x X , containing X plus additional "pseudo- counts" or biases. This is similar to the inclusion of priors in Bayesian estimation. Formulated this way, the biases can be easily incorporated in standard IB algorithms by adding additional pseudo-counts x with prior probability p(x ) = i(c). 2.2 Constraints on relations between centroids We want our discretization to capture correlations between strongly- and weakly- conserved variants of the same symbol. This can be done with standard IB using separate classes for the alternatives. However, since the distributions of other amino acids in these two variants are likely to be related, it is preferable to define a single shared prior for both variants, and to learn a model capturing their correlation. Friedman et al.[1] describe multivariate information bottleneck (mIB), an extension of information bottleneck to joint distributions over several correlated input and cluster variables. For profile discretization, we define two compression variables connected as in Friedman's "parallel IB": an amino acid class C {A, C, . . .} with an associated prior, and a strength S {0, 1}. Since this model correlates strong and weak variants of each category, it requires fewer priors than simple IB. It also has fewer parameters: a multivariate model with ns strengths and nc classes has as many categories as a univariate one with nc = nsnc classes, but has only ns +nc -2 free parameters for each x, instead of nsnc - 1. Sean O'Rourke, Gal Chechik, Robin Friedman, Eleazar Eskin |
NIPS | 4 |
| 2004 | From profiles to patterns and back again: a branch and bound algorithm for finding near optimal motif profilesabstractAn important part of deciphering gene regulatory mechanisms is discovering transcription factor binding sites. In many cases, these sites can be detected because they are often overrepresented in genomic sequences. The detection of the overrepresented signals in sequences, or motif-finding has become a central problem in computational biology. There are two major computational frameworks for attacking the motif finding problem which differ in their representation of the signals. The most popular is the profile or PSSM (Position Specific Scoring Matrix) representation. The goal of these algorithms is to obtain probabilistic representations of the overrepresented signals. Another is the consensus pattern or pattern with mismatches representation which represents a signal as discrete consensus pattern and allows some mismatches to occur in each instance of the pattern. The advantage of profiles is the expressiveness of their representation while the advantage of the consensus pattern approach is the existence of efficient algorithms that guarantee discovery of the best patterns. In this paper we present a unified framework for motif finding which encompasses both the profile representation and the consensus pattern representation. We prove that the problem of discovering the best profiles can be solved by considering a degenerate version of the problem of finding the best consensus patterns. The main advantage of our framework is that it motivates a novel algorithm, MITRA-PSSM, which discovers profiles, yet provides some of the guarantees of discovering the best signals. The algorithm searches for best profiles with respect to information content which is the same criterion of popular algorithms such as MEME and CONSENSUS. MITRA-PSSM is specifically designed for searching for profiles in this framework and introduces a novel notion of scoring consensus patterns, discrete information content. MITRA-PSSM is available for public use via webserver at http://www.calit2.net/compbio/mitra/. Eleazar Eskin |
RECOMB | 1 |
| 2004 | Haplotype reconstruction from genotype data using Imperfect PhylogenyabstractUNLABELLED: Critical to the understanding of the genetic basis for complex diseases is the modeling of human variation. Most of this variation can be characterized by single nucleotide polymorphisms (SNPs) which are mutations at a single nucleotide position. To characterize the genetic variation between different people, we must determine an individual's haplotype or which nucleotide base occurs at each position of these common SNPs for each chromosome. In this paper, we present results for a highly accurate method for haplotype resolution from genotype data. Our method leverages a new insight into the underlying structure of haplotypes that shows that SNPs are organized in highly correlated 'blocks'. In a few recent studies, considerable parts of the human genome were partitioned into blocks, such that the majority of the sequenced genotypes have one of about four common haplotypes in each block. Our method partitions the SNPs into blocks, and for each block, we predict the common haplotypes and each individual's haplotype. We evaluate our method over biological data. Our method predicts the common haplotypes perfectly and has a very low error rate (<2% over the data) when taking into account the predictions for the uncommon haplotypes. Our method is extremely efficient compared with previous methods such as PHASE and HAPLOTYPER. Its efficiency allows us to find the block partition of the haplotypes, to cope with missing data and to work with large datasets. AVAILABILITY: The algorithm is available via a Web server at http://www.calit2.net/compbio/hap/ Eran Halperin, Eleazar Eskin |
Bioinform. | 2 |
| 2004 | Mismatch string kernels for discriminative protein classificationabstractMOTIVATION: Classification of proteins sequences into functional and structural families based on sequence homology is a central problem in computational biology. Discriminative supervised machine learning approaches provide good performance, but simplicity and computational efficiency of training and prediction are also important concerns. RESULTS: We introduce a class of string kernels, called mismatch kernels, for use with support vector machines (SVMs) in a discriminative approach to the problem of protein classification and remote homology detection. These kernels measure sequence similarity based on shared occurrences of fixed-length patterns in the data, allowing for mutations between patterns. Thus, the kernels provide a biologically well-motivated way to compare protein sequences without relying on family-based generative models such as hidden Markov models. We compute the kernels efficiently using a mismatch tree data structure, allowing us to calculate the contributions of all patterns occurring in the data in one pass while traversing the tree. When used with an SVM, the kernels enable fast prediction on test sequences. We report experiments on two benchmark SCOP datasets, where we show that the mismatch kernel used with an SVM classifier performs competitively with state-of-the-art methods for homology detection, particularly when very few training examples are available. Examination of the highest-weighted patterns learned by the SVM classifier recovers biologically important motifs in protein families and superfamilies. Christina S. Leslie, Eleazar Eskin, Adiel Cohen, Jason Weston, William Stafford Noble |
Bioinform. | 2 |
| 2003 | Laplace PropagationabstractWe present a novel method for approximate inference in Bayesian mod- els and regularized risk functionals. It is based on the propagation of mean and variance derived from the Laplace approximation of condi- tional probabilities in factorizing distributions, much akin to Minka’s Expectation Propagation. In the jointly normal case, it coincides with the latter and belief propagation, whereas in the general case, it provides an optimization strategy containing Support Vector chunking, the Bayes Committee Machine, and Gaussian Process chunking as special cases. Alexander J. Smola, S. V. N. Vishwanathan, Eleazar Eskin |
NIPS | 3 |
| 2003 | Large scale reconstruction of haplotypes from genotype dataabstractCritical to the understanding of the genetic basis for complex diseases is the modeling of human variation. Most of this variation can be characterized by single nucleotide polymorphisms (SNPs) which are mutations at a single nucleotide position. To characterize an individual's variation, we must determine an individual's haplotype or which nucleotide base occurs at each position of these common SNPs for each chromosome. In this paper, we present results for a highly accurate method for haplotype resolution from genotype data. Our method leverages a new insight into the underlying structure of haplotypes which shows that SNPs are organized in highly correlated "blocks". The majority of individuals have one of about four common haplotypes in each block. Our method partitions the SNPs into blocks and for each block, we predict the common haplotypes and each individual's haplotype. We evaluate our method over biological data. Our method predicts the common haplotypes perfectly and has a very low error rate (0.47%) when taking into account the predictions for the uncommon haplotypes. Our method is extremely efficient compared to previous methods, (a matter of seconds where previous methods needed hours). Its efficiency allows us to find the block partition of the haplotypes, to cope with missing data and to work with large data sets such as genotypes for thousands of SNPs for hundreds of individuals. The algorithm is available via webserver at http://www.cs.columbia.edu/compbio/hap. Eleazar Eskin, Eran Halperin, Richard M. Karp |
RECOMB | 1 |
| 2002 | A Kernel Approach for Learning from almost Orthogonal Patterns
Bernhard Schölkopf, Jason Weston, Eleazar Eskin, Christina S. Leslie, William Stafford Noble |
ECML | 3 |
| 2002 | Finding composite regulatory patterns in DNA sequencesabstractPattern discovery in unaligned DNA sequences is a fundamental problem in computational biology with important applications in finding regulatory signals. Current approaches to pattern discovery focus on monad patterns that correspond to relatively short contiguous strings. However, many of the actual regulatory signals are composite patterns that are groups of monad patterns that occur near each other. A difficulty in discovering composite patterns is that one or both of the component monad patterns in the group may be 'too weak'. Since the traditional monad-based motif finding algorithms usually output one (or a few) high scoring patterns, they often fail to find composite regulatory signals consisting of weak monad parts. In this paper, we present a MITRA (MIsmatch TRee Algorithm) approach for discovering composite signals. We demonstrate that MITRA performs well for both monad and composite patterns by presenting experiments over biological and synthetic data. Eleazar Eskin, Pavel A. Pevzner |
ISMB | 1 |
| 2002 | Mismatch String Kernels for SVM Protein ClassificationabstractWe introduce a class of string kernels, called mismatch kernels, for use with support vector machines (SVMs) in a discriminative approach to the protein classification problem. These kernels measure sequence sim- ilarity based on shared occurrences of -length subsequences, counted with up to mismatches, and do not rely on any generative model for the positive training sequences. We compute the kernels efficiently using a mismatch tree data structure and report experiments on a benchmark SCOP dataset, where we show that the mismatch kernel used with an SVM classifier performs as well as the Fisher kernel, the most success- ful method for remote homology detection, while achieving considerable computational savings. Christina S. Leslie, Eleazar Eskin, Jason Weston, William Stafford Noble |
NIPS | 2 |
| 2002 | MET: an experimental system for Malicious Email TrackingabstractDespite the use of state of the art methods to protect against malicious programs, they continue to threaten and damage computer systems around the world. In this paper we present MET, the Malicious Email Tracking system, designed to automatically report statistics on the flow behavior of malicious software delivered via email attachments both at a local and global level. MET can help reduce the spread of malicious software worldwide, especially self-replicating viruses, as well as provide further insight toward minimizing damage caused by malicious programs in the future. In addition, the system can help system administrators detect all of the points of entry of a malicious email into a network. The core of MET's operation is a database of statistics about the trajectory of email attachments in and out of a network system, and the culling together of these statistics across networks to present a global view of the spread of the malicious software. From a statistical perspective sampling only a small amount of traffic (for example, .1 %) of a very large email stream is sufficient to detect suspicious or otherwise new email viruses that may be undetected by standard signature-based scanners. Therefore, relatively few MET installations would be necessary to gather sufficient data in order to provide broad protection services. Small scale simulations are presented to demonstrate MET in operation and suggests how detection of new virus propagations via flow statistics can be automated. Manasi Bhattacharyya, Shlomo Hershkop, Eleazar Eskin |
NSPW | 3 |
| 2002 | A Kernel Approach for Learning from Almost Orthogonal Patterns
Bernhard Schölkopf, Jason Weston, Eleazar Eskin, Christina S. Leslie, William Stafford Noble |
PKDD | 3 |
| 2002 | Detecting Malicious Software by Monitoring Anomalous Windows Registry Accesses
Frank Apap, Andrew Honig, Shlomo Hershkop, Eleazar Eskin, Salvatore J. Stolfo |
RAID | 4 |
| 2001 | Data Mining Methods for Detection of New Malicious ExecutablesabstractA serious security threat today is malicious executables, especially new, unseen malicious executables often arriving as email attachments. These new malicious executables are created at the rate of thousands every year and pose a serious security threat. Current anti-virus systems attempt to detect these new malicious programs with heuristics generated by hand. This approach is costly and oftentimes ineffective. We present a data mining framework that detects new, previously unseen malicious executables accurately and automatically. The data mining framework automatically found patterns in our data set and used these patterns to detect a set of new malicious binaries. Comparing our detection methods with a traditional signature-based method, our method more than doubles the current detection rates for new malicious executables. Matthew G. Schultz, Eleazar Eskin, Erez Zadok, Salvatore J. Stolfo |
S&P | 2 |
| 2000 | Anomaly Detection over Noisy Data using Learned Probability Distributions
Eleazar Eskin |
ICML | 1 |
| 2000 | Protein Family Classification Using Sparse Markov Transducers
Eleazar Eskin, William Stafford Noble, Yoram Singer |
ISMB | 1 |
| 1999 | Detecting Text Similarity over Short Passages: Exploring Linguistic Feature Combinations via Machine Learning
Vasileios Hatzivassiloglou, Judith L. Klavans, Eleazar Eskin |
EMNLP | 3 |
| 1999 | Genetic programming applied to Othello: introducing students to machine learning researchabstractIn this paper we describe and analyze a three week assignment that was given in a Machine Learning course at Columbia University. The assignment presented students with an introduction to machine learning research. The assignment required students to apply Genetic Programming to evolve algorithms that play the board game Othello. The students were provided with an implemented experimental approach as a starting point. The students were required to perform their own experimental modifications corresponding to research issues in machine learning. The results of student experiments were good both in terms of research and in terms of student learning. All relevant code, documentation and information about GPOthello is available at the following url: http://www.cs.columbia.edu/~evs/ml/othello.html. Eleazar Eskin, Eric V. Siegel |
SIGCSE | 1 |