Evan E. Eichler

dblp:14/3719 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-8246-4014ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 needLR: long-read structural variant annotation with population-scale frequency estimation
abstract
SUMMARY: We present needLR, a structural variant (SV) annotation tool that can be used for filtering and prioritization of candidate pathogenic SVs from long-read sequencing data using population allele frequencies, annotations for genomic context, and gene-phenotype associations. When using population data from 500 presumably healthy individuals to evaluate nine test cases with known pathogenic SVs, needLR assigned allele frequencies to over 97.5% of all detected SVs and reduced the average number of novel genic SVs to 121 per case while retaining all known pathogenic variants. AVAILABILITY AND IMPLEMENTATION: needLR is implemented in bash with dependencies including Truvari v4.2.2, BEDTools v2.31.1, and BCFtools v1.19. Source code, documentation, and pre-computed population allele frequency data are freely available at https://github.com/jgust1/needLR under an MIT license and archived on Zenodo at https://zenodo.org/records/19463479.
Jonas A Gustafson, Jiadong Lin, Miranda P. G. Zalusky, Evan E. Eichler, Danny E. Miller
Bioinform.4
2025 SVbyEye: a visual tool to characterize structural variation among whole-genome assemblies
abstract
MOTIVATION: We are now in the era of being able to routinely generate highly contiguous (near telomere-to-telomere) genome assemblies of human and nonhuman species. Complex structural variation and regions of rapid evolutionary turnover are being discovered for the first time. Thus, efficient and informative visualization tools are needed to evaluate and directly observe structural differences between two or more genomes. RESULTS: We developed SVbyEye, an open-source R package to visualize and annotate sequence-to-sequence alignments along with various functionalities to process these alignments. The tool facilitates the characterization of complex structural variants in the context of sequence homology helping resolve the mechanisms underlying their formation. AVAILABILITY AND IMPLEMENTATION: SVbyEye is available on GitHub (https://github.com/daewoooo/SVbyEye) and via Zenodo (https://doi.org/10.5281/zenodo.15303553).
David Porubsky, Xavi Guitart, Dongahn Yoo, Philip C. Dishuck, William T. Harvey, Evan E. Eichler
Bioinform.6
2024 Identification and annotation of centromeric hypomethylated regions with CDR-Finder
abstract
MOTIVATION: Centromeres are chromosomal regions historically understudied with sequencing technologies due to their repetitive nature and short-read mapping limitations. However, recent improvements in long-read sequencing allow for the investigation of complex regions of the genome at the sequence and epigenetic levels. RESULTS: Here, we present Centromere Dip Region (CDR)-Finder: a tool to identify regions of hypomethylation within the centromeres of high-quality, contiguous genome assemblies. These regions are typically associated with a unique type of chromatin containing the histone H3 variant CENP-A, which marks the location of the kinetochore. CDR-Finder identifies the CDRs in large and short centromeres and generates a BED file indicating the location of the CDRs within the centromere. It also outputs a plot for visualization, validation, and downstream analysis. AVAILABILITY AND IMPLEMENTATION: CDR-Finder is available at https://github.com/EichlerLab/CDR-Finder.
Francesco Kumara Mastrorosa, Keisuke Oshima, Allison N. Rozanski, William T. Harvey, Evan E. Eichler, Glennis Logsdon
Bioinform.5
2023 GAVISUNK: genome assembly validation via inter-SUNK distances in Oxford Nanopore reads
abstract
MOTIVATION: Highly contiguous de novo phased diploid genome assemblies are now feasible for large numbers of species and individuals. Methods are needed to validate assembly accuracy and detect misassemblies with orthologous sequencing data to allow for confident downstream analyses. RESULTS: We developed GAVISUNK, an open-source pipeline that detects misassemblies and produces a set of reliable regions genome-wide by assessing concordance of distances between unique k-mers in Pacific Biosciences high-fidelity assemblies and raw Oxford Nanopore Technologies reads. AVAILABILITY AND IMPLEMENTATION: GAVISUNK is available at https://github.com/pdishuck/GAVISUNK. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Philip C. Dishuck, Allison N. Rozanski, Glennis Logsdon, David Porubsky, Evan E. Eichler
Bioinform.5
2022 StainedGlass: interactive visualization of massive tandem repeat structures with identity heatmaps
abstract
SUMMARY: The visualization and analysis of genomic repeats is typically accomplished using dot plots; however, the emergence of telomere-to-telomere assemblies with multi-megabase repeats requires new visualization strategies. Here, we introduce StainedGlass, which can generate publication-quality figures and interactive visualizations that depict the identity and orientation of multi-megabase tandem repeat structures at a genome-wide scale. The tool can rapidly reveal higher-order structures and improve the inference of evolutionary history for some of the most complex regions of genomes. AVAILABILITY AND IMPLEMENTATION: StainedGlass is implemented using Snakemake and available open source under the MIT license at https://mrvollger.github.io/StainedGlass/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mitchell R. Vollger, Peter Kerpedjiev, Adam M. Phillippy, Evan E. Eichler
Bioinform.4
2018 Strand-seq enables reliable separation of long reads by chromosome via expectation maximization
abstract
Motivation: Current sequencing technologies are able to produce reads orders of magnitude longer than ever possible before. Such long reads have sparked a new interest in de novo genome assembly, which removes reference biases inherent to re-sequencing approaches and allows for a direct characterization of complex genomic variants. However, even with latest algorithmic advances, assembling a mammalian genome from long error-prone reads incurs a significant computational burden and does not preclude occasional misassemblies. Both problems could potentially be mitigated if assembly could commence for each chromosome separately. Results: To address this, we show how single-cell template strand sequencing (Strand-seq) data can be leveraged for this purpose. We introduce a novel latent variable model and a corresponding Expectation Maximization algorithm, termed SaaRclust, and demonstrates its ability to reliably cluster long reads by chromosome. For each long read, this approach produces a posterior probability distribution over all chromosomes of origin and read directionalities. In this way, it allows to assess the amount of uncertainty inherent to sparse Strand-seq data on the level of individual reads. Among the reads that our algorithm confidently assigns to a chromosome, we observed more than 99% correct assignments on a subset of Pacific Bioscience reads with 30.1× coverage. To our knowledge, SaaRclust is the first approach for the in silico separation of long reads by chromosome prior to assembly. Availability and implementation: https://github.com/daewoooo/SaaRclust.
Maryam Ghareghani, David Porubsky, Ashley D. Sanders, Sascha Meiers, Evan E. Eichler, Jan O. Korbel, Tobias Marschall
Bioinform.5
2017 Resolving Multicopy Duplications de novo Using Polyploid Phasing
Mark J. P. Chaisson, Sudipto Mukherjee 0001, Sreeram Kannan, Evan E. Eichler
RECOMB4
2014 ORMAN: Optimal resolution of ambiguous RNA-Seq multimappings in the presence of novel isoforms
abstract
MOTIVATION: RNA-Seq technology is promising to uncover many novel alternative splicing events, gene fusions and other variations in RNA transcripts. For an accurate detection and quantification of transcripts, it is important to resolve the mapping ambiguity for those RNA-Seq reads that can be mapped to multiple loci: >17% of the reads from mouse RNA-Seq data and 50% of the reads from some plant RNA-Seq data have multiple mapping loci. In this study, we show how to resolve the mapping ambiguity in the presence of novel transcriptomic events such as exon skipping and novel indels towards accurate downstream analysis. We introduce ORMAN ( O ptimal R esolution of M ultimapping A mbiguity of R N A-Seq Reads), which aims to compute the minimum number of potential transcript products for each gene and to assign each multimapping read to one of these transcripts based on the estimated distribution of the region covering the read. ORMAN achieves this objective through a combinatorial optimization formulation, which is solved through well-known approximation algorithms, integer linear programs and heuristics. RESULTS: On a simulated RNA-Seq dataset including a random subset of transcripts from the UCSC database, the performance of several state-of-the-art methods for identifying and quantifying novel transcripts, such as Cufflinks, IsoLasso and CLIIQ, is significantly improved through the use of ORMAN. Furthermore, in an experiment using real RNA-Seq reads, we show that ORMAN is able to resolve multimapping to produce coverage values that are similar to the original distribution, even in genes with highly non-uniform coverage. AVAILABILITY: ORMAN is available at http://orman.sf.net
Phuong Dao, Ibrahim Numanagic, Yen-Yi Lin, Faraz Hach, Emre Karakoç, Nilgun Donmez, Colin C. Collins, Evan E. Eichler, Süleyman Cenk Sahinalp
Bioinform.8
2012 Sensitive and fast mapping of di-base encoded reads
abstract
Bioinformatics (2011) 27(4), 1915–1921. The authors find it worth mentioning that the parameters used to run the PerM mapper were not optimal to achieve full sensitivity. Based on the new recommendations of the developers of PerM, we used the latest version of PerM (v. 0.3.6), and updated two parameters as follows: –seed F2 (full sensitivity for 1 SNPs); -v 2 (number of mismatches); -k 1 000 000 (maximum number of alignment for a read); -A (report all possible mapping for a reads). Previously, we have used ‘–seed S20 -k 10000 -v 4’. With this update, PerM now achieves full sensitivity in our simulation experiment. With real datasets (Table 6), PerM tends to map more reads compared with Bowtie, but maps slightly less than Mapreads and SOCS. We would like to apologize for the previous parameter sets we used for PerM, due to our misinterpretation of its documentation. We now update the relevant rows in Tables 3 and 6 as follows. Performance of PerM with simulated datasets considering the new parameters Reads are simulated from human reference genome build 35 (chromosome 1). Set 1: no errors; Set 2: color errors; Set 3: substitutions. Performance of PerM with real datasets using the new parameters
Farhad Hormozdiari, Faraz Hach, Süleyman Cenk Sahinalp, Evan E. Eichler, Can Alkan
Bioinform.4
2011 Simultaneous Structural Variation Discovery in Multiple Paired-End Sequenced Genomes
Fereydoun Hormozdiari, Iman Hajirasouliha, Andrew W. McPherson, Evan E. Eichler, Süleyman Cenk Sahinalp
RECOMB4
2011 Sensitive and fast mapping of di-base encoded reads
abstract
MOTIVATION: Discovering variation among high-throughput sequenced genomes relies on efficient and effective mapping of sequence reads. The speed, sensitivity and accuracy of read mapping are crucial to determining the full spectrum of single nucleotide variants (SNVs) as well as structural variants (SVs) in the donor genomes analyzed. RESULTS: We present drFAST, a read mapper designed for di-base encoded 'color-space' sequences generated with the AB SOLiD platform. drFAST is specially designed for better delineation of structural variants, including segmental duplications, and is able to return all possible map locations and underlying sequence variation of short reads within a user-specified distance threshold. We show that drFAST is more sensitive in comparison to all commonly used aligners such as Bowtie, BFAST and SHRiMP. drFAST is also faster than both BFAST and SHRiMP and achieves a mapping speed comparable to Bowtie. AVAILABILITY: The source code for drFAST is available at http://drfast.sourceforge.net
Farhad Hormozdiari, Faraz Hach, Süleyman Cenk Sahinalp, Evan E. Eichler, Can Alkan
Bioinform.4
2010 Detection and characterization of novel sequence insertions using paired-end next-generation sequencing
abstract
MOTIVATION: In the past few years, human genome structural variation discovery has enjoyed increased attention from the genomics research community. Many studies were published to characterize short insertions, deletions, duplications and inversions, and associate copy number variants (CNVs) with disease. Detection of new sequence insertions requires sequence data, however, the 'detectable' sequence length with read-pair analysis is limited by the insert size. Thus, longer sequence insertions that contribute to our genetic makeup are not extensively researched. RESULTS: We present NovelSeq: a computational framework to discover the content and location of long novel sequence insertions using paired-end sequencing data generated by the next-generation sequencing platforms. Our framework can be built as part of a general sequence analysis pipeline to discover multiple types of genetic variation (SNPs, structural variation, etc.), thus it requires significantly less-computational resources than de novo sequence assembly. We apply our methods to detect novel sequence insertions in the genome of an anonymous donor and validate our results by comparing with the insertions discovered in the same genome using various sources of sequence data. AVAILABILITY: The implementation of the NovelSeq pipeline is available at http://compbio.cs.sfu.ca/strvar.htm CONTACT: [email protected]; [email protected]
Iman Hajirasouliha, Fereydoun Hormozdiari, Can Alkan, Jeffrey M. Kidd, Inanç Birol, Evan E. Eichler, Süleyman Cenk Sahinalp
Bioinform.6
2010 Next-generation VariationHunter: combinatorial algorithms for transposon insertion discovery
abstract
UNLABELLED: Recent years have witnessed an increase in research activity for the detection of structural variants (SVs) and their association to human disease. The advent of next-generation sequencing technologies make it possible to extend the scope of structural variation studies to a point previously unimaginable as exemplified by the 1000 Genomes Project. Although various computational methods have been described for the detection of SVs, no such algorithm is yet fully capable of discovering transposon insertions, a very important class of SVs to the study of human evolution and disease. In this article, we provide a complete and novel formulation to discover both loci and classes of transposons inserted into genomes sequenced with high-throughput sequencing technologies. In addition, we also present 'conflict resolution' improvements to our earlier combinatorial SV detection algorithm (VariationHunter) by taking the diploid nature of the human genome into consideration. We test our algorithms with simulated data from the Venter genome (HuRef) and are able to discover >85% of transposon insertion events with precision of >90%. We also demonstrate that our conflict resolution algorithm (denoted as VariationHunter-CR) outperforms current state of the art (such as original VariationHunter, BreakDancer and MoDIL) algorithms when tested on the genome of the Yoruba African individual (NA18507). AVAILABILITY: The implementation of algorithm is available at http://compbio.cs.sfu.ca/strvar.htm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fereydoun Hormozdiari, Iman Hajirasouliha, Phuong Dao, Faraz Hach, Deniz Yörükoglu, Can Alkan, Evan E. Eichler, Süleyman Cenk Sahinalp
Bioinform.7
2010 Targeted interrogation of copy number variation using SCIMMkit
abstract
SUMMARY: Copy number variants (CNVs) contribute substantially to human genomic diversity, and development of accurate and efficient methods for CNV genotyping is a central problem in exploring human genotype-phenotype associations. SCIMMkit provides a robust, integrated implementation of three previously validated algorithms [SCIMM (SNP-Conditional Mixture Modeling), SCIMM-Search and SCOUT (SNP-Conditional OUTlier detection)] for targeted interrogation of CNVs using Illumina Infinium II and GoldenGate SNP assays. SCIMMkit is applicable to standardized genome-wide SNP arrays and customized multiplexed SNP panels, providing economy, efficiency and flexibility in experimental design. AVAILABILITY: Source code and documentation are available for noncommercial use at http://droog.gs.washington.edu/scimmkit.
Troy Zerr, Gregory M. Cooper, Evan E. Eichler, Deborah A. Nickerson
Bioinform.3
2009 Combinatorial Algorithms for Structural Variation Detection in High Throughput Sequenced Genomes
Fereydoun Hormozdiari, Can Alkan, Evan E. Eichler, Süleyman Cenk Sahinalp
RECOMB3
2007 Organization and Evolution of Primate Centromeric DNA from Whole-Genome Shotgun Sequence Data
abstract
The major DNA constituent of primate centromeres is alpha satellite DNA. As much as 2%-5% of sequence generated as part of primate genome sequencing projects consists of this material, which is fragmented or not assembled as part of published genome sequences due to its highly repetitive nature. Here, we develop computational methods to rapidly recover and categorize alpha-satellite sequences from previously uncharacterized whole-genome shotgun sequence data. We present an algorithm to computationally predict potential higher-order array structure based on paired-end sequence data and then experimentally validate its organization and distribution by experimental analyses. Using whole-genome shotgun data from the human, chimpanzee, and macaque genomes, we examine the phylogenetic relationship of these sequences and provide further support for a model for their evolution and mutation over the last 25 million years. Our results confirm fundamental differences in the dispersal and evolution of centromeric satellites in the Old World monkey and ape lineages of evolution.
Can Alkan, Mario Ventura, Nicoletta Archidiacono, Mariano Rocchi, Süleyman Cenk Sahinalp, Evan E. Eichler
PLoS Comput. Biol.6
2002 Statistical Identification of Uniformly Mutated Segments within Repeats
Süleyman Cenk Sahinalp, Evan E. Eichler, Paul W. Goldberg, Petra Berenbrink, Tom Friedetzky, Funda Ergün
CPM2
2002 Recent duplication, evolution and assembly of the human genome
abstract
It has been estimated that 5% of the human genome consists of interspersed duplicated material that has arisen over the last 30 million years of evolution. Two categories of recent duplicated segments can be distinguished: segmental duplications between non-homologous chromosomes (transchromosomal duplications) and duplications largely restricted to a particular chromosome (chromosome-specific duplications). A large proportion of these duplications exhibit an extraordinarily high degree of sequence identity at the nucleotide level (>95%) spanning large (1--100 kb) genomic distances. Through processes of paralogous recombination, these same regions are targets for rapid evolutionary turnover among the genomes of closely related primates. The dynamic nature of these regions in terms of recurrent chromosomal structural rearrangement and their ability to create fusion genes from juxtaposed cassettes suggests that duplicative transposition has been an important force in the evolution of our genome. Cycles of segmental duplication over periods of evolutionary time may provide the underlying mechanism for domain accretion and the increased modular complexity of the vertebrate proteome. Further, our data suggest that a small fraction of important human genes may have emerged recently through duplication processes and will not possess definitive orthologues in the genomes of model organisms. I will discuss computational methods developed in my laboratory to 1) unambiguously identify recent genomic duplicates within the human genome and 2) to assess their importance in hominoid gene innovation. The impact of this chromosomal architecture for assembly of the final draft sequence will be discussed.
Evan E. Eichler
RECOMB1