Mark Diekhans

dblp:62/2692 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0002-0430-0989ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 1 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
10 papers
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 50% Hardware accelerators and domain-specific architectures · 38% Parallel and multicore computing · 12%

Topics — the 23 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
genomics
0.832023
RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023
LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009
The UCSC Known Genes · Bioinform. 2006
Bioinformatics and computational biology › genomics › genomic data management
genomic data sharing
0.712023
RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023
Bioinformatics and computational biology
cancer genomics
0.322013
CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013
CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancer · Bioinform. 2011
Bioinformatics and computational biology › genomics
variant annotation
0.212013
CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013
Bioinformatics and computational biology › statistical genetics
variant prioritization
0.212013
CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013
Bioinformatics and computational biology
comparative genomics
0.222010
Cactus Graphs for Genome Comparisons · RECOMB 2010
Scoring two-species local alignments to try to statistically separate neutrally evolving from selected DNA segments · RECOMB 2003
Bioinformatics and computational biology
genome annotation
0.122008
Using native and syntenically mapped cDNA alignments to improve de novo gene finding · Bioinform. 2008
The UCSC Known Genes · Bioinform. 2006
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis
0.112011
CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancer · Bioinform. 2011
Bioinformatics and computational biology › comparative genomics
genome comparison
0.112010
Cactus Graphs for Genome Comparisons · RECOMB 2010
Bioinformatics and computational biology › genomics › variant annotation
SNP annotation
0.112009
LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009
Bioinformatics and computational biology › genome annotation
gene prediction
0.112008
Using native and syntenically mapped cDNA alignments to improve de novo gene finding · Bioinform. 2008
Bioinformatics and computational biology › transcriptomics › RNA splicing analysis
splice variant prediction
0.112008
Using native and syntenically mapped cDNA alignments to improve de novo gene finding · Bioinform. 2008
Bioinformatics and computational biology › protein function prediction
functional impact prediction
0.112005
LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources · Bioinform. 2005
Bioinformatics and computational biology › genome annotation
genomic variant annotation
0.112005
LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources · Bioinform. 2005
Hardware accelerators and domain-specific architectures
bioinformatics accelerator
0.112005
The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005
Processor architecture and microarchitecture › SIMD
SIMD processor
0.112005
The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005
Bioinformatics and computational biology › comparative genomics › conservation analysis
conserved region identification
0.012003
Scoring two-species local alignments to try to statistically separate neutrally evolving from selected DNA segments · RECOMB 2003
Bioinformatics and computational biology › structural bioinformatics › protein structure representation
protein structure visualization
0.012009
LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009
Bioinformatics and computational biology
structural biology
0.012009
LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection
0.011999
Using the Fisher Kernel Method to Detect Remote Protein Homologies · ISMB 1999
Bioinformatics and computational biology › genomics › genome visualization
genome browser
0.012006
The UCSC Known Genes · Bioinform. 2006
Processor architecture and microarchitecture
SIMD
0.012005
The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005
Parallel and multicore computing › data parallelism
SIMD vectorization
0.012005
The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005

Methods — techniques the papers use, named apart from their topics

matrix slicing · 0.7predictive scoring · 0.2predictive feature database · 0.1syntenic mapping · 0.1evolutionary conservation · 0.1EST alignment · 0.1cross-referencing · 0.1automated pipeline · 0.1protein structure modeling · 0.1performance analysis · 0.1pathway mapping · 0.1architectural design · 0.1
YearPublicationVenuePosition
2023 RNAget: an API to securely retrieve RNA quantifications
abstract
SUMMARY: Large-scale sharing of genomic quantification data requires standardized access interfaces. In this Global Alliance for Genomics and Health project, we developed RNAget, an API for secure access to genomic quantification data in matrix form. RNAget provides for slicing matrices to extract desired subsets of data and is applicable to all expression matrix-format data, including RNA sequencing and microarrays. Further, it generalizes to quantification matrices of other sequence-based genomics such as ATAC-seq and ChIP-seq. AVAILABILITY AND IMPLEMENTATION: https://ga4gh-rnaseq.github.io/schema/docs/index.html.
Sean Upchurch, Emilio Palumbo, Jeremy Adams, David Bujold, Guillaume Bourque, Jared Nedzel, Keenan Graham, Meenakshi S. Kagda, Pedro Assis, Benjamin C. Hitz, Emilio Righi, Roderic Guigó, Barbara J. Wold, Alvis Brazma, Julia Burchard, Joe Capka, Michael Cherry, Laura Clarke, Brian Craft, Manolis Dermitzakis, Mark Diekhans, John Dursi, Michael Sean Fitzsimons, Zac Flaming, Romina Garrido, Alfred Gil, Paul Godden, Matt Green, Mitch Guttman, Brian Haas, Max Haeussler, Sten Linnarsson, Adam Lipski, Simonne Longerich, David R. Lougheed, Jonathan Manning, John C. Marioni, Christopher Meyer, Stephen B. Montgomery, Alyssa Morrow, Alfonso Muñoz-Pomer Fuentes, Jared L. Nedzel, Kevin Osborn, Francis Ouellette, Irene Papatheodorou, Dmitri D. Pervouchine, Arun K. Ramani, Jordi Rambla De Argila, Bashir Sadjad, David Steinberg, Jeremiah Talkar, Timothy Tickle, Kathy Tzeng, Saman Vaisipour, Sean Watford, Barbara Wold
Bioinform.21
2015 The NIH BD2K center for big data in translational genomics
abstract
The world's genomics data will never be stored in a single repository - rather, it will be distributed among many sites in many countries. No one site will have enough data to explain genotype to phenotype relationships in rare diseases; therefore, sites must share data. To accomplish this, the genetics community must forge common standards and protocols to make sharing and computing data among many sites a seamless activity. Through the Global Alliance for Genomics and Health, we are pioneering the development of shared application programming interfaces (APIs) to connect the world's genome repositories. In parallel, we are developing an open source software stack (ADAM) that uses these APIs. This combination will create a cohesive genome informatics ecosystem. Using containers, we are facilitating the deployment of this software in a diverse array of environments. Through benchmarking efforts and big data driver projects, we are ensuring ADAM's performance and utility.
Benedict Paten, Mark Diekhans, Brian J. Druker, Stephen H. Friend, Justin Guinney, Nadine Gassner, Mitchell Guttman, W. James Kent, Patrick Mantey, Adam A. Margolin, Matt Massie, Adam M. Novak, Frank A. Nothaft, Lior Pachter, David A. Patterson 0001, Maciej Smuga-Otto, Joshua M. Stuart, Laura J. van't Veer, Barbara J. Wold, David Haussler
J. Am. Medical Informatics Assoc.2
2013 CRAVAT: cancer-related analysis of variants toolkit
abstract
SUMMARY: Advances in sequencing technology have greatly reduced the costs incurred in collecting raw sequencing data. Academic laboratories and researchers therefore now have access to very large datasets of genomic alterations but limited time and computational resources to analyse their potential biological importance. Here, we provide a web-based application, Cancer-Related Analysis of Variants Toolkit, designed with an easy-to-use interface to facilitate the high-throughput assessment and prioritization of genes and missense alterations important for cancer tumorigenesis. Cancer-Related Analysis of Variants Toolkit provides predictive scores for germline variants, somatic mutations and relative gene importance, as well as annotations from published literature and databases. Results are emailed to users as MS Excel spreadsheets and/or tab-separated text files. AVAILABILITY: http://www.cravat.us/
Christopher Douville, Hannah Carter, Rick Kim, Noushin Niknafs, Mark Diekhans, Peter D. Stenson, David N. Cooper, Michael C. Ryan, Rachel Karchin
Bioinform.5
2011 CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancer
abstract
SUMMARY: Thousands of cancer exomes are currently being sequenced, yielding millions of non-synonymous single nucleotide variants (SNVs) of possible relevance to disease etiology. Here, we provide a software toolkit to prioritize SNVs based on their predicted contribution to tumorigenesis. It includes a database of precomputed, predictive features covering all positions in the annotated human exome and can be used either stand-alone or as part of a larger variant discovery pipeline. AVAILABILITY AND IMPLEMENTATION: MySQL database, source code and binaries freely available for academic/government use at http://wiki.chasmsoftware.org, Source in Python and C++. Requires 32 or 64-bit Linux system (tested on Fedora Core 8,10,11 and Ubuntu 10), 2.5*≤ Python <3.0*, MySQL server >5.0, 60 GB available hard disk space (50 MB for software and data files, 40 GB for MySQL database dump when uncompressed), 2 GB of RAM.
Wing Chung Wong, Dewey Kim, Hannah Carter, Mark Diekhans, Michael C. Ryan, Rachel Karchin
Bioinform.4
2010 Cactus Graphs for Genome Comparisons
Benedict Paten, Mark Diekhans, Dent Earl, John St. John, Jian Ma 0004, Bernard B. Suh, David Haussler
RECOMB2
2009 LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures
abstract
SUMMARY: LS-SNP/PDB is a new WWW resource for genome-wide annotation of human non-synonymous (amino acid changing) SNPs. It serves high-quality protein graphics rendered with UCSF Chimera molecular visualization software. The system is kept up-to-date by an automated, high-throughput build pipeline that systematically maps human nsSNPs onto Protein Data Bank structures and annotates several biologically relevant features. AVAILABILITY: LS-SNP/PDB is available at (http://ls-snp.icm.jhu.edu/ls-snp-pdb) and via links from protein data bank (PDB) biology and chemistry tabs, UCSC Genome Browser Gene Details and SNP Details pages and PharmGKB Gene Variants Downloads/Cross-References pages.
Michael C. Ryan, Mark Diekhans, Stephanie Lien, Yun Liu 0013, Rachel Karchin
Bioinform.2
2008 Using native and syntenically mapped cDNA alignments to improve de novo gene finding
abstract
MOTIVATION: Computational annotation of protein coding genes in genomic DNA is a widely used and essential tool for analyzing newly sequenced genomes. However, current methods suffer from inaccuracy and do poorly with certain types of genes. Including additional sources of evidence of the existence and structure of genes can improve the quality of gene predictions. For many eukaryotic genomes, expressed sequence tags (ESTs) are available as evidence for genes. Related genomes that have been sequenced, annotated, and aligned to the target genome provide evidence of existence and structure of genes. RESULTS: We incorporate several different evidence sources into the gene finder AUGUSTUS. The sources of evidence are gene and transcript annotations from related species syntenically mapped to the target genome using TransMap, evolutionary conservation of DNA, mRNA and ESTs of the target species, and retroposed genes. The predictions include alternative splice variants where evidence supports it. Using only ESTs we were able to correctly predict at least one splice form exactly correct in 57% of human genes. Also using evidence from other species and human mRNAs, this number rises to 77%. Syntenic mapping is well-suited to annotate genomes closely related to genomes that are already annotated or for which extensive transcript evidence is available. Native cDNA evidence is most helpful when the alignments are used as compound information rather than independent positionwise information. AVAILABILITY: AUGUSTUS is open source and available at http://augustus.gobics.de. The gene predictions for human can be browsed and downloaded at the UCSC Genome Browser (http://genome.ucsc.edu).
Mario Stanke, Mark Diekhans, Robert Baertsch, David Haussler
Bioinform.2
2007 Comparative Genomics Search for Losses of Long-Established Genes on the Human Lineage
abstract
Taking advantage of the complete genome sequences of several mammals, we developed a novel method to detect losses of well-established genes in the human genome through syntenic mapping of gene structures between the human, mouse, and dog genomes. Unlike most previous genomic methods for pseudogene identification, this analysis is able to differentiate losses of well-established genes from pseudogenes formed shortly after segmental duplication or generated via retrotransposition. Therefore, it enables us to find genes that were inactivated long after their birth, which were likely to have evolved nonredundant biological functions before being inactivated. The method was used to look for gene losses along the human lineage during the approximately 75 million years (My) since the common ancestor of primates and rodents (the euarchontoglire crown group). We identified 26 losses of well-established genes in the human genome that were all lost at least 50 My after their birth. Many of them were previously characterized pseudogenes in the human genome, such as GULO and UOX. Our methodology is highly effective at identifying losses of single-copy genes of ancient origin, allowing us to find a few well-known pseudogenes in the human genome missed by previous high-throughput genome-wide studies. In addition to confirming previously known gene losses, we identified 16 previously uncharacterized human pseudogenes that are definitive losses of long-established genes. Among them is ACYL3, an ancient enzyme present in archaea, bacteria, and eukaryotes, but lost approximately 6 to 8 Mya in the ancestor of humans and chimps. Although losses of well-established genes do not equate to adaptive gene losses, they are a useful proxy to use when searching for such genetic changes. This is especially true for adaptive losses that occurred more than 250,000 years ago, since any genetic evidence of the selective sweep indicative of such an event has been erased.
Jingchun Zhu, J. Zachary Sanborn, Mark Diekhans, Craig B. Lowe, Tom H. Pringle, David Haussler
PLoS Comput. Biol.3
2006 The UCSC Known Genes
abstract
The University of California Santa Cruz (UCSC) Known Genes dataset is constructed by a fully automated process, based on protein data from Swiss-Prot/TrEMBL (UniProt) and the associated mRNA data from Genbank. The detailed steps of this process are described. Extensive cross-references from this dataset to other genomic and proteomic data were constructed. For each known gene, a details page is provided containing rich information about the gene, together with extensive links to other relevant genomic, proteomic and pathway data. As of July 2005, the UCSC Known Genes are available for human, mouse and rat genomes. The Known Genes serves as a foundation to support several key programs: the Genome Browser, Proteome Browser, Gene Sorter and Table Browser offered at the UCSC website. All the associated data files and program source code are also available. They can be accessed at http://genome.ucsc.edu. The genomic coverage of UCSC Known Genes, RefSeq, Ensembl Genes, H-Invitational and CCDS is analyzed. Although UCSC Known Genes offers the highest genomic and CDS coverage among major human and mouse gene sets, more detailed analysis suggests all of them could be further improved.
Fan Hsu, W. James Kent, Hiram Clawson, Robert M. Kuhn, Mark Diekhans, David Haussler
Bioinform.5
2005 LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources
abstract
MOTIVATION: The NCBI dbSNP database lists over 9 million single nucleotide polymorphisms (SNPs) in the human genome, but currently contains limited annotation information. SNPs that result in amino acid residue changes (nsSNPs) are of critical importance in variation between individuals, including disease and drug sensitivity. RESULTS: We have developed LS-SNP, a genomic scale software pipeline to annotate nsSNPs. LS-SNP comprehensively maps nsSNPs onto protein sequences, functional pathways and comparative protein structure models, and predicts positions where nsSNPs destabilize proteins, interfere with the formation of domain-domain interfaces, have an effect on protein-ligand binding or severely impact human health. It currently annotates 28,043 validated SNPs that produce amino acid residue substitutions in human proteins from the SwissProt/TrEMBL database. Annotations can be viewed via a web interface either in the context of a genomic region or by selecting sets of SNPs, genes, proteins or pathways. These results are useful for identifying candidate functional SNPs within a gene, haplotype or pathway and in probing molecular mechanisms responsible for functional impacts of nsSNPs. AVAILABILITY: http://www.salilab.org/LS-SNP CONTACT: [email protected] SUPPLEMENTARY INFORMATION: http://salilab.org/LS-SNP/supp-info.pdf.
Rachel Karchin, Mark Diekhans, Libusha Kelly, Daryl J. Thomas, Ursula Pieper 0001, Narayanan Eswar, David Haussler, Andrej Sali
Bioinform.2
2005 The UCSC Kestrel Parallel Processor
abstract
The architectural landscape of high-performance computing stretches from superscalar uniprocessor to explicitly parallel systems, to dedicated hardware implementations of algorithms. Single-purpose hardware can achieve the highest performance and uniprocessors can be the most programmable. Between these extremes, programmable and reconfigurable architectures provide a wide range of choice in flexibility, programmability, computational density, and performance. The UCSC Kestrel parallel processor strives to attain single-purpose performance while maintaining user programmability. Kestrel is a single-instruction stream, multiple-data stream (SIMD) parallel processor with a 512-element linear array of 8-bit processing elements. The system design focuses on efficient high-throughput DNA and protein sequence analysis, but its programmability enables high performance on computational chemistry, image processing, machine learning, and other applications. The Kestrel system has had unexpected longevity in its utility due to a careful design and analysis process. Experience with the system leads to the conclusion that programmable SIMD architectures can excel in both programmability and performance. This work presents the architecture, implementation, applications, and observations of the Kestrel project at the University of California at Santa Cruz.
Andrea Di Blas, David M. Dahle, Mark Diekhans, Leslie Grate, Jeffrey D. Hirschberg, Kevin Karplus, Hansjörg Keller, Mark Kendrick, Francisco J. Mesa-Martinez, David Pease, Eric Rice, Angela Schultz, Don Speck, Richard Hughey
IEEE Trans. Parallel Distributed Syst.3
2003 Scoring two-species local alignments to try to statistically separate neutrally evolving from selected DNA segments
abstract
We construct several score functions for use in locating unusually conserved regions in a genome-wide search of aligned DNA from two species. We test these functions on regions of the human genome aligned to the mouse genome. These score functions are derived from properties of neutrally evolving sites on the mouse and human genome, and can be adjusted to the local background rate of conservation. The aim of these functions is to try to identify regions of the human genome that are conserved by evolutionary selection, because they have an important function, rather than by chance. We use them to get a very rough estimate of the amount of DNA in the human genome that is under selection.
Krishna M. Roskin, Mark Diekhans, David Haussler
RECOMB2
1999 Using the Fisher Kernel Method to Detect Remote Protein Homologies
Tommi S. Jaakkola, Mark Diekhans, David Haussler
ISMB2