VLDB 2026 Research / reviewers in the wild / expert
Mark Diekhans
dblp:62/2692
· DBLP profile ↗
13ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0002-0430-0989ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 1 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 50% Hardware accelerators and domain-specific architectures · 38% Parallel and multicore computing · 12% |
Topics — the 23 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
genomics |
0.8 | 3 | 2023 | RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023 LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009 The UCSC Known Genes · Bioinform. 2006 |
Bioinformatics and computational biology › genomics › genomic data management
genomic data sharing |
0.7 | 1 | 2023 | RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023 |
Bioinformatics and computational biology
cancer genomics |
0.3 | 2 | 2013 | CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013 CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancer · Bioinform. 2011 |
Bioinformatics and computational biology › genomics
variant annotation |
0.2 | 1 | 2013 | CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013 |
Bioinformatics and computational biology › statistical genetics
variant prioritization |
0.2 | 1 | 2013 | CRAVAT: cancer-related analysis of variants toolkit · Bioinform. 2013 |
Bioinformatics and computational biology
comparative genomics |
0.2 | 2 | 2010 | Cactus Graphs for Genome Comparisons · RECOMB 2010 Scoring two-species local alignments to try to statistically separate neutrally evolving from selected DNA segments · RECOMB 2003 |
Bioinformatics and computational biology
genome annotation |
0.1 | 2 | 2008 | Using native and syntenically mapped cDNA alignments to improve de novo gene finding · Bioinform. 2008 The UCSC Known Genes · Bioinform. 2006 |
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis |
0.1 | 1 | 2011 | CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancer · Bioinform. 2011 |
Bioinformatics and computational biology › comparative genomics
genome comparison |
0.1 | 1 | 2010 | Cactus Graphs for Genome Comparisons · RECOMB 2010 |
Bioinformatics and computational biology › genomics › variant annotation
SNP annotation |
0.1 | 1 | 2009 | LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009 |
Bioinformatics and computational biology › genome annotation
gene prediction |
0.1 | 1 | 2008 | Using native and syntenically mapped cDNA alignments to improve de novo gene finding · Bioinform. 2008 |
Bioinformatics and computational biology › transcriptomics › RNA splicing analysis
splice variant prediction |
0.1 | 1 | 2008 | Using native and syntenically mapped cDNA alignments to improve de novo gene finding · Bioinform. 2008 |
Bioinformatics and computational biology › protein function prediction
functional impact prediction |
0.1 | 1 | 2005 | LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources · Bioinform. 2005 |
Bioinformatics and computational biology › genome annotation
genomic variant annotation |
0.1 | 1 | 2005 | LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sources · Bioinform. 2005 |
Hardware accelerators and domain-specific architectures
bioinformatics accelerator |
0.1 | 1 | 2005 | The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005 |
Processor architecture and microarchitecture › SIMD
SIMD processor |
0.1 | 1 | 2005 | The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005 |
Bioinformatics and computational biology › comparative genomics › conservation analysis
conserved region identification |
0.0 | 1 | 2003 | Scoring two-species local alignments to try to statistically separate neutrally evolving from selected DNA segments · RECOMB 2003 |
Bioinformatics and computational biology › structural bioinformatics › protein structure representation
protein structure visualization |
0.0 | 1 | 2009 | LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009 |
Bioinformatics and computational biology
structural biology |
0.0 | 1 | 2009 | LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structures · Bioinform. 2009 |
Bioinformatics and computational biology › sequence analysis › homology detection
remote homology detection |
0.0 | 1 | 1999 | Using the Fisher Kernel Method to Detect Remote Protein Homologies · ISMB 1999 |
Bioinformatics and computational biology › genomics › genome visualization
genome browser |
0.0 | 1 | 2006 | The UCSC Known Genes · Bioinform. 2006 |
Processor architecture and microarchitecture
SIMD |
0.0 | 1 | 2005 | The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005 |
Parallel and multicore computing › data parallelism
SIMD vectorization |
0.0 | 1 | 2005 | The UCSC Kestrel Parallel Processor · IEEE Trans. Parallel Distributed Syst. 2005 |
Methods — techniques the papers use, named apart from their topics
matrix slicing · 0.7predictive scoring · 0.2predictive feature database · 0.1syntenic mapping · 0.1evolutionary conservation · 0.1EST alignment · 0.1cross-referencing · 0.1automated pipeline · 0.1protein structure modeling · 0.1performance analysis · 0.1pathway mapping · 0.1architectural design · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | RNAget: an API to securely retrieve RNA quantificationsabstractSUMMARY: Large-scale sharing of genomic quantification data requires standardized access interfaces. In this Global Alliance for Genomics and Health project, we developed RNAget, an API for secure access to genomic quantification data in matrix form. RNAget provides for slicing matrices to extract desired subsets of data and is applicable to all expression matrix-format data, including RNA sequencing and microarrays. Further, it generalizes to quantification matrices of other sequence-based genomics such as ATAC-seq and ChIP-seq. AVAILABILITY AND IMPLEMENTATION: https://ga4gh-rnaseq.github.io/schema/docs/index.html. Sean Upchurch, Emilio Palumbo, Jeremy Adams, David Bujold, Guillaume Bourque, Jared Nedzel, Keenan Graham, Meenakshi S. Kagda, Pedro Assis, Benjamin C. Hitz, Emilio Righi, Roderic Guigó, Barbara J. Wold, Alvis Brazma, Julia Burchard, Joe Capka, Michael Cherry, Laura Clarke, Brian Craft, Manolis Dermitzakis, Mark Diekhans, John Dursi, Michael Sean Fitzsimons, Zac Flaming, Romina Garrido, Alfred Gil, Paul Godden, Matt Green, Mitch Guttman, Brian Haas, Max Haeussler, Sten Linnarsson, Adam Lipski, Simonne Longerich, David R. Lougheed, Jonathan Manning, John C. Marioni, Christopher Meyer, Stephen B. Montgomery, Alyssa Morrow, Alfonso Muñoz-Pomer Fuentes, Jared L. Nedzel, Kevin Osborn, Francis Ouellette, Irene Papatheodorou, Dmitri D. Pervouchine, Arun K. Ramani, Jordi Rambla De Argila, Bashir Sadjad, David Steinberg, Jeremiah Talkar, Timothy Tickle, Kathy Tzeng, Saman Vaisipour, Sean Watford, Barbara Wold |
Bioinform. | 21 |
| 2015 | The NIH BD2K center for big data in translational genomicsabstractThe world's genomics data will never be stored in a single repository - rather, it will be distributed among many sites in many countries. No one site will have enough data to explain genotype to phenotype relationships in rare diseases; therefore, sites must share data. To accomplish this, the genetics community must forge common standards and protocols to make sharing and computing data among many sites a seamless activity. Through the Global Alliance for Genomics and Health, we are pioneering the development of shared application programming interfaces (APIs) to connect the world's genome repositories. In parallel, we are developing an open source software stack (ADAM) that uses these APIs. This combination will create a cohesive genome informatics ecosystem. Using containers, we are facilitating the deployment of this software in a diverse array of environments. Through benchmarking efforts and big data driver projects, we are ensuring ADAM's performance and utility. Benedict Paten, Mark Diekhans, Brian J. Druker, Stephen H. Friend, Justin Guinney, Nadine Gassner, Mitchell Guttman, W. James Kent, Patrick Mantey, Adam A. Margolin, Matt Massie, Adam M. Novak, Frank A. Nothaft, Lior Pachter, David A. Patterson 0001, Maciej Smuga-Otto, Joshua M. Stuart, Laura J. van't Veer, Barbara J. Wold, David Haussler |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | CRAVAT: cancer-related analysis of variants toolkitabstractSUMMARY: Advances in sequencing technology have greatly reduced the costs incurred in collecting raw sequencing data. Academic laboratories and researchers therefore now have access to very large datasets of genomic alterations but limited time and computational resources to analyse their potential biological importance. Here, we provide a web-based application, Cancer-Related Analysis of Variants Toolkit, designed with an easy-to-use interface to facilitate the high-throughput assessment and prioritization of genes and missense alterations important for cancer tumorigenesis. Cancer-Related Analysis of Variants Toolkit provides predictive scores for germline variants, somatic mutations and relative gene importance, as well as annotations from published literature and databases. Results are emailed to users as MS Excel spreadsheets and/or tab-separated text files. AVAILABILITY: http://www.cravat.us/ Christopher Douville, Hannah Carter, Rick Kim, Noushin Niknafs, Mark Diekhans, Peter D. Stenson, David N. Cooper, Michael C. Ryan, Rachel Karchin |
Bioinform. | 5 |
| 2011 | CHASM and SNVBox: toolkit for detecting biologically important single nucleotide mutations in cancerabstractSUMMARY: Thousands of cancer exomes are currently being sequenced, yielding millions of non-synonymous single nucleotide variants (SNVs) of possible relevance to disease etiology. Here, we provide a software toolkit to prioritize SNVs based on their predicted contribution to tumorigenesis. It includes a database of precomputed, predictive features covering all positions in the annotated human exome and can be used either stand-alone or as part of a larger variant discovery pipeline. AVAILABILITY AND IMPLEMENTATION: MySQL database, source code and binaries freely available for academic/government use at http://wiki.chasmsoftware.org, Source in Python and C++. Requires 32 or 64-bit Linux system (tested on Fedora Core 8,10,11 and Ubuntu 10), 2.5*≤ Python <3.0*, MySQL server >5.0, 60 GB available hard disk space (50 MB for software and data files, 40 GB for MySQL database dump when uncompressed), 2 GB of RAM. Wing Chung Wong, Dewey Kim, Hannah Carter, Mark Diekhans, Michael C. Ryan, Rachel Karchin |
Bioinform. | 4 |
| 2010 | Cactus Graphs for Genome Comparisons
Benedict Paten, Mark Diekhans, Dent Earl, John St. John, Jian Ma 0004, Bernard B. Suh, David Haussler |
RECOMB | 2 |
| 2009 | LS-SNP/PDB: annotated non-synonymous SNPs mapped to Protein Data Bank structuresabstractSUMMARY: LS-SNP/PDB is a new WWW resource for genome-wide annotation of human non-synonymous (amino acid changing) SNPs. It serves high-quality protein graphics rendered with UCSF Chimera molecular visualization software. The system is kept up-to-date by an automated, high-throughput build pipeline that systematically maps human nsSNPs onto Protein Data Bank structures and annotates several biologically relevant features. AVAILABILITY: LS-SNP/PDB is available at (http://ls-snp.icm.jhu.edu/ls-snp-pdb) and via links from protein data bank (PDB) biology and chemistry tabs, UCSC Genome Browser Gene Details and SNP Details pages and PharmGKB Gene Variants Downloads/Cross-References pages. Michael C. Ryan, Mark Diekhans, Stephanie Lien, Yun Liu 0013, Rachel Karchin |
Bioinform. | 2 |
| 2008 | Using native and syntenically mapped cDNA alignments to improve de novo gene findingabstractMOTIVATION: Computational annotation of protein coding genes in genomic DNA is a widely used and essential tool for analyzing newly sequenced genomes. However, current methods suffer from inaccuracy and do poorly with certain types of genes. Including additional sources of evidence of the existence and structure of genes can improve the quality of gene predictions. For many eukaryotic genomes, expressed sequence tags (ESTs) are available as evidence for genes. Related genomes that have been sequenced, annotated, and aligned to the target genome provide evidence of existence and structure of genes. RESULTS: We incorporate several different evidence sources into the gene finder AUGUSTUS. The sources of evidence are gene and transcript annotations from related species syntenically mapped to the target genome using TransMap, evolutionary conservation of DNA, mRNA and ESTs of the target species, and retroposed genes. The predictions include alternative splice variants where evidence supports it. Using only ESTs we were able to correctly predict at least one splice form exactly correct in 57% of human genes. Also using evidence from other species and human mRNAs, this number rises to 77%. Syntenic mapping is well-suited to annotate genomes closely related to genomes that are already annotated or for which extensive transcript evidence is available. Native cDNA evidence is most helpful when the alignments are used as compound information rather than independent positionwise information. AVAILABILITY: AUGUSTUS is open source and available at http://augustus.gobics.de. The gene predictions for human can be browsed and downloaded at the UCSC Genome Browser (http://genome.ucsc.edu). Mario Stanke, Mark Diekhans, Robert Baertsch, David Haussler |
Bioinform. | 2 |
| 2007 | Comparative Genomics Search for Losses of Long-Established Genes on the Human LineageabstractTaking advantage of the complete genome sequences of several mammals, we developed a novel method to detect losses of well-established genes in the human genome through syntenic mapping of gene structures between the human, mouse, and dog genomes. Unlike most previous genomic methods for pseudogene identification, this analysis is able to differentiate losses of well-established genes from pseudogenes formed shortly after segmental duplication or generated via retrotransposition. Therefore, it enables us to find genes that were inactivated long after their birth, which were likely to have evolved nonredundant biological functions before being inactivated. The method was used to look for gene losses along the human lineage during the approximately 75 million years (My) since the common ancestor of primates and rodents (the euarchontoglire crown group). We identified 26 losses of well-established genes in the human genome that were all lost at least 50 My after their birth. Many of them were previously characterized pseudogenes in the human genome, such as GULO and UOX. Our methodology is highly effective at identifying losses of single-copy genes of ancient origin, allowing us to find a few well-known pseudogenes in the human genome missed by previous high-throughput genome-wide studies. In addition to confirming previously known gene losses, we identified 16 previously uncharacterized human pseudogenes that are definitive losses of long-established genes. Among them is ACYL3, an ancient enzyme present in archaea, bacteria, and eukaryotes, but lost approximately 6 to 8 Mya in the ancestor of humans and chimps. Although losses of well-established genes do not equate to adaptive gene losses, they are a useful proxy to use when searching for such genetic changes. This is especially true for adaptive losses that occurred more than 250,000 years ago, since any genetic evidence of the selective sweep indicative of such an event has been erased. Jingchun Zhu, J. Zachary Sanborn, Mark Diekhans, Craig B. Lowe, Tom H. Pringle, David Haussler |
PLoS Comput. Biol. | 3 |
| 2006 | The UCSC Known GenesabstractThe University of California Santa Cruz (UCSC) Known Genes dataset is constructed by a fully automated process, based on protein data from Swiss-Prot/TrEMBL (UniProt) and the associated mRNA data from Genbank. The detailed steps of this process are described. Extensive cross-references from this dataset to other genomic and proteomic data were constructed. For each known gene, a details page is provided containing rich information about the gene, together with extensive links to other relevant genomic, proteomic and pathway data. As of July 2005, the UCSC Known Genes are available for human, mouse and rat genomes. The Known Genes serves as a foundation to support several key programs: the Genome Browser, Proteome Browser, Gene Sorter and Table Browser offered at the UCSC website. All the associated data files and program source code are also available. They can be accessed at http://genome.ucsc.edu. The genomic coverage of UCSC Known Genes, RefSeq, Ensembl Genes, H-Invitational and CCDS is analyzed. Although UCSC Known Genes offers the highest genomic and CDS coverage among major human and mouse gene sets, more detailed analysis suggests all of them could be further improved. Fan Hsu, W. James Kent, Hiram Clawson, Robert M. Kuhn, Mark Diekhans, David Haussler |
Bioinform. | 5 |
| 2005 | LS-SNP: large-scale annotation of coding non-synonymous SNPs based on multiple information sourcesabstractMOTIVATION: The NCBI dbSNP database lists over 9 million single nucleotide polymorphisms (SNPs) in the human genome, but currently contains limited annotation information. SNPs that result in amino acid residue changes (nsSNPs) are of critical importance in variation between individuals, including disease and drug sensitivity. RESULTS: We have developed LS-SNP, a genomic scale software pipeline to annotate nsSNPs. LS-SNP comprehensively maps nsSNPs onto protein sequences, functional pathways and comparative protein structure models, and predicts positions where nsSNPs destabilize proteins, interfere with the formation of domain-domain interfaces, have an effect on protein-ligand binding or severely impact human health. It currently annotates 28,043 validated SNPs that produce amino acid residue substitutions in human proteins from the SwissProt/TrEMBL database. Annotations can be viewed via a web interface either in the context of a genomic region or by selecting sets of SNPs, genes, proteins or pathways. These results are useful for identifying candidate functional SNPs within a gene, haplotype or pathway and in probing molecular mechanisms responsible for functional impacts of nsSNPs. AVAILABILITY: http://www.salilab.org/LS-SNP CONTACT: [email protected] SUPPLEMENTARY INFORMATION: http://salilab.org/LS-SNP/supp-info.pdf. Rachel Karchin, Mark Diekhans, Libusha Kelly, Daryl J. Thomas, Ursula Pieper 0001, Narayanan Eswar, David Haussler, Andrej Sali |
Bioinform. | 2 |
| 2005 | The UCSC Kestrel Parallel ProcessorabstractThe architectural landscape of high-performance computing stretches from superscalar uniprocessor to explicitly parallel systems, to dedicated hardware implementations of algorithms. Single-purpose hardware can achieve the highest performance and uniprocessors can be the most programmable. Between these extremes, programmable and reconfigurable architectures provide a wide range of choice in flexibility, programmability, computational density, and performance. The UCSC Kestrel parallel processor strives to attain single-purpose performance while maintaining user programmability. Kestrel is a single-instruction stream, multiple-data stream (SIMD) parallel processor with a 512-element linear array of 8-bit processing elements. The system design focuses on efficient high-throughput DNA and protein sequence analysis, but its programmability enables high performance on computational chemistry, image processing, machine learning, and other applications. The Kestrel system has had unexpected longevity in its utility due to a careful design and analysis process. Experience with the system leads to the conclusion that programmable SIMD architectures can excel in both programmability and performance. This work presents the architecture, implementation, applications, and observations of the Kestrel project at the University of California at Santa Cruz. Andrea Di Blas, David M. Dahle, Mark Diekhans, Leslie Grate, Jeffrey D. Hirschberg, Kevin Karplus, Hansjörg Keller, Mark Kendrick, Francisco J. Mesa-Martinez, David Pease, Eric Rice, Angela Schultz, Don Speck, Richard Hughey |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2003 | Scoring two-species local alignments to try to statistically separate neutrally evolving from selected DNA segmentsabstractWe construct several score functions for use in locating unusually conserved regions in a genome-wide search of aligned DNA from two species. We test these functions on regions of the human genome aligned to the mouse genome. These score functions are derived from properties of neutrally evolving sites on the mouse and human genome, and can be adjusted to the local background rate of conservation. The aim of these functions is to try to identify regions of the human genome that are conserved by evolutionary selection, because they have an important function, rather than by chance. We use them to get a very rough estimate of the amount of DNA in the human genome that is under selection. Krishna M. Roskin, Mark Diekhans, David Haussler |
RECOMB | 2 |
| 1999 | Using the Fisher Kernel Method to Detect Remote Protein Homologies
Tommi S. Jaakkola, Mark Diekhans, David Haussler |
ISMB | 2 |