EDBT 2026 Demo / reviewers in the wild / expert
Jill P. Mesirov
dblp:62/2857
· DBLP profile ↗
27ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-9755-2818ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 3 since 2021Systems, architecture and hardware · 6Artificial intelligence and machine learning · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
19 papers |
Bioinformatics and computational biology · 97% Computational science and engineering · 3% | |
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
High-performance computing · 52% Parallel and multicore computing · 37% Interconnection networks and networks-on-chip · 11% |
Topics — the 30 heaviest of 46, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › genomics
genome visualization |
1.9 | 4 | 2026 | igv-reports: embedding interactive genomic visualizations in HTML reports to aid variant review · Bioinform. 2026 igv.js: an embeddable JavaScript implementation of the Integrative Genomics Viewer (IGV) · Bioinform. 2023 Genome-wide analysis and visualization of copy number with CNVpytor in igv.js · Bioinform. 2024 |
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number variation detection |
0.8 | 1 | 2024 | Genome-wide analysis and visualization of copy number with CNVpytor in igv.js · Bioinform. 2024 |
Bioinformatics and computational biology › population genetics
genetic variation analysis |
0.8 | 1 | 2024 | Genome-wide analysis and visualization of copy number with CNVpytor in igv.js · Bioinform. 2024 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene set enrichment analysis |
0.3 | 3 | 2012 | AraPath: a knowledgebase for pathway analysis in Arabidopsis · Bioinform. 2012 Molecular signatures database (MSigDB) 3.0 · Bioinform. 2011 GSEA-P: a desktop application for Gene Set Enrichment Analysis · Bioinform. 2007 |
Bioinformatics and computational biology › transcriptomics
alternative splicing analysis |
0.2 | 1 | 2015 | Quantitative visualization of alternative exon expression from RNA-seq data · Bioinform. 2015 |
Bioinformatics and computational biology › transcriptomics
isoform expression visualization |
0.2 | 1 | 2015 | Quantitative visualization of alternative exon expression from RNA-seq data · Bioinform. 2015 |
Bioinformatics and computational biology
transcriptomics |
0.2 | 1 | 2015 | Quantitative visualization of alternative exon expression from RNA-seq data · Bioinform. 2015 |
Bioinformatics and computational biology
gene expression analysis |
0.2 | 4 | 2007 | GSEA-P: a desktop application for Gene Set Enrichment Analysis · Bioinform. 2007 Comparative gene marker selection suite · Bioinform. 2006 GeneCluster 2.0: an advanced toolset for bioarray analysis · Bioinform. 2004 |
Bioinformatics and computational biology › systems bioinformatics
pathway analysis |
0.1 | 1 | 2012 | AraPath: a knowledgebase for pathway analysis in Arabidopsis · Bioinform. 2012 |
Bioinformatics and computational biology
genomics |
0.1 | 1 | 2011 | Molecular signatures database (MSigDB) 3.0 · Bioinform. 2011 |
Bioinformatics and computational biology › single-cell analysis
cytometry data analysis |
0.1 | 1 | 2010 | Automated High-Dimensional Flow Cytometric Data Analysis · RECOMB 2010 |
Bioinformatics and computational biology
genome annotation |
0.1 | 2 | 2005 | Improving genome annotations using phylogenetic profile anomaly detection · Bioinform. 2005 GeneCruiser: a web service for the annotation of microarray data · Bioinform. 2005 |
Computational science and engineering
high-dimensional data analysis |
0.1 | 1 | 2010 | Automated High-Dimensional Flow Cytometric Data Analysis · RECOMB 2010 |
Bioinformatics and computational biology › gene expression analysis › gene expression classification
marker gene selection |
0.1 | 2 | 2006 | Comparative gene marker selection suite · Bioinform. 2006 GeneCluster 2.0: an advanced toolset for bioarray analysis · Bioinform. 2004 |
Bioinformatics and computational biology › gene expression analysis
sample classification |
0.1 | 2 | 2004 | GeneCluster 2.0: an advanced toolset for bioarray analysis · Bioinform. 2004 Class prediction and discovery using gene expression data · RECOMB 2000 |
Bioinformatics and computational biology › transcriptomics
RNA-seq analysis |
0.1 | 1 | 2015 | Quantitative visualization of alternative exon expression from RNA-seq data · Bioinform. 2015 |
Bioinformatics and computational biology › sequence alignment
genome alignment |
0.1 | 1 | 2006 | Combo: a whole genome comparative browser · Bioinform. 2006 |
Computational science and engineering
statistical significance testing |
0.1 | 1 | 2006 | Comparative gene marker selection suite · Bioinform. 2006 |
Bioinformatics and computational biology
phylogenetics |
0.1 | 1 | 2005 | Improving genome annotations using phylogenetic profile anomaly detection · Bioinform. 2005 |
Bioinformatics and computational biology › phylogenetics
phylogenetic profiling |
0.1 | 1 | 2005 | Improving genome annotations using phylogenetic profile anomaly detection · Bioinform. 2005 |
Bioinformatics and computational biology › genomics
plant genomics |
0.0 | 1 | 2012 | AraPath: a knowledgebase for pathway analysis in Arabidopsis · Bioinform. 2012 |
Bioinformatics and computational biology › gene expression analysis
class discovery |
0.0 | 1 | 2000 | Class prediction and discovery using gene expression data · RECOMB 2000 |
Bioinformatics and computational biology
comparative genomics |
0.0 | 1 | 2000 | Human and mouse gene structure: comparative analysis and application to exon prediction · RECOMB 2000 |
Bioinformatics and computational biology › genome annotation › gene prediction
exon prediction |
0.0 | 1 | 2000 | Human and mouse gene structure: comparative analysis and application to exon prediction · RECOMB 2000 |
Bioinformatics and computational biology › genome annotation
gene prediction |
0.0 | 1 | 2000 | Human and mouse gene structure: comparative analysis and application to exon prediction · RECOMB 2000 |
Bioinformatics and computational biology › genomics
genome sequencing |
0.0 | 1 | 2000 | Sequencing a genome by walking with clone-end sequences: a mathematical analysis (abstract) · RECOMB 2000 |
Bioinformatics and computational biology
sequence alignment |
0.0 | 1 | 2000 | Human and mouse gene structure: comparative analysis and application to exon prediction · RECOMB 2000 |
Bioinformatics and computational biology › genomics › genome sequencing
whole genome shotgun sequencing |
0.0 | 1 | 2000 | Sequencing a genome by walking with clone-end sequences: a mathematical analysis (abstract) · RECOMB 2000 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.0 | 1 | 2006 | Comparative gene marker selection suite · Bioinform. 2006 |
Bioinformatics and computational biology › data integration
database integration |
0.0 | 1 | 2005 | GeneCruiser: a web service for the annotation of microarray data · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
HTML report generation · 1.0read depth analysis · 0.8b-allele frequency analysis · 0.8javascript · 0.7browser-based rendering · 0.7gene set enrichment analysis · 0.1automated data analysis · 0.1statistical enrichment testing · 0.1statistical significance testing · 0.1phylogenetic profiling · 0.1vortex-in-cell method · 0.0random vortex method · 0.0backpropagation · 0.0verlet neighbor list · 0.0linked list · 0.0coarse-grained cell · 0.0hamiltonian path routing · 0.0communication bandwidth optimization · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | igv-reports: embedding interactive genomic visualizations in HTML reports to aid variant reviewabstractSUMMARY: We present igv-reports, a command-line tool to create standalone HTML pages embedding interactive genomic visualizations of read alignments and associated annotations to support variant inspection workflows. The reports contain all data and code required for visualization of the variant sites, with no dependencies on the input data files. AVAILABILITY AND IMPLEMENTATION: igv-reports is a command-line application written in Python. It is freely available at https://github.com/igvteam/igv-reports under an MIT license. James T. Robinson, Helga Thorvaldsdóttir, Jill P. Mesirov |
Bioinform. | 3 |
| 2024 | Genome-wide analysis and visualization of copy number with CNVpytor in igv.jsabstractSUMMARY: Copy number variation (CNV) and alteration (CNA) analysis is a crucial component in many genomic studies and its applications span from basic research to clinic diagnostics and personalized medicine. CNVpytor is a tool featuring a read depth-based caller and combined read depth and B-allele frequency (BAF) based 2D caller to find CNVs and CNAs. The tool stores processed intermediate data and CNV/CNA calls in a compact HDF5 file-pytor file. Here, we describe a new track in igv.js that utilizes pytor and whole genome variant files as input for on-the-fly read depth and BAF visualization, CNV/CNA calling and analysis. Embedding into HTML pages and Jupiter Notebooks enables convenient remote data access and visualization simplifying interpretation and analysis of omics data. AVAILABILITY AND IMPLEMENTATION: The CNVpytor track is integrated with igv.js and available at https://github.com/igvteam/igv.js. The documentation is available at https://github.com/igvteam/igv.js/wiki/cnvpytor. Usage can be tested in the IGV-Web app at https://igv.org/app and also on https://github.com/abyzovlab/CNVpytor. Arijit Panda, Milovan Suvakov, Helga Thorvaldsdóttir, Jill P. Mesirov, James T. Robinson, Alexej Abyzov |
Bioinform. | 4 |
| 2023 | igv.js: an embeddable JavaScript implementation of the Integrative Genomics Viewer (IGV)abstractSUMMARY: igv.js is an embeddable JavaScript implementation of the Integrative Genomics Viewer (IGV). It can be easily dropped into any web page with a single line of code and has no external dependencies. The viewer runs completely in the web browser, with no backend server and no data pre-processing required. AVAILABILITY AND IMPLEMENTATION: The igv.js JavaScript component can be installed from NPM at https://www.npmjs.com/package/igv. The source code is available at https://github.com/igvteam/igv.js under the MIT open-source license. IGV-Web, the end-user application built around igv.js, is available at https://igv.org/app. The source code is available at https://github.com/igvteam/igv-webapp under the MIT open-source license. SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online. James T. Robinson, Helga Thorvaldsdóttir, Douglass Turner, Jill P. Mesirov |
Bioinform. | 4 |
| 2015 | Quantitative visualization of alternative exon expression from RNA-seq dataabstractAbstract Motivation: Analysis of RNA sequencing (RNA-Seq) data revealed that the vast majority of human genes express multiple mRNA isoforms, produced by alternative pre-mRNA splicing and other mechanisms, and that most alternative isoforms vary in expression between human tissues. As RNA-Seq datasets grow in size, it remains challenging to visualize isoform expression across multiple samples. Results: To help address this problem, we present Sashimi plots, a quantitative visualization of aligned RNA-Seq reads that enables quantitative comparison of exon usage across samples or experimental conditions. Sashimi plots can be made using the Broad Integrated Genome Viewer or with a stand-alone command line program. Availability and implementation: Software code and documentation freely available here: http://miso.readthedocs.org/en/fastmiso/sashimi.html Contact: [email protected], [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Yarden Katz, Eric T. Wang, Jacob Silterra, Schraga Schwartz, Bang Wong, Helga Thorvaldsdóttir, James T. Robinson, Jill P. Mesirov, Edoardo M. Airoldi, Christopher B. Burge |
Bioinform. | 8 |
| 2013 | Integrative Genomics Viewer (IGV): high-performance genomics data visualization and explorationabstractData visualization is an essential component of genomic data analysis. However, the size and diversity of the data sets produced by today's sequencing and array-based profiling methods present major challenges to visualization tools. The Integrative Genomics Viewer (IGV) is a high-performance viewer that efficiently handles large heterogeneous data sets, while providing a smooth and intuitive user experience at all levels of genome resolution. A key characteristic of IGV is its focus on the integrative nature of genomic studies, with support for both array-based and next-generation sequencing data, and the integration of clinical and phenotypic data. Although IGV is often used to view genomic data from public sources, its primary emphasis is to support researchers who wish to visualize and explore their own data sets or those from colleagues. To that end, IGV supports flexible loading of local and remote data sets, and is optimized to provide high-performance data visualization and exploration on standard desktop systems. IGV is freely available for download from http://www.broadinstitute.org/igv, under a GNU LGPL open-source license. Helga Thorvaldsdóttir, James T. Robinson, Jill P. Mesirov |
Briefings Bioinform. | 3 |
| 2012 | AraPath: a knowledgebase for pathway analysis in ArabidopsisabstractUNLABELLED: Studying plants using high-throughput genomics technologies is becoming routine, but interpretation of genome-wide expression data in terms of biological pathways remains a challenge, partly due to the lack of pathway databases. To create a knowledgebase for plant pathway analysis, we collected 1683 lists of differentially expressed genes from 397 gene-expression studies, which constitute a molecular signature database of various genetic and environmental perturbations of Arabidopsis. In addition, we extracted 1909 gene sets from various sources such as Gene Ontology, KEGG, AraCyc, Plant Ontology, predicted target genes of microRNAs and transcription factors, and computational gene clusters defined by meta-analysis. With this knowledgebase, we applied Gene Set Enrichment Analysis to an expression profile of cold acclimation and identified expected functional categories and pathways. Our results suggest that the AraPath database can be used to generate specific, testable hypotheses regarding plant molecular pathways from gene expression data. AVAILABILITY: http://bioinformatics.sdstate.edu/arapath/. Liming Lai, Arthur Liberzon, Jason Hennessey, Gaixin Jiang, Jianli Qi, Jill P. Mesirov, Steven X. Ge |
Bioinform. | 6 |
| 2011 | Molecular signatures database (MSigDB) 3.0abstractMOTIVATION: Well-annotated gene sets representing the universe of the biological processes are critical for meaningful and insightful interpretation of large-scale genomic data. The Molecular Signatures Database (MSigDB) is one of the most widely used repositories of such sets. RESULTS: We report the availability of a new version of the database, MSigDB 3.0, with over 6700 gene sets, a complete revision of the collection of canonical pathways and experimental signatures from publications, enhanced annotations and upgrades to the web site. AVAILABILITY AND IMPLEMENTATION: MSigDB is freely available for non-commercial use at http://www.broadinstitute.org/msigdb. Arthur Liberzon, Aravind Subramanian, Reid Pinchback, Helga Thorvaldsdóttir, Pablo Tamayo, Jill P. Mesirov |
Bioinform. | 6 |
| 2010 | Automated High-Dimensional Flow Cytometric Data Analysis
Saumyadipta Pyne, Xinli Hu, Elizabeth Rossin, Tsung-I Lin, Lisa Maier, Clare Baecher-Allan, Geoffrey J. McLachlan, Pablo Tamayo, David Hafler, Philip L. De Jager, Jill P. Mesirov |
RECOMB | 12 |
| 2008 | ISMB 2008 TorontoabstractISCB) presents the Sixteenth International Conference on Intelligent Systems for Molecular Biology (ISMB 2008), to be held in Toronto, Canada, July 19-23, 2008.Now in the final phases of scheduling selected presentations, demonstrations, and posters, the organizers are preparing what will likely be recognized as the premier conference on computational biology in 2008.ISMB 2008 (http://www.iscb.org/ismb2008/)will follow the road paved by the ISMB/ ECCB 2007 (http://www.iscb.org/ismbeccb2007/) in Vienna in the attempt to specifically encourage increased participation from previously under-represented disciplines of computational biology.This conference will feature the best of the computer and life sciences through a variety of core sessions running in multiple parallel tracks, along with single-tracked Keynote Presentations, posters on display throughout the duration of the conference, and an extensive commercial exposition.The first day (July 18) of the meeting is reserved for two-day Special Interest Group (SIG) and Satellite meetings, the second day (July 19) runs SIGs for the first time in parallel with Tutorials and the Student Council Symposium, and for the first time two SIGs are running in parallel with the main ISMB meeting (July 20-23). Michal Linial, Jill P. Mesirov, B. J. Morrison McKay, Burkhard Rost |
PLoS Comput. Biol. | 2 |
| 2007 | GSEA-P: a desktop application for Gene Set Enrichment AnalysisabstractUNLABELLED: Gene Set Enrichment Analysis (GSEA) is a computational method that assesses whether an a priori defined set of genes shows statistically significant, concordant differences between two biological states. We report the availability of a new version of the Java based software (GSEA-P 2.0) that represents a major improvement on the previous release through the addition of a leading edge analysis component, seamless integration with the Molecular Signature Database (MSigDB) and an embedded browser that allows users to search for gene sets and map them to a variety of microarray platform formats. This functionality makes it possible for users to directly import gene sets from MSigDB for analysis with GSEA. We have also improved the visualizations in GSEA-P 2.0 and added links to a new form of concise gene set annotations called Gene Set Cards. These additions, as well as other improvements suggested by over 3500 users who have downloaded the software over the past year have been incorporated into this new release of the GSEA-P Java desktop program. AVAILABILITY: GSEA-P 2.0 is freely available for academic and commercial users and can be downloaded from http://www.broad.mit.edu/GSEA Aravind Subramanian, Heidi Kuehn, Joshua Gould, Pablo Tamayo, Jill P. Mesirov |
Bioinform. | 5 |
| 2007 | Portraits of breast cancer progressionabstractBACKGROUND: Clustering analysis of microarray data is often criticized for giving ambiguous results because of sensitivity to data perturbation or clustering techniques used. In this paper, we describe a new method based on principal component analysis and ensemble consensus clustering that avoids these problems. RESULTS: We illustrate the method on a public microarray dataset from 36 breast cancer patients of whom 31 were diagnosed with at least two of three pathological stages of disease (atypical ductal hyperplasia (ADH), ductal carcinoma in situ (DCIS) and invasive ductal carcinoma (IDC). Our method identifies an optimum set of genes and divides the samples into stable clusters which correlate with clinical classification into Luminal, Basal-like and Her2+ subtypes. Our analysis reveals a hierarchical portrait of breast cancer progression and identifies genes and pathways for each stage, grade and subtype. An intriguing observation is that the disease phenotype is distinguishable in ADH and progresses along distinct pathways for each subtype. The genetic signature for disease heterogeneity across subtypes is greater than the heterogeneity of progression from DCIS to IDC within a subtype, suggesting that the disease subtypes have distinct progression pathways. Our method identifies six disease subtype and one normal clusters. The first split separates the normal samples from the cancer samples. Next, the cancer cluster splits into low grade (pathological grades 1 and 2) and high grade (pathological grades 2 and 3) while the normal cluster is unchanged. Further, the low grade cluster splits into two subclusters and the high grade cluster into four. The final six disease clusters are mapped into one Luminal A, three Luminal B, one Basal-like and one Her2+. CONCLUSION: We confirm that the cancer phenotype can be identified in early stage because the genes altered in this stage progressively alter further as the disease progresses through DCIS into IDC. We identify six subtypes of disease which have distinct genetic signatures and remain separated in the clustering hierarchy. Our findings suggest that the heterogeneity of disease across subtypes is higher than the heterogeneity of the disease progression within a subtype, indicating that the subtypes are in fact distinct diseases. Gul S. Dalgin, Gabriela Alexe, Daniel Scanfeld, Pablo Tamayo, Jill P. Mesirov, Shridar Ganesan, Charles DeLisi, Gyan Bhanot |
BMC Bioinform. | 5 |
| 2006 | Combo: a whole genome comparative browserabstractSUMMARY: Combo is a comparative genome browser that provides a dynamic view of whole genome alignments along with their associated annotations. Combo provides two different visualization perspectives. The perpendicular (dot plot) view provides a dot plot of genome alignments synchronized with a display of genome annotations along each axis. The parallel view displays two genome annotations horizontally, synchronized through a panel displaying local alignments as trapezoids. Users can zoom to any resolution, from whole chromosomes to individual bases. They can select, highlight and view detailed information from specific alignments and annotations. Combo is an organism agnostic and can import data from a variety of file formats. AVAILABILITY: Combo is integrated as part of the Argo Genome Browser which also provides single-genome browsing and editing capabilities. Argo is written in Java, runs on multiple platforms and is freely available for download at http://www.broad.mit.edu/annotation/argo/. Reinhard Engels, Tamara Yu, Christopher B. Burge, Jill P. Mesirov, Dave DeCaprio, James E. Galagan |
Bioinform. | 4 |
| 2006 | Comparative gene marker selection suiteabstractMOTIVATION: An important step in analyzing expression profiles from microarray data is to identify genes that can discriminate between distinct classes of samples. Many statistical approaches for assigning significance values to genes have been developed. The Comparative Marker Selection suite consists of three modules that allow users to apply and compare different methods of computing significance for each marker gene, a viewer to assess the results, and a tool to create derivative datasets and marker lists based on user-defined significance criteria. AVAILABILITY: The Comparative Marker Selection application suite is freely available as a GenePattern module. The GenePattern analysis environment is freely available at http://www.broad.mit.edu/genepattern. Joshua Gould, Gad Getz, Stefano Monti, Michael Reich, Jill P. Mesirov |
Bioinform. | 5 |
| 2005 | GeneCruiser: a web service for the annotation of microarray dataabstractSUMMARY: GeneCruiser is a web service allowing users to annotate their genomic data by mapping microarray feature identifiers to gene identifiers from databases, such as UniGene, while providing links to web resources, such as the UCSC Genome Browser. It relies on a regularly updated database that retrieves and indexes the mappings between microarray probes and genomic databases. Genes are identified using the Life Sciences Identifier standard. AVAILABILITY: GeneCruiser is freely available in the following forms: Web service and Web application, http://www.genecruiser.org; GenePattern, GeneCruiser access has been integrated into our microarray analysis platform, GenePattern. http://www.genepattern.org. Ted Liefeld, Michael Reich, Joshua Gould, Peili Zhang, Pablo Tamayo, Jill P. Mesirov |
Bioinform. | 6 |
| 2005 | Improving genome annotations using phylogenetic profile anomaly detectionabstractMOTIVATION: A promising strategy for refining genome annotations is to detect features that conflict with known functional or evolutionary relationships between groups of genes. Previous work in this area has been focused on investigating the absence of 'housekeeping' genes or components of well-studied pathways. We have sought to develop a method for improving new annotations that can automatically synthesize and use the information available in a database of other annotated genomes. RESULTS: We show that a probabilistic model of phylogenetic profiles, trained from a database of curated genome annotations, can be used to reliably detect errors in new annotations. We use our method to identify 22 genes that were missed in previously published annotations of prokaryotic genomes. AVAILABILITY: The method was evaluated using MATLAB and open source software referenced in this work. Scripts and datasets are available from the authors upon request. CONTACT: [email protected]. Tarjei S. Mikkelsen, James E. Galagan, Jill P. Mesirov |
Bioinform. | 3 |
| 2004 | GeneCluster 2.0: an advanced toolset for bioarray analysisabstractSUMMARY: GeneCluster 2.0 is a software package for analyzing gene expression and other bioarray data, giving users a variety of methods to build and evaluate class predictors, visualize marker lists, cluster data and validate results. GeneCluster 2.0 greatly expands the data analysis capabilities of GeneCluster 1.0 by adding classification, class discovery and permutation test methods. It includes algorithms for building and testing supervised models using weighted voting and k-nearest neighbor algorithms, a module for systematically finding and evaluating clustering via self-organizing maps, and modules for marker gene selection and heat map visualization that allow users to view and sort samples and genes by many criteria. GeneCluster 2.0 is a stand-alone Java application and runs on any platform that supports the Java Runtime Environment version 1.3.1 or greater. AVAILABILITY: http://www.broad.mit.edu/cancer/software Michael Reich, K. Ohm, Michael Angelo, Pablo Tamayo, Jill P. Mesirov |
Bioinform. | 5 |
| 2003 | Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data
Stefano Monti, Pablo Tamayo, Jill P. Mesirov, Todd R. Golub |
Mach. Learn. | 3 |
| 2000 | Sequencing a genome by walking with clone-end sequences: a mathematical analysis (abstract)abstractOne important approach to sequencing a large genome is (i) to sequence a collection of non-overlapping `seed' chosen from a genomic library of large-insert clones (such as bacterial artificial chromosome (BACs)) and then (ii) to take successive `walking' steps by selecting and sequencing minimally overlapping clones, using information such as clone-end sequences to identify the overlaps. We analyze the strategic issues involved in using this approach. We derive formulas showing how two key factors, the initial density of seed clones and the depth of the genomic library used for walking, affect the cost and time of a sequencing project—that is, the amount of redundant sequencing and the number of steps to cover the vast majority of the genome. We also discuss a variant strategy in which a second genomic library with clones having a somewhat smaller insert size is used to close gaps. This approach can dramatically decrease the amount of redundant sequencing, without affecting the rate at which the genome is covered. Serafim Batzoglou, Bonnie Berger, Jill P. Mesirov, Eric S. Lander |
RECOMB | 3 |
| 2000 | Human and mouse gene structure: comparative analysis and application to exon predictionabstractWe describe a novel analytical approach to gene recognition based on cross-species comparison We first undertook a comparison of orthologous genomic look from human and mouse, studying the extent of similarity in the number, size and sequence of exons and introns We then developed an approach for recognizing genes within such orthologous regions, by first aligning the regions using an iterative global alignment system and then identifying genes based on conservation of exonic features at aligned positions in both species The alignment and gene recognition are performed by new programs called GLASS and ROSETTA, respectively ROSETTA performed well at exact identification of coding exons in 117 orthologous pairs tested. Serafim Batzoglou, Lior Pachter, Jill P. Mesirov, Bonnie Berger, Eric S. Lander |
RECOMB | 3 |
| 2000 | Class prediction and discovery using gene expression dataabstractClassification of patient samples is a crucial aspect of cancer diagnosis and treatment. We present a method for classifying samples by computational analysis of gene expression data. We consider the classification problem in two parts: class discovery and class prediction. Class discovery refers to the process of dividing samples into reproducible classes that have similar behavior or properties, while class prediction places new samples into already known classes. We describe a method for performing class prediction and illustrate its strength by correctly classifying bone marrow and blood samples from acute leukemia patients. We also describe how to use our predictor to validate newly discovered classes, and we demonstrate how this technique could have discovered the key distinctions among leukemias if they were not already known. This proof-of-concept experiment paves the way for a wealth of future work on the molecular classification and understanding of disease. Donna K. Slonim, Pablo Tamayo, Jill P. Mesirov, Todd R. Golub, Eric S. Lander |
RECOMB | 3 |
| 1991 | Computing turbulent flow in complex geometries on a massively parallel processorabstractIn this paper, we present parallel implementations of two methods for computing turbulent flow in complex geometries.Both methods are based on the random vortex method, which is particularly suited for computing complex, viscous, incompressible flow across a wide range of flow regimes and characteristics.The lirst method is a full vortex method, designed to accurately simulate such fluid phenomenon as vortex shedding, merger, and rollup, as well quantitative features of the flow.The second method, based on a "vortexin-cell" method, is an extremely fast version which can offer qualitative portrayal of the dominant fluid structures and mechanisms useful in the design stage.Both methods are non-standard, containing few of the positive attributes commonly associated with methods that easily lend themselves to massively parallel implementations.They are Lagrangian schemes, in which the position of each computational element is affeeted by all others at each time step.The efficient execution of these methods on a Connection Machine CM-2 requires parallel N-body solvers, parallel elliptic solvers, and pamllel data structures for the adaptive creation of computational elements on the boundary of the confining region.provide timing runs.Both are connected to the Realtime Interactive Visualization Environment described in [10]. James A. Sethian, Jean-Philippe Brunet, Adam Greenberg, Jill P. Mesirov |
SC | 4 |
| 1991 | Parallel approaches to short range molecular dynamics simulationsabstractWe discuss different approaches to the short -range variab le-neighbor Molecular Dynamlcs problem on parallel mach ines and in particular the Connection Machzne CM-2. Tradltwnal Molecular Dynamics methods are reviewed and the computational require ments of parallel algorithms are analyzed. Three ap proaches based on parallel extensions of lznked lists, Verlet neighbor lzsts and coarse-grained cells are p resented and their advantages and disadvantages dis cussed. Performance evaluation and comparisons with other algorithms and machines are provided. Pablo Tamayo, Jill P. Mesirov, Bruce M. Boghosian |
SC | 2 |
| 1990 | An optional hypercube direct N-body solver on the connection machineabstractThe authors have designed and implemented a hypercube algorithm for direct N-body solvers on the Connection Machine CM-2. The algorithm is optimal in the sense that, as long as there is sufficient data, it uses the full communication bandwidth of a hypercube of any dimension. When the number of bodies per node is large enough, the communication time for the implementation is negligible, i.e., less than 2%. In particular, this means that one obtains close to optimal speedup in the regime. To obtain this performance, 'rotated and translated Gray codes' which result in time-wise edge disjoint Hamiltonian paths on the hypercube are used. Timings are presented for a collection of interacting point vortices in two dimensions. The computation of the velocities of 14,000 vortices in 32-bit precision takes 2 seconds on a 16K CM-2.> Jean-Philippe Brunet, Alan Edelman, Jill P. Mesirov |
SC | 3 |
| 1990 | The backpropagation algorithm on grid and hypercube architectures
Xiru Zhang, Michael McKenna, Jill P. Mesirov, David L. Waltz |
Parallel Comput. | 3 |
| 1989 | An Efficient Implementation of the Back-propagation Algorithm on the Connection Machine CM-2
Xiru Zhang, Michael McKenna, Jill P. Mesirov, David L. Waltz |
NIPS | 3 |
| 1989 | Protein structure prediction by a data-level parallel algorithmabstractWe have developed a software system, PHI-PSI, on the Connection Machine that uses a parallel algorithm to retrieve and use information from a database of 112 known protein structures (selected from the Brookhaven Protein Databank) to predict the structures of other proteins. The φ and ψ angles of each amino acid (the angles each amino acid forms with its immediate neighbors) in a protein are used to represent its 3-D structure. PHI-PSI's algorithm is based on the idea of Memory-based reasoning (MBR) [10] and extends it to include a recursive procedure to refine its initial prediction and a “window” of varying sizes to look at different contexts of an input. PHI-PSI has been tested with all the available data. Initial results show that it performs better than distribution-based guesses for most of the φ and ψ angle values. Xiru Zhang, David L. Waltz, Jill P. Mesirov |
SC | 3 |
| 1989 | Study of protein sequence comparison metrics on the connection machine CM-2
Eric S. Lander, Jill P. Mesirov, Washington Taylor |
J. Supercomput. | 2 |