Lincoln Stein

dblp:s/LincolnStein · also Lincoln D. Stein · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0002-1983-4588ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 22 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2024 FAIR Header Reference genome: a TRUSTworthy standard
abstract
The lack of interoperable data standards among reference genome data-sharing platforms inhibits cross-platform analysis while increasing the risk of data provenance loss. Here, we describe the FAIR bioHeaders Reference genome (FHR), a metadata standard guided by the principles of Findability, Accessibility, Interoperability and Reuse (FAIR) in addition to the principles of Transparency, Responsibility, User focus, Sustainability and Technology. The objective of FHR is to provide an extensive set of data serialisation methods and minimum data field requirements while still maintaining extensibility, flexibility and expressivity in an increasingly decentralised genomic data ecosystem. The effort needed to implement FHR is low; FHR's design philosophy ensures easy implementation while retaining the benefits gained from recording both machine and human-readable provenance.
Adam Wright, Mark D. Wilkinson, Chris Mungall, Scott Cain, Stephen Richards, Paul W. Sternberg, Ellen Provin, Jonathan L. Jacobs, Scott Geib, Daniela Raciti, Karen Yook, Lincoln Stein, David C. Molik
Briefings Bioinform.12
2023 JBrowse Jupyter: a Python interface to JBrowse 2
abstract
MOTIVATION: JBrowse Jupyter is a package that aims to close the gap between Python programming and genomic visualization. Web-based genome browsers are routinely used for publishing and inspecting genome annotations. Historically they have been deployed at the end of bioinformatics pipelines, typically decoupled from the analysis itself. However, emerging technologies such as Jupyter notebooks enable a more rapid iterative cycle of development, analysis and visualization. RESULTS: We have developed a package that provides a Python interface to JBrowse 2's suite of embeddable components, including the primary Linear Genome View. The package enables users to quickly set up, launch and customize JBrowse views from Jupyter notebooks. In addition, users can share their data via Google's Colab notebooks, providing reproducible interactive views. AVAILABILITY AND IMPLEMENTATION: JBrowse Jupyter is released under the Apache License and is available for download on PyPI. Source code and demos are available on GitHub at https://github.com/GMOD/jbrowse-jupyter.
Teresa De Jesus Martinez, Elliot A. Hershberg, Emma Guo, Garrett Stevens, Colin M. Diesh, Peter Xie, Caroline Bridge, Scott Cain, Robin Haw, Robert M. Buels, Lincoln Stein, Ian H. Holmes
Bioinform.11
2023 Fourteen quick tips for crowdsourcing geographically linked data for public health advocacy
abstract
This article presents 14 quick tips to build a team to crowdsource data for public health advocacy. It includes tips around team building and logistics, infrastructure setup, media and industry outreach, and project wrap-up and archival for posterity.
Joshua Atienza, Anjalee Benedict, Lincoln Stein, Kashif Pirzada, Cheryl White, Shraddha Pai
PLoS Comput. Biol.3
2021 JBrowseR: an R interface to the JBrowse 2 genome browser
abstract
MOTIVATION: Genome browsers are an essential tool in genome analysis. Modern genome browsers enable complex and interactive visualization of a wide variety of genomic data modalities. While such browsers are very powerful, they can be challenging to configure and program for bioinformaticians lacking expertise in web development. RESULTS: We have developed an R package that provides an interface to the JBrowse 2 genome browser. The package can be used to configure and customize the browser entirely with R code. The browser can be deployed from the R console, or embedded in Shiny applications or R Markdown documents. AVAILABILITY AND IMPLEMENTATION: JBrowseR is available for download from CRAN, and the source code is openly available from the Github repository at https://github.com/GMOD/JBrowseR/.
Elliot A. Hershberg, Garrett Stevens, Colin M. Diesh, Peter Xie, Teresa De Jesus Martinez, Robert M. Buels, Lincoln Stein, Ian H. Holmes
Bioinform.7
2020 JBrowse Connect: A server API to connect JBrowse instances and users
abstract
We describe JBrowse Connect, an optional expansion to the JBrowse genome browser, targeted at developers. JBrowse Connect allows live messaging, notifications for new annotation tracks, heavy-duty analyses initiated by the user from within the browser, and other dynamic features. We present example applications of JBrowse Connect that allow users 1) to specify and execute BLAST searches by either running on the same host as the webserver, with a self-contained BLAST module leveraging NCBI Blast+ commands, or via a managed Galaxy instance that can optionally run on a different host, and 2) to run the primer design service Primer3. JBrowse Connect allows users to track job progress and view results in the context of the browser. The software is available under a choice of open source licenses including LGPL and the Artistic License.
Eric Yao, Robert M. Buels, Lincoln Stein, Taner Z. Sen, Ian H. Holmes
PLoS Comput. Biol.3
2019 Whole genomes define concordance of matched primary, xenograft, and organoid models of pancreas cancer
abstract
Pancreatic ductal adenocarcinoma (PDAC) has the worst prognosis among solid malignancies and improved therapeutic strategies are needed to improve outcomes. Patient-derived xenografts (PDX) and patient-derived organoids (PDO) serve as promising tools to identify new drugs with therapeutic potential in PDAC. For these preclinical disease models to be effective, they should both recapitulate the molecular heterogeneity of PDAC and validate patient-specific therapeutic sensitivities. To date however, deep characterization of the molecular heterogeneity of PDAC PDX and PDO models and comparison with matched human tumour remains largely unaddressed at the whole genome level. We conducted a comprehensive assessment of the genetic landscape of 16 whole-genome pairs of tumours and matched PDX, from primary PDAC and liver metastasis, including a unique cohort of 5 'trios' of matched primary tumour, PDX, and PDO. We developed a pipeline to score concordance between PDAC models and their paired human tumours for genomic events, including mutations, structural variations, and copy number variations. Tumour-model comparisons of mutations displayed single-gene concordance across major PDAC driver genes, but relatively poor agreement across the greater mutational load. Genome-wide and chromosome-centric analysis of structural variation (SV) events highlights previously unrecognized concordance across chromosomes that demonstrate clustered SV events. We found that polyploidy presented a major challenge when assessing copy number changes; however, ploidy-corrected copy number states suggest good agreement between donor-model pairs. Collectively, our investigations highlight that while PDXs and PDOs may serve as tractable and transplantable systems for probing the molecular properties of PDAC, these models may best serve selective analyses across different levels of genomic complexity.
Deena M. A. Gendoo, Robert E. Denroche, Nikolina Radulovich, Gun Ho Jang, Mathieu Lemire, Sandra Fischer, Dianne Chadwick, Ilinca M. Lungu, Emin Ibrahimov, Ping-Jiang Cao, Lincoln Stein, Julie M. Wilson, John M. S. Bartlett, Ming-Sound Tsao, Neesha Dhani, David Hedley, Steven Gallinger, Benjamin Haibe-Kains
PLoS Comput. Biol.12
2018 Reactome diagram viewer: data structures and strategies to boost performance
abstract
Motivation: Reactome is a free, open-source, open-data, curated and peer-reviewed knowledgebase of biomolecular pathways. For web-based pathway visualization, Reactome uses a custom pathway diagram viewer that has been evolved over the past years. Here, we present comprehensive enhancements in usability and performance based on extensive usability testing sessions and technology developments, aiming to optimize the viewer towards the needs of the community. Results: The pathway diagram viewer version 3 achieves consistently better performance, loading and rendering of 97% of the diagrams in Reactome in less than 1 s. Combining the multi-layer html5 canvas strategy with a space partitioning data structure minimizes CPU workload, enabling the introduction of new features that further enhance user experience. Through the use of highly optimized data structures and algorithms, Reactome has boosted the performance and usability of the new pathway diagram viewer, providing a robust, scalable and easy-to-integrate solution to pathway visualization. As graph-based visualization of complex data is a frequent challenge in bioinformatics, many of the individual strategies presented here are applicable to a wide range of web-based bioinformatics resources. Availability and implementation: Reactome is available online at: https://reactome.org. The diagram viewer is part of the Reactome pathway browser (https://reactome.org/PathwayBrowser/) and also available as a stand-alone widget at: https://reactome.org/dev/diagram/. The source code is freely available at: https://github.com/reactome-pwp/diagram. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Antonio Fabregat, Konstantinos Sidiropoulos, Guilherme Viteri, Pablo Marín-García, Peipei Ping, Lincoln Stein, Peter D'Eustachio, Henning Hermjakob
Bioinform.6
2018 Reactome graph database: Efficient access to complex pathway data
abstract
Reactome is a free, open-source, open-data, curated and peer-reviewed knowledgebase of biomolecular pathways. One of its main priorities is to provide easy and efficient access to its high quality curated data. At present, biological pathway databases typically store their contents in relational databases. This limits access efficiency because there are performance issues associated with queries traversing highly interconnected data. The same data in a graph database can be queried more efficiently. Here we present the rationale behind the adoption of a graph database (Neo4j) as well as the new ContentService (REST API) that provides access to these data. The Neo4j graph database and its query language, Cypher, provide efficient access to the complex Reactome data model, facilitating easy traversal and knowledge discovery. The adoption of this technology greatly improved query efficiency, reducing the average query time by 93%. The web service built on top of the graph database provides programmatic access to Reactome data by object oriented queries, but also supports more complex queries that take advantage of the new underlying graph-based data storage. By adopting graph database technology we are providing a high performance pathway data resource to the community. The Reactome graph database use case shows the power of NoSQL database engines for complex biological data types.
Antonio Fabregat, Florian Korninger, Guilherme Viteri, Konstantinos Sidiropoulos, Pablo Marín-García, Peipei Ping, Guanming Wu, Lincoln Stein, Peter D'Eustachio, Henning Hermjakob
PLoS Comput. Biol.8
2017 Reactome enhanced pathway visualization
abstract
MOTIVATION: Reactome is a free, open-source, open-data, curated and peer-reviewed knowledge base of biomolecular pathways. Pathways are arranged in a hierarchical structure that largely corresponds to the GO biological process hierarchy, allowing the user to navigate from high level concepts like immune system to detailed pathway diagrams showing biomolecular events like membrane transport or phosphorylation. Here, we present new developments in the Reactome visualization system that facilitate navigation through the pathway hierarchy and enable efficient reuse of Reactome visualizations for users' own research presentations and publications. RESULTS: For the higher levels of the hierarchy, Reactome now provides scalable, interactive textbook-style diagrams in SVG format, which are also freely downloadable and editable. Repeated diagram elements like 'mitochondrion' or 'receptor' are available as a library of graphic elements. Detailed lower-level diagrams are now downloadable in editable PPTX format as sets of interconnected objects. AVAILABILITY AND IMPLEMENTATION: http://reactome.org. CONTACT: [email protected] or [email protected].
Konstantinos Sidiropoulos, Guilherme Viteri, Cristoffer Sevilla, Steven Jupe, Marissa Webber, Marija Orlic-Milacic, Bijay Jassal, Bruce May, Veronica Shamovsky, Corina Duenas, Karen Rothfels, Lisa Matthews, Heeyeon Song, Lincoln Stein, Robin Haw, Peter D'Eustachio, Peipei Ping, Henning Hermjakob, Antonio Fabregat
Bioinform.14
2017 Reactome pathway analysis: a high-performance in-memory approach
abstract
BACKGROUND: Reactome aims to provide bioinformatics tools for visualisation, interpretation and analysis of pathway knowledge to support basic research, genome analysis, modelling, systems biology and education. Pathway analysis methods have a broad range of applications in physiological and biomedical research; one of the main problems, from the analysis methods performance point of view, is the constantly increasing size of the data samples. RESULTS: Here, we present a new high-performance in-memory implementation of the well-established over-representation analysis method. To achieve the target, the over-representation analysis method is divided in four different steps and, for each of them, specific data structures are used to improve performance and minimise the memory footprint. The first step, finding out whether an identifier in the user's sample corresponds to an entity in Reactome, is addressed using a radix tree as a lookup table. The second step, modelling the proteins, chemicals, their orthologous in other species and their composition in complexes and sets, is addressed with a graph. The third and fourth steps, that aggregate the results and calculate the statistics, are solved with a double-linked tree. CONCLUSION: Through the use of highly optimised, in-memory data structures and algorithms, Reactome has achieved a stable, high performance pathway analysis service, enabling the analysis of genome-wide datasets within seconds, allowing interactive exploration and analysis of high throughput data. The proposed pathway analysis approach is available in the Reactome production web site either via the AnalysisService for programmatic access or the user submission interface integrated into the PathwayBrowser. Reactome is an open data and open source project and all of its source code, including the one described here, is available in the AnalysisTools repository in the Reactome GitHub ( https://github.com/reactome/ ).
Antonio Fabregat, Konstantinos Sidiropoulos, Guilherme Viteri, Oscar Forner-Martinez, Pablo Marín-García, Vicente Arnau, Peter D'Eustachio, Lincoln Stein, Henning Hermjakob
BMC Bioinform.8
2015 A genome-wide association study platform built on iPlant cyber-infrastructure
abstract
Summary We demonstrate a flexible genome‐wide association study platform built upon the iPlant Collaborative Cyber‐infrastructure. The platform supports big data management, sharing, and large‐scale study of both genotype and phenotype data on clusters. End users can add their own analysis tools and create customized analysis workflows through the graphical user interfaces in both iPlant Discovery Environment and BioExtract server. Copyright © 2014 John Wiley & Sons, Ltd.
Doreen Ware, Carol Lushbough, Nirav C. Merchant, Lincoln Stein
Concurr. Comput. Pract. Exp.5
2014 Inferring clonal evolution of tumors from single nucleotide somatic mutations
abstract
BACKGROUND: High-throughput sequencing allows the detection and quantification of frequencies of somatic single nucleotide variants (SNV) in heterogeneous tumor cell populations. In some cases, the evolutionary history and population frequency of the subclonal lineages of tumor cells present in the sample can be reconstructed from these SNV frequency measurements. But automated methods to do this reconstruction are not available and the conditions under which reconstruction is possible have not been described. RESULTS: We describe the conditions under which the evolutionary history can be uniquely reconstructed from SNV frequencies from single or multiple samples from the tumor population and we introduce a new statistical model, PhyloSub, that infers the phylogeny and genotype of the major subclonal lineages represented in the population of cancer cells. It uses a Bayesian nonparametric prior over trees that groups SNVs into major subclonal lineages and automatically estimates the number of lineages and their ancestry. We sample from the joint posterior distribution over trees to identify evolutionary histories and cell population frequencies that have the highest probability of generating the observed SNV frequency data. When multiple phylogenies are consistent with a given set of SNV frequencies, PhyloSub represents the uncertainty in the tumor phylogeny using a "partial order plot". Experiments on a simulated dataset and two real datasets comprising tumor samples from acute myeloid leukemia and chronic lymphocytic leukemia patients demonstrate that PhyloSub can infer both linear (or chain) and branching lineages and its inferences are in good agreement with ground truth, where it is available. CONCLUSIONS: PhyloSub can be applied to frequencies of any "binary" somatic mutation, including SNVs as well as small insertions and deletions. The PhyloSub and partial order plot software is available from https://github.com/morrislab/phylosub/.
Wei Jiao, Shankar Vembu, Amit G. Deshwar, Lincoln Stein, Quaid Morris
BMC Bioinform.4
2013 Using GBrowse 2.0 to visualize and share next-generation sequence data
abstract
GBrowse is a mature web-based genome browser that is suitable for deployment on both public and private web sites. It supports most of genome browser features, including qualitative and quantitative (wiggle) tracks, track uploading, track sharing, interactive track configuration, semantic zooming and limited smooth track panning. As of version 2.0, GBrowse supports next-generation sequencing (NGS) data by providing for the direct display of SAM and BAM sequence alignment files. SAM/BAM tracks provide semantic zooming and support both local and remote data sources. This article provides step-by-step instructions for configuring GBrowse to display NGS data.
Lincoln Stein
Briefings Bioinform.1
2011 PeakRanger: A cloud-enabled peak caller for ChIP-seq data
abstract
BACKGROUND: Chromatin immunoprecipitation (ChIP), coupled with massively parallel short-read sequencing (seq) is used to probe chromatin dynamics. Although there are many algorithms to call peaks from ChIP-seq datasets, most are tuned either to handle punctate sites, such as transcriptional factor binding sites, or broad regions, such as histone modification marks; few can do both. Other algorithms are limited in their configurability, performance on large data sets, and ability to distinguish closely-spaced peaks. RESULTS: In this paper, we introduce PeakRanger, a peak caller software package that works equally well on punctate and broad sites, can resolve closely-spaced peaks, has excellent performance, and is easily customized. In addition, PeakRanger can be run in a parallel cloud computing environment to obtain extremely high performance on very large data sets. We present a series of benchmarks to evaluate PeakRanger against 10 other peak callers, and demonstrate the performance of PeakRanger on both real and synthetic data sets. We also present real world usages of PeakRanger, including peak-calling in the modENCODE project. CONCLUSIONS: Compared to other peak callers tested, PeakRanger offers improved resolution in distinguishing extremely closely-spaced peaks. PeakRanger has above-average spatial accuracy in terms of identifying the precise location of binding events. PeakRanger also has excellent sensitivity and specificity in all benchmarks evaluated. In addition, PeakRanger offers significant improvements in run time when running on a single processor system, and very marked improvements when allowed to take advantage of the MapReduce parallel environment offered by a cloud computing resource. PeakRanger can be downloaded at the official site of modENCODE project: http://www.modencode.org/software/ranger/
Robert Grossman, Lincoln Stein
BMC Bioinform.3
2010 Localizing triplet periodicity in DNA and cDNA sequences
abstract
BACKGROUND: The protein-coding regions (coding exons) of a DNA sequence exhibit a triplet periodicity (TP) due to fact that coding exons contain a series of three nucleotide codons that encode specific amino acid residues. Such periodicity is usually not observed in introns and intergenic regions. If a DNA sequence is divided into small segments and a Fourier Transform is applied on each segment, a strong peak at frequency 1/3 is typically observed in the Fourier spectrum of coding segments, but not in non-coding regions. This property has been used in identifying the locations of protein-coding genes in unannotated sequence. The method is fast and requires no training. However, the need to compute the Fourier Transform across a segment (window) of arbitrary size affects the accuracy with which one can localize TP boundaries. Here, we report a technique that provides higher-resolution identification of these boundaries, and use the technique to explore the biological correlates of TP regions in the genome of the model organism C. elegans. RESULTS: Using both simulated TP signals and the real C. elegans sequence F56F11 as an example, we demonstrate that, (1) Modified Wavelet Transform (MWT) can better define the boundary of TP region than the conventional Short Time Fourier Transform (STFT); (2) The scale parameter (a) of MWT determines the precision of TP boundary localization: bigger values of a give sharper TP boundaries but result in a lower signal to noise ratio; (3) RNA splicing sites have weaker TP signals than coding region; (4) TP signals in coding region can be destroyed or recovered by frame-shift mutations; (5) 6 bp periodicities in introns and intergenic region can generate false positive signals and it can be removed with 6 bp MWT. CONCLUSIONS: MWT can provide more precise TP boundaries than STFT and the boundaries can be further refined by bigger scale MWT. Subtraction of 6 bp periodicity signals reduces the number of false positives. Experimentally-introduced frame-shift mutations help recover TP signal that have been lost by possible ancient frame-shifts. More importantly, TP signal has the potential to be used to detect the splice junctions in fully spliced mRNA sequence.
Lincoln Stein
BMC Bioinform.2
2009 CMap 1.01: a comparative mapping application for the Internet
abstract
UNLABELLED: CMap is a web-based tool for displaying and comparing maps of any type and from any species. A user can compare an unlimited number of maps, view pair-wise comparisons of known correspondences, and search for maps or for features by name, species, type and accession. CMap is freely available, can run on a variety of database engines and uses only free and open software components. AVAILABILITY: http://www.gmod.org/cmap
Ken Youens-Clark, Benjamin Faga, Immanuel Yap, Lincoln Stein, Doreen Ware
Bioinform.4
2008 nGASP - the nematode genome annotation assessment project
abstract
BACKGROUND: While the C. elegans genome is extensively annotated, relatively little information is available for other Caenorhabditis species. The nematode genome annotation assessment project (nGASP) was launched to objectively assess the accuracy of protein-coding gene prediction software in C. elegans, and to apply this knowledge to the annotation of the genomes of four additional Caenorhabditis species and other nematodes. Seventeen groups worldwide participated in nGASP, and submitted 47 prediction sets across 10 Mb of the C. elegans genome. Predictions were compared to reference gene sets consisting of confirmed or manually curated gene models from WormBase. RESULTS: The most accurate gene-finders were 'combiner' algorithms, which made use of transcript- and protein-alignments and multi-genome alignments, as well as gene predictions from other gene-finders. Gene-finders that used alignments of ESTs, mRNAs and proteins came in second. There was a tie for third place between gene-finders that used multi-genome alignments and ab initio gene-finders. The median gene level sensitivity of combiners was 78% and their specificity was 42%, which is nearly the same accuracy reported for combiners in the human genome. C. elegans genes with exons of unusual hexamer content, as well as those with unusually many exons, short exons, long introns, a weak translation start signal, weak splice sites, or poorly conserved orthologs posed the greatest difficulty for gene-finders. CONCLUSION: This experiment establishes a baseline of gene prediction accuracy in Caenorhabditis genomes, and has guided the choice of gene-finders for the annotation of newly sequenced genomes of Caenorhabditis and other nematode species. We have created new gene sets for C. briggsae, C. remanei, C. brenneri, C. japonica, and Brugia malayi using some of the best-performing gene-finders.
Avril Coghlan, Tristan J. Fiedler, Sheldon J. McKay, Paul Flicek, Todd W. Harris, Darin Blasiar, Lincoln Stein
BMC Bioinform.7
2006 Look-Align: an interactive web-based multiple sequence alignment viewer with polymorphism analysis support
abstract
UNLABELLED: We have developed Look-Align, an interactive web-based viewer to display pre-computed multiple sequence alignments. Although initially developed to support the visualization needs of the maize diversity website Panzea (http://www.panzea.org), the viewer is a generic stand-alone tool that can be easily integrated into other websites. AVAILABILITY: Look-Align is written in Perl using open-source components and is available under an open-source license. Live installation and download information can be found at the Panzea website (http://www.panzea.org/software/alignment_viewer.html). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: The Supplementary information includes sample lists of multiple sequence alignment software and sample screenshots of the viewer.
Payan Canaran, Lincoln Stein, Doreen Ware
Bioinform.2
2005 SynBrowse: a synteny browser for comparative sequence analysis
abstract
MOTIVATION: The recent efforts of various sequence projects to sequence deeply into various phylogenies provide great resources for comparative sequence analysis. A generic and portable tool is essential for scientists to visualize and analyze sequence comparisons. RESULTS: We have developed SynBrowse, a synteny browser for visualizing and analyzing genome alignments both within and between species. It is intended to help scientists study macrosynteny, microsynteny and homologous genes between sequences. It can also aid with the identification of uncharacterized genes, putative regulatory elements and novel structural features of a species. SynBrowse is a GBrowse (the Generic Genome Browser) family software tool that runs on top of the open source BioPerl modules. It consists of two components: a web-based front end and a set of relational database back ends. Each database stores pre-computed alignments from a focus sequence to reference sequences in addition to the genome annotations of the focus sequence. The user interface lets end users select a key comparative alignment type and search for syntenic blocks between two sequences and zoom in to view the relationships among the corresponding genome annotations in detail. SynBrowse is portable with simple installation, flexible configuration, convenient data input and easy integration with other components of a model organism system. AVAILABILITY: The software is available at http://www.gmod.org CONTACT: [email protected]
Xiaokang Pan, Lincoln Stein, Volker Brendel
Bioinform.2
2004 Applying Semantic Web Services to Bioinformatics: Experiences Gained, Lessons Learnt
Phillip Lord, Sean Bechhofer, Mark D. Wilkinson, Gary S. Schiltz, Damian Gessler, Duncan Hull, Carole A. Goble, Lincoln Stein
ISWC8
2001 The Distributed Annotation System
abstract
BACKGROUND: Currently, most genome annotation is curated by centralized groups with limited resources. Efforts to share annotations transparently among multiple groups have not yet been satisfactory. RESULTS: Here we introduce a concept called the Distributed Annotation System (DAS). DAS allows sequence annotations to be decentralized among multiple third-party annotators and integrated on an as-needed basis by client-side software. The communication between client and servers in DAS is defined by the DAS XML specification. Annotations are displayed in layers, one per server. Any client or server adhering to the DAS XML specification can participate in the system; we describe a simple prototype client and server example. CONCLUSIONS: The DAS specification is being used experimentally by Ensembl, WormBase, and the Berkeley Drosophila Genome Project. Continued success will depend on the readiness of the research community to adopt DAS and provide annotations. All components are freely available from the project website http://www.biodas.org/.
Robin D. Dowell, Rodney M. Jokerst, Allen Day, Sean R. Eddy, Lincoln Stein
BMC Bioinform.5
1999 SBOX: Put CGI Scripts in a Box
Lincoln Stein
USENIX ATC, General Track1
1998 The LabFlow System for Workflow Management in Large Scale Biology Research Laboratories
Nathan Goodman, Steve Rozen, Lincoln Stein
ISMB3
1998 The LabBase system for data management in large scale biology research laboratories
abstract
MOTIVATION: The development of laboratory information management systems (LIMSs) for large scale biology research projects can be a challenging problem. Many such projects generate complex datasets via complex procedures that undergo continuous refinement. A key software challenge is to simplify the database-development task so that databases can be built and modified quickly enough to keep pace with changing project-requirements. RESULTS: LabBase extends the facilities offered by relational database systems to simplify the task of creating databases for large scale biology research projects. LabBase provides a structural object data model, similar to ACEDB, and adds to this the concepts of Materials, Steps, and States: Materials are objects representing the identifiable things that participate in a laboratory protocol; Steps are objects reporting the results of a laboratory or analytical procedure; and States are objects denoting places in a laboratory protocol. The system provides a data definition language for succinctly defining laboratory databases, and operations for conveniently storing and retrieving data in such databases. The system also provides support for workflow management. LabBase is implemented in Perl5 and provides a natural interface for laboratory application programs written in Perl. AVAILABILITY: The software is freely available. Contact the authors. CONTACT: [email protected]
Nathan Goodman, Steve Rozen, Lincoln Stein, A. G. Smith
Bioinform.3
1997 Building human genome maps with radiation hybrids
abstract
Genome maps are crucial tools in human genetic research, providing known landmarks for locating disease genes and frameworks for large-scale sequencing.Radiation hybrid mapping is one technique for building genome maps.In this paper, we describe the methods used to build radiation hybrid maps of the entire human genome.We present the hidden Markov model that we employ to estimate the likelihood of a map despite uncertainty about the data, and we discuss the problem of searching for maximum-likelihood maps.We describe the graph algorithms used to find sparse but reliable init,ial maps and our methods of extending them.Finally, we show results validating our software on simulated data, and we describe our genome-wide human radiat,ion hybrid maps and the evidence supporting them.
Donna K. Slonim, Leonid Kruglyak, Lincoln Stein, Eric S. Lander
RECOMB3
1994 Building a Laboratory Information System Around a C++-Based Object-Oriented DBMS
Nathan Goodman, Steve Rozen, Lincoln Stein
VLDB3