Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Peter W. Rose

dblp:50/1609 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
3since 2021 · last 2024
0000-0001-9981-9750ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
8 papers
Bioinformatics and computational biology · 81% Medical and health informatics · 19%
Computer graphics and multimedia
1 paper
Rendering · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › knowledge representation in biology
biomedical knowledge graph
1.422024
Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024
The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023
Bioinformatics and computational biology
structural bioinformatics
0.952017
BioJava-ModFinder: identification of protein modifications in 3D structures from the Protein Data Bank · Bioinform. 2017
Integrating genomic information with protein sequence and 3D atomic level structure at the RCSB protein data bank · Bioinform. 2016
RCSB PDB Mobile: iOS and Android mobile apps to provide data access and visualization to the RCSB Protein Data Bank · Bioinform. 2015
Bioinformatics and computational biology
knowledge graph
0.922024
The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023
Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024
Bioinformatics and computational biology › biomedical text mining
biomedical question answering
0.812024
Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024
Medical and health informatics
retrieval-augmented generation
0.812024
Biomedical knowledge graph-optimized prompt generation for large language models · Bioinform. 2024
Bioinformatics and computational biology › data integration
biomedical data integration
0.712023
The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023
Medical and health informatics
precision medicine
0.712023
The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information · Bioinform. 2023
Bioinformatics and computational biology
protein structure analysis
0.322017
BioJava-ModFinder: identification of protein modifications in 3D structures from the Protein Data Bank · Bioinform. 2017
BioJava: an open-source framework for bioinformatics in 2012 · Bioinform. 2012
Bioinformatics and computational biology › molecular informatics
molecular visualization
0.312018
NGL viewer: web-based molecular graphics for large complexes · Bioinform. 2018
Bioinformatics and computational biology › structural bioinformatics › protein structure representation
protein structure visualization
0.212015
RCSB PDB Mobile: iOS and Android mobile apps to provide data access and visualization to the RCSB Protein Data Bank · Bioinform. 2015
Bioinformatics and computational biology › bioinformatics infrastructure
bioinformatics framework
0.112012
BioJava: an open-source framework for bioinformatics in 2012 · Bioinform. 2012
Bioinformatics and computational biology › protein structure analysis
protein structure alignment
0.112010
Pre-calculated protein structure alignments at the RCSB PDB website · Bioinform. 2010
Bioinformatics and computational biology › structural bioinformatics › protein structure database
protein data bank curation
0.112017
BioJava-ModFinder: identification of protein modifications in 3D structures from the Protein Data Bank · Bioinform. 2017
Bioinformatics and computational biology
genome annotation
0.112016
Integrating genomic information with protein sequence and 3D atomic level structure at the RCSB protein data bank · Bioinform. 2016
Bioinformatics and computational biology
genomics
0.112016
Integrating genomic information with protein sequence and 3D atomic level structure at the RCSB protein data bank · Bioinform. 2016
Bioinformatics and computational biology
sequence alignment
0.012012
BioJava: an open-source framework for bioinformatics in 2012 · Bioinform. 2012
Bioinformatics and computational biology
sequence analysis
0.012012
BioJava: an open-source framework for bioinformatics in 2012 · Bioinform. 2012
Bioinformatics and computational biology › structural bioinformatics
protein structure database
0.012010
Pre-calculated protein structure alignments at the RCSB PDB website · Bioinform. 2010

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 0.8large language model · 0.8embedding-based context pruning · 0.8ontology-based integration · 0.7binary compression format · 0.7WebGL · 0.7REST API · 0.7structure scanning · 0.3annotation integration · 0.3data integration · 0.2
YearPublicationVenuePosition
2024 Biomedical knowledge graph-optimized prompt generation for large language models
abstract
MOTIVATION: Large language models (LLMs) are being adopted at an unprecedented rate, yet still face challenges in knowledge-intensive domains such as biomedicine. Solutions such as pretraining and domain-specific fine-tuning add substantial computational overhead, requiring further domain-expertise. Here, we introduce a token-optimized and robust Knowledge Graph-based Retrieval Augmented Generation (KG-RAG) framework by leveraging a massive biomedical KG (SPOKE) with LLMs such as Llama-2-13b, GPT-3.5-Turbo, and GPT-4, to generate meaningful biomedical text rooted in established knowledge. RESULTS: Compared to the existing RAG technique for Knowledge Graphs, the proposed method utilizes minimal graph schema for context extraction and uses embedding methods for context pruning. This optimization in context extraction results in more than 50% reduction in token consumption without compromising the accuracy, making a cost-effective and robust RAG implementation on proprietary LLMs. KG-RAG consistently enhanced the performance of LLMs across diverse biomedical prompts by generating responses rooted in established knowledge, accompanied by accurate provenance and statistical evidence (if available) to substantiate the claims. Further benchmarking on human curated datasets, such as biomedical true/false and multiple-choice questions (MCQ), showed a remarkable 71% boost in the performance of the Llama-2 model on the challenging MCQ dataset, demonstrating the framework's capacity to empower open-source models with fewer parameters for domain-specific questions. Furthermore, KG-RAG enhanced the performance of proprietary GPT models, such as GPT-3.5 and GPT-4. In summary, the proposed framework combines explicit and implicit knowledge of KG and LLM in a token optimized fashion, thus enhancing the adaptability of general-purpose LLMs to tackle domain-specific questions in a cost-effective fashion. AVAILABILITY AND IMPLEMENTATION: SPOKE KG can be accessed at https://spoke.rbvi.ucsf.edu/neighborhood.html. It can also be accessed using REST-API (https://spoke.rbvi.ucsf.edu/swagger/). KG-RAG code is made available at https://github.com/BaranziniLab/KG_RAG. Biomedical benchmark datasets used in this study are made available to the research community in the same GitHub repository.
Karthik Soman, Peter W. Rose, John Scotter Morris, Rabia E. Akbas, Brett Smith, Braian Peetoom, Catalina Villouta-Reyes, Gabriel Cerono, Yongmei Shi, Angela Rizk-Jackson, Sharat Israni, Charlotte A. Nelson, Sui Huang, Sergio Baranzini
Bioinform.2
2023 The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information
abstract
MOTIVATION: Knowledge graphs (KGs) are being adopted in industry, commerce and academia. Biomedical KG presents a challenge due to the complexity, size and heterogeneity of the underlying information. RESULTS: In this work, we present the Scalable Precision Medicine Open Knowledge Engine (SPOKE), a biomedical KG connecting millions of concepts via semantically meaningful relationships. SPOKE contains 27 million nodes of 21 different types and 53 million edges of 55 types downloaded from 41 databases. The graph is built on the framework of 11 ontologies that maintain its structure, enable mappings and facilitate navigation. SPOKE is built weekly by python scripts which download each resource, check for integrity and completeness, and then create a 'parent table' of nodes and edges. Graph queries are translated by a REST API and users can submit searches directly via an API or a graphical user interface. Conclusions/Significance: SPOKE enables the integration of seemingly disparate information to support precision medicine efforts. AVAILABILITY AND IMPLEMENTATION: The SPOKE neighborhood explorer is available at https://spoke.rbvi.ucsf.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
John Scotter Morris, Karthik Soman, Rabia E. Akbas, Xiaoyuan Zhou, Brett Smith, Elaine C. Meng, Conrad C. Huang, Gabriel Cerono, Gundolf Schenk, Angela Rizk-Jackson, Adil Harroud, Lauren M. Sanders, Sylvain V. Costes, Krish Bharat, Arjun Chakraborty, Alexander R. Pico, Taline Mardirossian, Michael J. Keiser, Alice Tang, Josef Hardi, Yongmei Shi, Mark A. Musen, Sharat Israni, Sui Huang, Peter W. Rose, Charlotte A. Nelson, Sergio Baranzini
Bioinform.25
2021 Ten simple rules to cultivate transdisciplinary collaboration in data science
abstract
Author(s): Sahneh, Faryad; Balk, Meghan A; Kisley, Marina; Chan, Chi-kwan; Fox, Mercury; Nord, Brian; Lyons, Eric; Swetnam, Tyson; Huppenkothen, Daniela; Sutherland, Will; Walls, Ramona L; Quinn, Daven P; Tarin, Tonantzin; LeBauer, David; Ribes, David; Birnie, Dunbar P; Lushbough, Carol; Carr, Eric; Nearing, Grey; Fischer, Jeremy; Tyle, Kevin; Carrasco, Luis; Lang, Meagan; Rose, Peter W; Rushforth, Richard R; Roy, Samapriya; Matheson, Thomas; Lee, Tina; Brown, C Titus; Teal, Tracy K; Papeș, Monica; Kobourov, Stephen; Merchant, Nirav | Editor(s): Schwartz, Russell
Faryad Sahneh, Meghan A. Balk, Marina Kisley, Chi-Kwan Chan, Mercury Fox, Brian Nord, Eric Lyons 0002, Tyson Lee Swetnam, Daniela Huppenkothen, Will Sutherland, Ramona L. Walls, Daven P. Quinn, Tonantzin Tarin, David S. LeBauer, David Ribes, Dunbar P. Birnie III, Carol Lushbough, Eric Carr, Grey Nearing, Jeremy Fischer, Kevin Tyle, Luis Carrasco, Meagan Lang, Peter W. Rose, Richard R. Rushforth, Samapriya Roy, Thomas Matheson, Tina Lee, C. Titus Brown, Tracy K. Teal, Monica Papes, Stephen G. Kobourov, Nirav C. Merchant
PLoS Comput. Biol.24
2019 Analyzing the symmetrical arrangement of structural repeats in proteins with CE-Symm
abstract
Many proteins fold into highly regular and repetitive three dimensional structures. The analysis of structural patterns and repeated elements is fundamental to understand protein function and evolution. We present recent improvements to the CE-Symm tool for systematically detecting and analyzing the internal symmetry and structural repeats in proteins. In addition to the accurate detection of internal symmetry, the tool is now capable of i) reporting the type of symmetry, ii) identifying the smallest repeating unit, iii) describing the arrangement of repeats with transformation operations and symmetry axes, and iv) comparing the similarity of all the internal repeats at the residue level. CE-Symm 2.0 helps the user investigate proteins with a robust and intuitive sequence-to-structure analysis, with many applications in protein classification, functional annotation and evolutionary studies. We describe the algorithmic extensions of the method and demonstrate its applications to the study of interesting cases of protein evolution.
Spencer Bliven, Aleix Lafita, Peter W. Rose, Guido Capitani, Andreas Prlic, Philip E. Bourne
PLoS Comput. Biol.3
2019 BioJava 5: A community driven open-source bioinformatics library
abstract
BioJava is an open-source project that provides a Java library for processing biological data. The project aims to simplify bioinformatic analyses by implementing parsers, data structures, and algorithms for common tasks in genomics, structural biology, ontologies, phylogenetics, and more. Since 2012, we have released two major versions of the library (4 and 5) that include many new features to tackle challenges with increasingly complex macromolecular structure data. BioJava requires Java 8 or higher and is freely available under the LGPL 2.1 license. The project is hosted on GitHub at https://github.com/biojava/biojava. More information and documentation can be found online on the BioJava website (http://www.biojava.org) and tutorial (https://github.com/biojava/biojava-tutorial). All inquiries should be directed to the GitHub page or the BioJava mailing list (http://lists.open-bio.org/mailman/listinfo/biojava-l).
Aleix Lafita, Spencer Bliven, Andreas Prlic, Dmytro Guzenko, Peter W. Rose, Anthony R. Bradley, Paolo Pavan, Douglas Myers-Turnbull, Yana Rose, Michael L. Heuer, Matt Larson, Stephen K. Burley, Jose M. Duarte
PLoS Comput. Biol.5
2019 Ten simple rules for writing and sharing computational analyses in Jupyter Notebooks
abstract
Author(s): Rule, Adam; Birmingham, Amanda; Zuniga, Cristal; Altintas, Ilkay; Huang, Shih-Cheng; Knight, Rob; Moshiri, Niema; Nguyen, Mai H; Rosenthal, Sara Brin; Pérez, Fernando; Rose, Peter W | Editor(s): Lewitter, Fran
Adam Rule, Amanda Birmingham, Cristal Zuñiga, Ilkay Altintas, Shih-Cheng Huang, Rob Knight 0001, Niema Moshiri, Mai H. Nguyen, Sara Brin Rosenthal, Peter W. Rose
PLoS Comput. Biol.11
2018 NGL viewer: web-based molecular graphics for large complexes
abstract
Motivation: The interactive visualization of very large macromolecular complexes on the web is becoming a challenging problem as experimental techniques advance at an unprecedented rate and deliver structures of increasing size. Results: We have tackled this problem by developing highly memory-efficient and scalable extensions for the NGL WebGL-based molecular viewer and by using Macromolecular Transmission Format (MMTF), a binary and compressed MMTF. These enable NGL to download and render molecular complexes with millions of atoms interactively on desktop computers and smartphones alike, making it a tool of choice for web-based molecular visualization in research and education. Availability and implementation: The source code is freely available under the MIT license at github.com/arose/ngl and distributed on NPM (npmjs.com/package/ngl). MMTF-JavaScript encoders and decoders are available at github.com/rcsb/mmtf-javascript.
Alexander S. Rose, Anthony R. Bradley, Yana Rose, Jose M. Duarte, Andreas Prlic, Peter W. Rose
Bioinform.6
2017 BioJava-ModFinder: identification of protein modifications in 3D structures from the Protein Data Bank
abstract
SUMMARY: We developed a new software tool, BioJava-ModFinder, for identifying protein modifications observed in 3D structures archived in the Protein Data Bank (PDB). Information on more than 400 types of protein modifications were collected and curated from annotations in PDB, RESID, and PSI-MOD. We divided these modifications into three categories: modified residues, attachment modifications, and cross-links. We have developed a systematic method to identify these modifications in 3D protein structures. We have integrated this package with the RCSB PDB web application and added protein modification annotations to the sequence diagram and structure display. By scanning all 3D structures in the PDB using BioJava-ModFinder, we identified more than 30 000 structures with protein modifications, which can be searched, browsed, and visualized on the RCSB PDB website. AVAILABILITY AND IMPLEMENTATION: BioJava-ModFinder is available as open source (LGPL license) at ( https://github.com/biojava/biojava/tree/master/biojava-modfinder ). The RCSB PDB can be accessed at http://www.rcsb.org . CONTACT: [email protected].
Jianjiong Gao, Andreas Prlic, Chunxiao Bi, Wolfgang Bluhm, Dong Xu 0002, Philip E. Bourne, Peter W. Rose
Bioinform.8
2017 MMTF - An efficient file format for the transmission, visualization, and analysis of macromolecular structures
abstract
Recent advances in experimental techniques have led to a rapid growth in complexity, size, and number of macromolecular structures that are made available through the Protein Data Bank. This creates a challenge for macromolecular visualization and analysis. Macromolecular structure files, such as PDB or PDBx/mmCIF files can be slow to transfer, parse, and hard to incorporate into third-party software tools. Here, we present a new binary and compressed data representation, the MacroMolecular Transmission Format, MMTF, as well as software implementations in several languages that have been developed around it, which address these issues. We describe the new format and its APIs and demonstrate that it is several times faster to parse, and about a quarter of the file size of the current standard format, PDBx/mmCIF. As a consequence of the new data representation, it is now possible to visualize structures with millions of atoms in a web browser, keep the whole PDB archive in memory or parse it within few minutes on average computers, which opens up a new way of thinking how to design and implement efficient algorithms in structural bioinformatics. The PDB archive is available in MMTF file format through web services and data that are updated on a weekly basis.
Anthony R. Bradley, Alexander S. Rose, Antonín Pavelka, Yana Rose, Jose M. Duarte, Andreas Prlic, Peter W. Rose
PLoS Comput. Biol.7
2016 Integrating genomic information with protein sequence and 3D atomic level structure at the RCSB protein data bank
abstract
The Protein Data Bank (PDB) now contains more than 120,000 three-dimensional (3D) structures of biological macromolecules. To allow an interpretation of how PDB data relates to other publicly available annotations, we developed a novel data integration platform that maps 3D structural information across various datasets. This integration bridges from the human genome across protein sequence to 3D structure space. We developed novel software solutions for data management and visualization, while incorporating new libraries for web-based visualization using SVG graphics. AVAILABILITY AND IMPLEMENTATION: The new views are available from http://www.rcsb.org and software is available from https://github.com/rcsb/. CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Andreas Prlic, Tara Kalro, Roshni Bhattacharya, Cole H. Christie, Stephen K. Burley, Peter W. Rose
Bioinform.6
2015 RCSB PDB Mobile: iOS and Android mobile apps to provide data access and visualization to the RCSB Protein Data Bank
abstract
SUMMARY: The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) resource provides tools for query, analysis and visualization of the 3D structures in the PDB archive. As the mobile Web is starting to surpass desktop and laptop usage, scientists and educators are beginning to integrate mobile devices into their research and teaching. In response, we have developed the RCSB PDB Mobile app for the iOS and Android mobile platforms to enable fast and convenient access to RCSB PDB data and services. Using the app, users from the general public to expert researchers can quickly search and visualize biomolecules, and add personal annotations via the RCSB PDB's integrated MyPDB service. AVAILABILITY AND IMPLEMENTATION: RCSB PDB Mobile is freely available from the Apple App Store and Google Play (http://www.rcsb.org).
Greg B. Quinn, Chunxiao Bi, Cole H. Christie, Kyle Pang, Andreas Prlic, Takanori Nakane, Christine Zardecki, Maria Voigt, Helen M. Berman, Philip E. Bourne, Peter W. Rose
Bioinform.11
2012 BioJava: an open-source framework for bioinformatics in 2012
abstract
UNLABELLED: BioJava is an open-source project for processing of biological data in the Java programming language. We have recently released a new version (3.0.5), which is a major update to the code base that greatly extends its functionality. RESULTS: BioJava now consists of several independent modules that provide state-of-the-art tools for protein structure comparison, pairwise and multiple sequence alignments, working with DNA and protein sequences, analysis of amino acid properties, detection of protein modifications and prediction of disordered regions in proteins as well as parsers for common file formats using a biologically meaningful data model. AVAILABILITY: BioJava is an open-source project distributed under the Lesser GPL (LGPL). BioJava can be downloaded from the BioJava website (http://www.biojava.org). BioJava requires Java 1.6 or higher. All inquiries should be directed to the BioJava mailing lists. Details are available at http://biojava.org/wiki/BioJava:MailingLists.
Andreas Prlic, Andy Yates, Spencer Bliven, Peter W. Rose, Julius O. B. Jacobsen, Peter V. Troshin, Mark Chapman, Jianjiong Gao, Chuan Hock Koh, Sylvain Foisy, Richard C. G. Holland, Gediminas Rimsa, Michael L. Heuer, Hannes Brandstätter-Müller, Philip E. Bourne, Scooter Willis
Bioinform.4
2010 Pre-calculated protein structure alignments at the RCSB PDB website
abstract
SUMMARY: With the continuous growth of the RCSB Protein Data Bank (PDB), providing an up-to-date systematic structure comparison of all protein structures poses an ever growing challenge. Here, we present a comparison tool for calculating both 1D protein sequence and 3D protein structure alignments. This tool supports various applications at the RCSB PDB website. First, a structure alignment web service calculates pairwise alignments. Second, a stand-alone application runs alignments locally and visualizes the results. Third, pre-calculated 3D structure comparisons for the whole PDB are provided and updated on a weekly basis. These three applications allow users to discover novel relationships between proteins available either at the RCSB PDB or provided by the user. AVAILABILITY AND IMPLEMENTATION: A web user interface is available at http://www.rcsb.org/pdb/workbench/workbench.do. The source code is available under the LGPL license from http://www.biojava.org. A source bundle, prepared for local execution, is available from http://source.rcsb.org CONTACT: [email protected]; [email protected].
Andreas Prlic, Spencer Bliven, Peter W. Rose, Wolfgang Bluhm, Chris Bizon, Adam Godzik, Philip E. Bourne
Bioinform.3
2010 Integration of open access literature into the RCSB Protein Data Bank using BioLit
abstract
BACKGROUND: Biological data have traditionally been stored and made publicly available through a variety of on-line databases, whereas biological knowledge has traditionally been found in the printed literature. With journals now on-line and providing an increasing amount of open access content, often free of copyright restriction, this distinction between database and literature is blurring. To exploit this opportunity we present the integration of open access literature with the RCSB Protein Data Bank (PDB). RESULTS: BioLit provides an enhanced view of articles with markup of semantic data and links to biological databases, based on the content of the article. For example, words matching to existing biological ontologies are highlighted and database identifiers are linked to their database of origin. Among other functions, it identifies PDB IDs that are mentioned in the open access literature, by parsing the full text for all research articles in PubMed Central (PMC) and exposing the results as simple XML Web Services. Here, we integrate BioLit results with the RCSB PDB website by using these services to find PDB IDs that are mentioned in research articles and subsequently retrieving abstract, figures, and text excerpts for those articles. A new RCSB PDB literature view permits browsing through the figures and abstracts of the articles that mention a given structure. The BioLit Web Services that are providing the underlying data are publicly accessible. A client library is provided that supports querying these services (Java). CONCLUSIONS: The integration between literature and websites, as demonstrated here with the RCSB PDB, provides a broader view for how a given structure has been analyzed and used. This approach detects the mention of a PDB structure even if it is not formally cited in the paper. Other structures related through the same literature references can also be identified, possibly providing new scientific insight. To our knowledge this is the first time that database and literature have been integrated in this way and it speaks to the opportunities afforded by open and free access to both database and literature content.
Andreas Prlic, Marco A. Martinez, Bojan Beran, Benjamin T. Yukich, Peter W. Rose, Philip E. Bourne, J. Lynn Fink
BMC Bioinform.6
2010 Will Widgets and Semantic Tagging Change Computational Biology?
abstract
We argue here, through the use of several examples from our work in support of structural biology, that the answer to the question posed by the title of this Perspective is a resounding yes. The discussion that follows is aimed primarily at those of the journal's readers who are biological resource developers and Web page developers interested in developing the richest possible Web pages. However, those of you who simply use biological resources might find this a helpful discussion in understanding what is on the horizon. Whatever your interest, please let us hear your opinion on the question posed by this Perspective through the associated comment feature.
Philip E. Bourne, Bojan Beran, Chunxiao Bi, Wolfgang Bluhm, Roland L. Dunbrack Jr., Andreas Prlic, Greg B. Quinn, Peter W. Rose, Raship Shah, Wendy Tao, Brian D. Weitzner, Benjamin T. Yukich
PLoS Comput. Biol.8