VLDB 2026 Research / reviewers in the wild / expert
Philip E. Bourne
dblp:b/PhilipEBourne
· DBLP profile ↗
123ranked-venue papers
36as first author
11since 2021 · last 2024
0000-0002-7618-7292ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 121 · 34 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Prop3D: A flexible, Python-based platform for machine learning with protein structural properties and biophysical dataabstractBACKGROUND: Machine learning (ML) has a rich history in structural bioinformatics, and modern approaches, such as deep learning, are revolutionizing our knowledge of the subtle relationships between biomolecular sequence, structure, function, dynamics and evolution. As with any advance that rests upon statistical learning approaches, the recent progress in biomolecular sciences is enabled by the availability of vast volumes of sufficiently-variable data. To be useful, such data must be well-structured, machine-readable, intelligible and manipulable. These and related requirements pose challenges that become especially acute at the computational scales typical in ML. Furthermore, in structural bioinformatics such data generally relate to protein three-dimensional (3D) structures, which are inherently more complex than sequence-based data. A significant and recurring challenge concerns the creation of large, high-quality, openly-accessible datasets that can be used for specific training and benchmarking tasks in ML pipelines for predictive modeling projects, along with reproducible splits for training and testing. RESULTS: Here, we report 'Prop3D', a platform that allows for the creation, sharing and extensible reuse of libraries of protein domains, featurized with biophysical and evolutionary properties that can range from detailed, atomically-resolved physicochemical quantities (e.g., electrostatics) to coarser, residue-level features (e.g., phylogenetic conservation). As a community resource, we also supply a 'Prop3D-20sf' protein dataset, obtained by applying our approach to CATH . We have developed and deployed the Prop3D framework, both in the cloud and on local HPC resources, to systematically and reproducibly create comprehensive datasets via the Highly Scalable Data Service ( HSDS ). Our datasets are freely accessible via a public HSDS instance, or they can be used with accompanying Python wrappers for popular ML frameworks. CONCLUSION: Prop3D and its associated Prop3D-20sf dataset can be of broad utility in at least three ways. Firstly, the Prop3D workflow code can be customized and deployed on various cloud-based compute platforms, with scalability achieved largely by saving the results to distributed HDF5 files via HSDS . Secondly, the linked Prop3D-20sf dataset provides a hand-crafted, already-featurized dataset of protein domains for 20 highly-populated CATH families; importantly, provision of this pre-computed resource can aid the more efficient development (and reproducible deployment) of ML pipelines. Thirdly, Prop3D-20sf's construction explicitly takes into account (in creating datasets and data-splits) the enigma of 'data leakage', stemming from the evolutionary relationships between proteins. Eli J. Draizen, John Readey, Cameron Mura, Philip E. Bourne |
BMC Bioinform. | 4 |
| 2023 | Ten simple rules for managing laboratory informationabstractInformation is the cornerstone of research, from experimental (meta)data and computational processes to complex inventories of reagents and equipment. These 10 simple rules discuss best practices for leveraging laboratory information management systems to transform this large information load into useful scientific findings. Casey-Tyler Berezin, Luis U. Aguilera, Sonja Billerbeck, Philip E. Bourne, Douglas Densmore, Paul S. Freemont, Thomas E. Gorochowski, Sarah I. Hernandez, Nathan J. Hillson, Connor R. King, Michael Köpke, Shuyi Ma, Katie M. Miller, Tae Seok Moon, Jason H. Moore, Brian Munsky, Chris J. Myers, Dequina A. Nicholas, Samuel J. Peccoud, Jean Peccoud |
PLoS Comput. Biol. | 4 |
| 2023 | End-to-end sequence-structure-function meta-learning predicts genome-wide chemical-protein interactions for dark proteinsabstractSystematically discovering protein-ligand interactions across the entire human and pathogen genomes is critical in chemical genomics, protein function prediction, drug discovery, and many other areas. However, more than 90% of gene families remain "dark"-i.e., their small-molecule ligands are undiscovered due to experimental limitations or human/historical biases. Existing computational approaches typically fail when the dark protein differs from those with known ligands. To address this challenge, we have developed a deep learning framework, called PortalCG, which consists of four novel components: (i) a 3-dimensional ligand binding site enhanced sequence pre-training strategy to encode the evolutionary links between ligand-binding sites across gene families; (ii) an end-to-end pretraining-fine-tuning strategy to reduce the impact of inaccuracy of predicted structures on function predictions by recognizing the sequence-structure-function paradigm; (iii) a new out-of-cluster meta-learning algorithm that extracts and accumulates information learned from predicting ligands of distinct gene families (meta-data) and applies the meta-data to a dark gene family; and (iv) a stress model selection step, using different gene families in the test data from those in the training and development data sets to facilitate model deployment in a real-world scenario. In extensive and rigorous benchmark experiments, PortalCG considerably outperformed state-of-the-art techniques of machine learning and protein-ligand docking when applied to dark gene families, and demonstrated its generalization power for target identifications and compound screenings under out-of-distribution (OOD) scenarios. Furthermore, in an external validation for the multi-target compound screening, the performance of PortalCG surpassed the rational design from medicinal chemists. Our results also suggest that a differentiable sequence-structure-function deep learning framework, where protein structural information serves as an intermediate layer, could be superior to conventional methodology where predicted protein structures were used for the compound screening. We applied PortalCG to two case studies to exemplify its potential in drug discovery: designing selective dual-antagonists of dopamine receptors for the treatment of opioid use disorder (OUD), and illuminating the understudied human genome for target diseases that do not yet have effective and safe therapeutics. Our results suggested that PortalCG is a viable solution to the OOD problem in exploring understudied regions of protein functional space. Tian Cai, Li Xie 0002, Shuo Zhang 0010, Muge Chen, Di He 0003, Amitesh Badkul, Hari Krishna Namballa, Michael Dorogan, Wayne W. Harding, Cameron Mura, Philip E. Bourne, Lei Xie 0006 |
PLoS Comput. Biol. | 12 |
| 2023 | Ten simple rules for humane data scienceabstractThe capabilities of data science can help us see hidden patterns, customize services, and advance biomedicine and science.As data science permeates industry and academia, a question that often arises is: How can we use these capabilities to genuinely help the world?Considering this question brings to mind terms such as "responsible data science" and "data for good."Inherent in these terms is the desire that data science improve the human condition-in other words, the desire to undertake humane data science.How can we do this?What follows are 10 simple rules to contemplate.(Note that we are referring to data science in a holistic way.Along with machine learning and AI methods, we include all aspects of the data acquisition, analysis, dissemination, and application pipeline, and their accompanying human and socioeconomic factors.) Hassan Masum, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2023 | Ten simple rules for serving as an editorabstractWe believe that readers, particularly those at relatively early career stages, could benefit from ten simple rules (TSR) on the topic of "serving as an editor."By this phrase, we mean that role which may variously be called "handling editor," "academic editor," "scientific editor" and so on-in other words, the individual who oversees the process of shepherding a written piece of scientific work from the point of manuscript submission through to peer review and, ultimately, either publication or rejection.We mean this in contradistinction to, say, serving as an editorial advisory board member, a "section editor," or an editor-in-chief or such.Those roles are of interest too, in terms of career and professional development; however, those types of positions generally begin later in one's career (e.g., as a well-established scientist), versus nearer the start of an independent career (e.g., late-postdoc or early-faculty), spurring us to focus this piece more on the context of "handling editor."The goal of this TSR is to offer guidance that can help you, the reader-in your current or future roles as a novice handling editor-be the type of editor whom you might have liked to have dealt with yourself, in your own experience and interactions thus far in publishing your work.This piece can be viewed as complementary to an early TSR for reviewers [1], and the closest material of which we are aware is Erren and Erren's "Simple Rules for Editors'?Here is One Rule to Tackle Neglected Problems of Publishing," published as a correspondence in this journal 15 years ago [2].That insightful piece suggested an "Editor Rule for Appropriate Recognition" as a way to improve the recognition and credit due to those who contribute to a submitted work, and yet who may be less visible or even entirely overlooked (possibly to the point of omission from an author list), perhaps because they are relatively young or less experienced than more senior authors.Specifically, Erren and Erren's proposed "rule" called on editors to ask authors for a written statement that (i) avows that no substantial contributors have been omitted from an authorship list; and (ii) explicitly delineates the oftentimes key role of more junior/overlooked authors (and/or researchers that the work cites) in formulating the hypotheses or rationale that underlies the submitted study.We agree with those proposed practices.In the present work, we focus chiefly on the mechanics and best practices of one's editorial responsibilities when handling a manuscript, starting at the initial point of being invited to serve as an editor. Cameron Mura, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2022 | Ten simple rules for improving communication among scientistsabstractCommunication is a fundamental part of scientific development and methodology. With the advancement of the internet and social networks, communication has become rapid and sometimes overwhelming, especially in science. It is important to provide scientists with useful, effective, and dynamic tools to establish and build a fluid communication framework that allows for scientific advancement. Therefore, in this article, we present advice and recommendations that can help promote and improve science communication while respecting an adequate balance in the degree of commitment toward collaborative work. We have developed 10 rules shown in increasing order of commitment that are grouped into 3 key categories: (1) speak (based on active participation); (2) join (based on joining scientific groups); and (3) assess (based on the analysis and retrospective consideration of the weaknesses and strengths). We include examples and resources that provide actionable strategies for involvement and engagement with science communication, from basic steps to more advanced, introspective, and long-term commitments. Overall, we aim to help spread science from within and encourage and engage scientists to become involved in science communication effectively and dynamically. Carla Bautista, Narjes Alfuraiji, Anna Drangowska-Way, Karishma Gangwani, Alida de Flamingh, Philip E. Bourne |
PLoS Comput. Biol. | 6 |
| 2022 | Ten simple rules for using entrepreneurship skills to improve research careers and cultureabstractThe academic path is not an easy one.The acknowledgment of the problems of the research precariat [1], the lack of stability and certainty in our careers, is not a new thing.The number of permanent positions in academic research has always limited the career prospects of many postdocs, but there are concerns the increasing supply of graduates is exacerbating this issue [2], and for scientists, job satisfaction is at an all-time low [3].Indeed, the current situation is forcing many of us to reevaluate our options [4].For many, this may mean leaving the academic world and moving to industry or enterprise, but this often can feel like failure [5].We argue that the establishment of a scientific career can be thought of as an entrepreneurial enterprise.We contend that an acknowledgment of this viewpoint and reevaluation of what constitutes value and impact in this system can contribute towards an improved research culture.With this in mind, we present some rules informed by the world of enterprise and business to help the early career researcher develop and navigate their careers.Entrepreneurship can be defined as tAU : PleasenotethatasperPLOSstyle; italicsshouldnotbeusedfore he means by which new organisations are formed with their resultant job and wealth creation [6].We can refocus this definition on the individual and define career entrepreneurship as the process of developing a career that results in increased value for the individual.Some common traits have been identified in entrepreneurship: risk taking, achievement, autonomy, self-efficacy (the ability to complete tasks), and locus of control (the control we believe we have on the outcomes of our lives [7]) [8], and these traits are also important markers and indicators for career progression in scientific research.Druker describes innovation as the ability to exploit change and argues that looking for change and exploiting it as an opportunity defines an entrepreneur [9].Change is a commodity that is not in short supply in the world of scientific research, with changing political and funding landscapes as well as new discoveries and knowledge.Through the definitions and perspectives offered above, we can prepare ourselves to be able to navigate and exploit changes for the benefit of our careers using traits and techniques also found in entrepreneurship and in innovators. Matt Bawn, David Dent, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2022 | Ten simple rules for good leadership
Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2022 | Ten simple rules for organizing a special session at a scientific conferenceabstractSpecial sessions are important parts of scientific meetings and conferences: They gather together researchers and students interested in a specific topic and can strongly contribute to the success of the conference itself. Moreover, they can be the first step for trainees and students to the organization of a scientific event. Organizing a special session, however, can be uneasy for beginners and students. Here, we provide ten simple rules to follow to organize a special session at a scientific conference. Davide Chicco, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2021 | Informatics-enabled citizen science to advance health equityabstractThe COVID-19 pandemic has once again highlighted the ubiquity and persistence of health inequities along with our inability to respond to them in a timely and effective manner. There is an opportunity to address the limitations of our current approaches through new models of informatics-enabled research and clinical practice that shift the norm from small- to large-scale patient engagement. We propose augmenting our approach to address health inequities through informatics-enabled citizen science, challenging the types of questions being asked, prioritized, and acted upon. We envision this democratization of informatics that builds upon the inclusive tradition of community-based participatory research (CBPR) as a logical and transformative step toward improving individual, community, and population health in a way that deeply reflects the needs of historically marginalized populations. Rupa Valdez, Don E. Detmer, Philip E. Bourne, Katherine K. Kim, Robin Austin, Anna McCollister-Slipp, Courtney C. Rogers, Karen C. Waters-Wicks |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Ten simple rules for starting (and sustaining) an academic data science initiativeabstractData science has emerged as a new paradigm for research. Readers of this journal might be tempted to say this is the research we have been doing all along. However, we contest that there is something fundamentally different in terms of the dimensions of data, diversity of disciplines, as well as the role of the private sector, than what has gone before. We take this position based on our collective experiences in, and observations of, calls to action around data science over the past 10 years. Those calls have resulted in many notable and successful responses from US universities.
A working definition of data science
Defining data science is like defining the internet—ask 10 people and you get 10 different answers. What most would likely agree on, at a high level of abstraction, is that it draws from statistics, computer science, and applied mathematics to operate on data from one or more domains leading to outcomes not achieved otherwise. The extent to which domain knowledge is incorporated in the work of data science varies, but it is essential for achieving meaningful outcomes. Outcomes that have implications to us as humans and our collective communities and society that, in turn, need to be addressed as part of the data life cycle [1]. In short, data science transcends traditional disciplinary boundaries to discover new insights not owned by any one existing discipline, driven by endless streams of digital data with the promise of translation to societal benefit. Micaela S. Parker, Arlyn E. Burgess, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2020 | Ten simple rules for more objective decision-making
Anthony C. Fletcher, Georges A. Wagner, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2020 | Ten simple rules for researchers while in isolation from a pandemicabstractThe scale and intensity of the coronavirus disease 2019 (COVID-19) worldwide pandemic is unprecedented in all our lifetimes This is written for all of us involved in scientific research - graduate student, postdoc, academic, staff scientist, in academia, government or industry Rule 3: Follow institutional guidance and provide feedback By institution, we mean everything from the government (federal, state, and local) to your workplace to your individual laboratory Rule 4: Embrace a new work habit and environment How science is conducted has changed, almost overnight [Extracted from the article] Copyright of PLoS Computational Biology is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission However, users may print, download, or email articles for individual use This abstract may be abridged No warranty is given about the accuracy of the copy Users should refer to the original published version of the material for the full abstract (Copyright applies to all Abstracts ) Hoe-Han Goh, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2020 | Ten simple rules for writing scientific op-ed articles
Hoe-Han Goh, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2020 | Ten simple rules for starting research in your late teens
Cameron Mura, Mike Chalupa, Abigail M. Newbury, Jack Chalupa, Philip E. Bourne |
PLoS Comput. Biol. | 5 |
| 2019 | Natural language processing of symptoms documented in free-text narratives of electronic health records: a systematic reviewabstractOBJECTIVE: Natural language processing (NLP) of symptoms from electronic health records (EHRs) could contribute to the advancement of symptom science. We aim to synthesize the literature on the use of NLP to process or analyze symptom information documented in EHR free-text narratives. MATERIALS AND METHODS: Our search of 1964 records from PubMed and EMBASE was narrowed to 27 eligible articles. Data related to the purpose, free-text corpus, patients, symptoms, NLP methodology, evaluation metrics, and quality indicators were extracted for each study. RESULTS: Symptom-related information was presented as a primary outcome in 14 studies. EHR narratives represented various inpatient and outpatient clinical specialties, with general, cardiology, and mental health occurring most frequently. Studies encompassed a wide variety of symptoms, including shortness of breath, pain, nausea, dizziness, disturbed sleep, constipation, and depressed mood. NLP approaches included previously developed NLP tools, classification methods, and manually curated rule-based processing. Only one-third (n = 9) of studies reported patient demographic characteristics. DISCUSSION: NLP is used to extract information from EHR free-text narratives written by a variety of healthcare providers on an expansive range of symptoms across diverse clinical specialties. The current focus of this field is on the development of methods to extract symptom information and the use of symptom information for disease classification tasks rather than the examination of symptoms themselves. CONCLUSION: Future NLP studies should concentrate on the investigation of symptoms and symptom documentation in EHR free-text narratives. Efforts should be undertaken to examine patient characteristics and make symptom-related NLP algorithms or pipelines and vocabularies openly available. Theresa A. Koleck, Caitlin N. Dreisbach, Philip E. Bourne, Suzanne Bakken |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Analyzing the symmetrical arrangement of structural repeats in proteins with CE-SymmabstractMany proteins fold into highly regular and repetitive three dimensional structures. The analysis of structural patterns and repeated elements is fundamental to understand protein function and evolution. We present recent improvements to the CE-Symm tool for systematically detecting and analyzing the internal symmetry and structural repeats in proteins. In addition to the accurate detection of internal symmetry, the tool is now capable of i) reporting the type of symmetry, ii) identifying the smallest repeating unit, iii) describing the arrangement of repeats with transformation operations and symmetry axes, and iv) comparing the similarity of all the internal repeats at the residue level. CE-Symm 2.0 helps the user investigate proteins with a robust and intuitive sequence-to-structure analysis, with many applications in protein classification, functional annotation and evolutionary studies. We describe the algorithmic extensions of the method and demonstrate its applications to the study of interesting cases of protein evolution. Spencer Bliven, Aleix Lafita, Peter W. Rose, Guido Capitani, Andreas Prlic, Philip E. Bourne |
PLoS Comput. Biol. | 6 |
| 2019 | Ten simple rules to aid in achieving a visionabstractIn a career that now spans 40 years, I have had, on several occasions, opportunities to turn loose ideas into a unified vision and the resources to implement that vision.What do I mean by a vision?A vision, at least in my mind, is the ability to see something important to the future, perhaps before others do.Fulfilling that vision does not have to change the whole world (although that would be nice) but only to impact others in a positive way.Consider what I perceive has been my own visioning to provide some context.My visioning began in the 1990s when, seeing what computation was bringing to the life sciences through work on the human genome project, I was lucky enough to be able to envision and establish a bioinformatics laboratory before the idea became mainstream.In the early 2000s, it was a collective vision for what an exemplar data resource, namely, the Research Collaboratory for Structural Bioinformatics (RCSB) Protein Data Bank (PDB), should achieve.Around 2005, it was something dear to this readership, a vision for a new journal, PLOS Computational Biology, for which I was cofounder and Founding Editor-in-Chief for 7 years.Around 2007, it was forming a company, SciVee.tv, to envision how digital media other than print could be used to communicate science.Around 2014, as the first Associate Director for Data Science (ADDS) for the United States National Institutes of Health (NIH), the vision was how big data could catalyze change in life sciences research.Finally, now in 2019, the vision is how one of the first academic Schools of Data Science should be established and run.These either were, or are, great opportunities to lay out a vision and act upon that blueprint.Not all were successful (PLOS Computational Biology was), but all were learnt from.So, what did I learn?Here is at least part of my life's lesson, in that now familiar and comfortable Ten Simple Rules format.In this article, the rules are generic and can be considered beyond our own discipline. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2019 | Ten Simple Rules for avoiding and resolving conflicts with your colleaguesabstractDuring the course of our personal and professional lives, we spend a significant amount of time communicating with others.In fact, communication is one of the most important, but possibly also one of the hardest, things we do, having the power to bring individuals and communities together or create divisions.Getting it right is therefore crucial.Modern technologies have had a significant impact on the ways in which we are now able to communicate, allowing us to share our thoughts with colleagues, family, or friends at the click of a button.But communicating more quickly does not always result in better communication-the technologies we use often divorce us from the visual clues that are so crucial to understanding each other's true meaning and make it easy to misinterpret each other's real intentions.For this reason, our interactions can sometimes be unexpectedly difficult or can go unaccountably wrong.Given that communication is vital to the health and productivity of relationships, how can we best make our interactions work, and how can we resolve situations when they arise?The following are 10 simple rules based on our experience that we hope will help.Many of these rules can apply to the kinds of communication we may have with colleagues, family, or friends.However, we focus this article on the professional environment: we begin with suggestions to help avoid disagreements or to help stop them turning into serious conflicts; we then reflect on steps that might help to resolve situations that have become confrontational.Most interactions with colleagues are cordial and are working towards a common goal.Sometimes, however, because of differing views, misinterpretation of something said, or just because you're having a bad day, communications can go awry and become heated; from this point, without resolution, awkward situations can quickly escalate.Practicing effective communication skills before a confrontation arises, or during a confrontation, is the topic of this article.For more general ideas about engaging in successful collaborations, see [1].To delve further into the area of conflict management in the work environment, see [2,3].To keep this contribution manageable, we have confined ourselves to peer-to-peer communication and not considered a larger ecosystem of interactions in which conflict occurs.We feel that is a separate contribution that should be written. Rule 1: Always treat people with equality and respectWhether interacting with your peers or not, treat people courteously.Don't prejudge individuals based on their rank or perceived academic abilities, or worse, their gender, race, or sexual orientation.Be polite, and treat everyone equally and fairly. Fran Lewitter, Philip E. Bourne, Terri K. Attwood |
PLoS Comput. Biol. | 2 |
| 2018 | Ten simple rules when considering retirementabstractYou know a field is maturing when its early proponents think about retiring or, alas, pass away.So it is with computational biology.At 65, I think about retiring more.Not so much about retirement per se but in terms of what I want to accomplish before I retire and what retirement means to me in the first place.In other words, retirement is complex, and these ten rules probably (and hopefully) just start a discourse.For most scientists, including computational biologists, it's not a situation of now that you're 65, please accept your plaque, exit stage left, and goodbye.The questions I'm now considering are far more varied and nuanced.It's about where I want to focus my life and my energies.How much do I want to continue to mentor younger folk?How do I keep ties with colleagues I like and respect?How can I give back as much as possible to society and a profession that has treated me so well?The focus here is on retirement as an option, not a requirement.After all, emeritus status in academia, government, and industry can typically go on in some form indefinitely.It must also be said ahead of the rules themselves that in drafting the rules and having them reviewed by those acknowledged below, as well as discussing them with colleagues, friends, and family, that I was entering a very personal space.Efforts to coopt coauthors, which I like to do to provide a broader perspective, resulted in something to diffuse.How individuals think, or indeed choose not to think, about retirement is very personal.As such, the rules are more personal than I would normally write or have written on other subjects.Therefore, it may be that these rules do not resonate with you directly, but I hope they will at least make you think about what retirement means to you. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2018 | One thousand simple rulesabstractWhat began as a one-off in 2005 as Ten Simple Rules for Getting Published [1] has, in thirteen years, now multiplied a hundredfold to become One Thousand Simple Rules for many aspects of one's professional development and led to Quick Tips in the journal's Education section.This milestone of a thousand rules has been reached thanks to the unselfish work of all stakeholders-authors, editors, reviewers, and readers.Let's face it, writing, editing or reviewing a Ten Simple Rules (TSR) article is not the same as a publication that advances a scientific field.What it is going to get you is the satisfaction of knowing you have passed on a part of your experience in a form that is easily understood and acted upon by those following in your footsteps and hence have a very different kind of positive impact on science.Thank you.On the other hand, as a reader the TSRs might contribute in some small way to you getting tenure, or whatever else you care about in your professional life."How to write a review, by M Pautasso.It changed my approach to writing reviews."Anonymous, Academic faculty "Ten Simple Rules for Reproducible Computational Research, and other articles related to computational biology, programming, etc.They really helped in organizing my projects and code more efficiently." Philip E. Bourne, Fran Lewitter, Scott Markel, Jason A. Papin |
PLoS Comput. Biol. | 1 |
| 2018 | Cloud computing applications for biomedical science: A perspectiveabstractBiomedical research has become a digital data-intensive endeavor, relying on secure and scalable computing, storage, and network infrastructure, which has traditionally been purchased, supported, and maintained locally. For certain types of biomedical applications, cloud computing has emerged as an alternative to locally maintained traditional computing approaches. Cloud computing offers users pay-as-you-go access to services such as hardware infrastructure, platforms, and software for solving common biomedical computational problems. Cloud computing services offer secure on-demand storage and analysis and are differentiated from traditional high-performance computing by their rapid availability and scalability of services. As such, cloud services are engineered to address big data problems and enhance the likelihood of data and analytics sharing, reproducibility, and reuse. Here, we provide an introductory perspective on cloud computing to help the reader determine its value to their own research. Vivek Navale, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2017 | BioJava-ModFinder: identification of protein modifications in 3D structures from the Protein Data BankabstractSUMMARY: We developed a new software tool, BioJava-ModFinder, for identifying protein modifications observed in 3D structures archived in the Protein Data Bank (PDB). Information on more than 400 types of protein modifications were collected and curated from annotations in PDB, RESID, and PSI-MOD. We divided these modifications into three categories: modified residues, attachment modifications, and cross-links. We have developed a systematic method to identify these modifications in 3D protein structures. We have integrated this package with the RCSB PDB web application and added protein modification annotations to the sequence diagram and structure display. By scanning all 3D structures in the PDB using BioJava-ModFinder, we identified more than 30 000 structures with protein modifications, which can be searched, browsed, and visualized on the RCSB PDB website. AVAILABILITY AND IMPLEMENTATION: BioJava-ModFinder is available as open source (LGPL license) at ( https://github.com/biojava/biojava/tree/master/biojava-modfinder ). The RCSB PDB can be accessed at http://www.rcsb.org . CONTACT: [email protected]. Jianjiong Gao, Andreas Prlic, Chunxiao Bi, Wolfgang Bluhm, Dong Xu 0002, Philip E. Bourne, Peter W. Rose |
Bioinform. | 7 |
| 2017 | Ten simple rules in considering a career in academia versus governmentabstractThis article is focused on a career point at which a higher degree is in hand-perhaps along with some practical experience-and it is time to make a career decision.One such decision might be between an academic scientific research career versus a non-research career in government service.There are many other opportunities, of course, and industry versus academia has been well covered previously in this series [1].With federal research funding as limited as it is, early-career scientific researchers are increasingly looking at nonacademic pathways; government service is one option.An example choice might be between accepting a postdoctoral fellowship or tenure track assistant professorship versus becoming a program officer for a funding agency, working in government relations, or working in government policy development.Obviously, these are only a couple of the many career choices available in academia and government.These rules are meant to be as generic as possible by recognizing the broad similarities and differences that exist in the 2 work environments.The rules do not cover the obvious differences, such as the ability to teach in academia but likely not in government.As indicated, academic research and government service both cover large amounts of career territory.While trying to be as evenhanded as possible between these 2 career paths, undoubtedly, bias stemming from my own experience creeps in, and it is important to understand from where my perspective derives.I have spent most of my career in academia as both a professor and a university administrator.More recently, I spent 3 years in the United States federal government, where I had both an administrative and research role, both in biomedicine.My experience is far from that needed to provide a complete picture of career options.For example, it does not address government service, federal or state, outside of the US.Nor does it truly address the myriad of options outside of working for a government funding agency focused on biomedical research.More problematic is having worked 3 years in government versus over 40 years in academia.Undoubtedly, it is a different article than if I had spent 40 years in government service and 3 recent years in academia.Keeping in mind these limitations and the fact that I have been strongly influenced by the excellent reviews of the first version of this article, what follow are the rules I have to offer, rules which are made as generic as I know how.Remember also that career options are not for life, and experience in government can be very useful to furthering a career in academia and vice versa.This is something that I can attest to, and which I try to capture. Rule 1: Public good means different thingsAs an academic, I rarely thought about public good, defined as a commodity or service provided without profit to all members of society.Yes, I did my research with the idea of Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2017 | Ten simple rules to consider regarding preprint submissionabstractdescription of a body of scientific work that has yet to be published in a journal.Typically, a preprint is a research article, editorial, review, etc. that is ready to be submitted to a journal for peer review or is under review.It could also be a commentary, a report of negative results, a large data set and its description, and more.Finally, it could also be a paper that has been peer reviewed and either is awaiting formal publication by a journal or was rejected, but the authors are willing to make the content public.In short, a preprint is a research output that has not completed a typical publication pipeline but is of value to the community and deserving of being easily discovered and accessed.We also note that the term preprint is an anomaly, since there may not be a print version at all.The rules that follow relate to all these preprint types unless otherwise noted.In 1991, physics (and later, other disciplines, including mathematics, computer science, and quantitative biology) began a tradition of making preprints available through arXiv [1].arXiv currently contains well over 1 million preprints.While late to the game [2], the availability of preprints in biomedicine has gained significant community attention recently [3,4] and led to the formation of a scientist-driven effort, ASAPbio [5], to promote their use.As a result of an ASAPbio meeting held in February of 2016, a paper was published [6] that describes the pros and cons of preprints from the perspective of the stakeholders-scientists, publishers, and funders.Here, we formulate the message specifically for scientists in the form of ten simple rules for considering using preprints as a communication mechanism. Rule 1: Preprints speed up disseminationA recent analysis highlighted that the median review time-the time between submission and acceptance of an article-is around 100 days, with a further 25 days or so spent preparing the work for publication [7].However, these figures-slow as they are-do not include the time researchers spend "shopping around" for a journal to publish their findings, which can induce rounds of editorial rejection before or after peer review.Stephen Royle, a cell biologist at the University of Warwick, undertook an analysis of his published papers over the past dozen years and concluded that the average time from first submission to publication was around 9 months [8].Royle's is one example of a well-studied phenomenon [9].In summary, at a time when technology allows research findings to be shared instantly, the time to access research output appears glacial and similar to the pre-internet era. Philip E. Bourne, Jessica K. Polka, Ronald D. Vale, Robert Kiley |
PLoS Comput. Biol. | 1 |
| 2016 | Drug repurposing to target Ebola virus replication and virulence using structural systems pharmacologyabstractBACKGROUND: The recent outbreak of Ebola has been cited as the largest in history. Despite this global health crisis, few drugs are available to efficiently treat Ebola infections. Drug repurposing provides a potentially efficient solution to accelerating the development of therapeutic approaches in response to Ebola outbreak. To identify such candidates, we use an integrated structural systems pharmacology pipeline which combines proteome-scale ligand binding site comparison, protein-ligand docking, and Molecular Dynamics (MD) simulation. RESULTS: One thousand seven hundred and sixty-six FDA-approved drugs and 259 experimental drugs were screened to identify those with the potential to inhibit the replication and virulence of Ebola, and to determine the binding modes with their respective targets. Initial screening has identified a number of promising hits. Notably, Indinavir; an HIV protease inhibitor, may be effective in reducing the virulence of Ebola. Additionally, an antifungal (Sinefungin) and several anti-viral drugs (e.g. Maraviroc, Abacavir, Telbivudine, and Cidofovir) may inhibit Ebola RNA-directed RNA polymerase through targeting the MTase domain. CONCLUSIONS: Identification of safe drug candidates is a crucial first step toward the determination of timely and effective therapeutic approaches to address and mitigate the impact of the Ebola global crisis and future outbreaks of pathogenic diseases. Further in vitro and in vivo testing to evaluate the anti-Ebola activity of these drugs is warranted. Che L. Martin, Raymond Fan, Philip E. Bourne, Lei Xie 0006 |
BMC Bioinform. | 4 |
| 2015 | Big data in biomedicine - An NIH perspectiveabstractSummary form only given. Biomedical research is becoming increasingly data driven, analytical and hence digital. In recognition of this evolution NIH has established the Office for Data Science with trans NIH responsibility for maximizing the value of this digital enterprise. This effort brings together communities, policy changes and new infrastructure to be applied to existing and new areas of research such as precision medicine. We will review these changes from the perspective of research advances that are underway and highlight how this community can further engage in these activities. Philip E. Bourne |
BIBM | 1 |
| 2015 | Detection of circular permutations within protein structures using CE-CPabstractMOTIVATION: Circular permutation is an important type of protein rearrangement. Natural circular permutations have implications for protein function, stability and evolution. Artificial circular permutations have also been used for protein studies. However, such relationships are difficult to detect for many sequence and structure comparison algorithms and require special consideration. RESULTS: We developed a new algorithm, called Combinatorial Extension for Circular Permutations (CE-CP), which allows the structural comparison of circularly permuted proteins. CE-CP was designed to be user friendly and is integrated into the RCSB Protein Data Bank. It was tested on two collections of circularly permuted proteins. Pairwise alignments can be visualized both in a desktop application or on the web using Jmol and exported to other programs in a variety of formats. AVAILABILITY AND IMPLEMENTATION: The CE-CP algorithm can be accessed through the RCSB website at http://www.rcsb.org/pdb/workbench/workbench.do. Source code is available under the LGPL 2.1 as part of BioJava 3 (http://biojava.org; http://github.com/biojava/biojava). CONTACT: [email protected] or [email protected]. Spencer Bliven, Philip E. Bourne, Andreas Prlic |
Bioinform. | 2 |
| 2015 | RCSB PDB Mobile: iOS and Android mobile apps to provide data access and visualization to the RCSB Protein Data BankabstractSUMMARY: The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) resource provides tools for query, analysis and visualization of the 3D structures in the PDB archive. As the mobile Web is starting to surpass desktop and laptop usage, scientists and educators are beginning to integrate mobile devices into their research and teaching. In response, we have developed the RCSB PDB Mobile app for the iOS and Android mobile platforms to enable fast and convenient access to RCSB PDB data and services. Using the app, users from the general public to expert researchers can quickly search and visualize biomolecules, and add personal annotations via the RCSB PDB's integrated MyPDB service. AVAILABILITY AND IMPLEMENTATION: RCSB PDB Mobile is freely available from the Apple App Store and Google Play (http://www.rcsb.org). Greg B. Quinn, Chunxiao Bi, Cole H. Christie, Kyle Pang, Andreas Prlic, Takanori Nakane, Christine Zardecki, Maria Voigt, Helen M. Berman, Philip E. Bourne, Peter W. Rose |
Bioinform. | 10 |
| 2015 | Editor's Choice: Achievements and challenges in structural bioinformatics and computational biophysicsabstractMOTIVATION: The field of structural bioinformatics and computational biophysics has undergone a revolution in the last 10 years. Developments that are captured annually through the 3DSIG meeting, upon which this article reflects. RESULTS: An increase in the accessible data, computational resources and methodology has resulted in an increase in the size and resolution of studied systems and the complexity of the questions amenable to research. Concomitantly, the parameterization and efficiency of the methods have markedly improved along with their cross-validation with other computational and experimental results. CONCLUSION: The field exhibits an ever-increasing integration with biochemistry, biophysics and other disciplines. In this article, we discuss recent achievements along with current challenges within the field. Ilan Samish, Philip E. Bourne, Rafael Najmanovich |
Bioinform. | 2 |
| 2015 | The NIH Big Data to Knowledge (BD2K) initiativeabstractUnderstanding the human condition is a Big Data problem. This statement is nicely illustrated by the articles that follow from seven of the Centers for Data Excellence that have been funded by the National Institutes of Health (NIH) Big Data to Knowledge (BD2K) initiative. BD2K is a trans-NIH program, funded by all Institutes and Centers at NIH as well as the NIH Common Fund; it is overseen by the NIH Office of Data Science within the NIH Office of the Director. The beginnings of BD2K have been described previously, 1 and the purpose here is to provide an overall context, from the perspective of NIH, for the emerging program, as exemplified by the work of the Data Centers of Excellence (7 of 12 described here), the Data Discovery Index Coordinating Consortium, the various training awards, and the various individual investigator awards that have been made. What is aptly described by the seven Center articles is the research being performed across a rich array of data types and emergent infrastructure—metadata, analysis tools, frameworks, web resources, and more—that focus on problems inherent in extracting knowledge from large amounts of data with varying degrees of structure. Also detailed are plans to train researchers to make the most of the new opportunities presented by these data science advances. Recognizing that what motivates researchers is the desire to understand the human condition, the BD2K initiative was designed so that driving biological problems are central to the effort, but with the solutions to those problems being deliverables in the form of new methods, tools, software, and training. The deliverables that emerge from the Centers should be FAIR 2 —that is, contributing to the ability to F ind, A ccess, I nteroperate, and R euse the products of this research. BD2K aims to have these digital objects exist, not in isolation, but rather as part of an emergent ecosystem that is shared with the biomedical research community at large. To this end, we have introduced the notion of the Commons , a shared virtual space that conforms to the FAIR principles. The Commons allows digital objects to be stored and computed upon by a broad community once found by the emergent data discovery index being developed by the BioCADDIE group. 3 The Commons pilots that are under way are based on public cloud resources, but other compute and storage resources (such as high-performance computing facilities and institutional facilities) are expected to join the Commons as it develops. This assumes that the initial Commons pilots provide a cost-effective, sustainable, and usable environment. The primary aim of pilots now under way is to test this notion. If the pilots are successful, the BD2K Centers described herein and individual investigator grants will seed the Commons with data and software, whereupon we can monitor usage, which is an important for determining the value of this research output. We are at the beginnings of an exciting initiative, and the plans of the BD2K Centers are an excellent first step in advancing biomedicine in the era of Big Data. Philip E. Bourne, Vivien Bonazzi, Michelle Dunn, Eric D. Green, Mark Guyer, George A. Komatsoulis, Jennie Larkin, Beth Russell |
J. Am. Medical Informatics Assoc. | 1 |
| 2015 | Confronting the Ethical Challenges of Big Data in Public Healthabstractaccelerate much needed cures and improvements Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2015 | Ten Years of PLoS‡ Computational Biology: A Decade of Appreciation and InnovationabstractLibrary of Science, has evolved its own gloss from the piquantly camelesque "PLoS" to the modern " Philip E. Bourne, Steven E. Brenner, Michael B. Eisen |
PLoS Comput. Biol. | 1 |
| 2015 | Ten Simple Rules for Lifelong Learning, According to HammingabstractA mathematician, like a painter or poet, is a maker of patterns.If his patterns are more permanent than theirs, it is because they are made with ideas."-G.H. Hardy, "A Mathematician's Apology" [1] Learning is a lifelong imperative for any scientist, and Richard Hamming provided timeless advice on how to achieve this.In this sequel to our 2007 contribution to the Ten Simple Rules series [2], we attempt to distil the essence of what this mathematician and computer science and telecommunications pioneer addressed in one of his talks [3] and in his book The Art of Doing Science and Engineering: Learning to Learn [4].Hamming developed both the talk and the book as a synthesis of his graduate course in engineering at the United States Naval Postgraduate School in Monterey, California.We have organized his authoritative advice into ten rules.We believe these will equip the reader to more confidently face the unremitting emergence of an exponentially increasing amount of new knowledge, coupled with the equally relentless obsolescence of established knowledge, in a world containing a greater number of scientists than ever before.Our rules promote a certain "style of thinking."They also emphasize orientation towards the future and-we hope-will help the reader learn how to learn while motivating him or her to continue learning throughout life. Thomas C. Erren, Tracy E. Slanger, J. Valérie Groß, Philip E. Bourne, Paul Cullen |
PLoS Comput. Biol. | 4 |
| 2014 | What Big Data means to meabstractAcross the world, there is much talk about Big Data and how it is going to change the way we do business. Conferences, workshops, and funding initiatives in Big Data attract hundreds to thousands of stakeholders. Interestingly, no two stakeholders would be likely to define Big Data in exactly the same way. An obvious definition is the appropriate description, integration and sustainability of very large datasets generated by high throughput experiments. Another equally accurate and obvious definition is a large collection of small disparate, unstructured datasets which, taken together, can be analyzed to find unusual trends. I would like to offer a somewhat different viewpoint, which expands and encompasses these two definitions. For me, Big Data represents the emergence of the digital enterprise—the ability for an organization to take full advantage of its digital assets—which collectively can be described as large amounts of data and more. In other words, lots of data is a sign that something more is going on that has yet to be fully articulated and understood. Pertinent to this audience is the health center, the biomedical research institution, the commercial health-related entity, or the government agency as the ‘enterprise,’ and digital assets that range from basic research data, to electronic health records, to courses online, to proposals submitted, to administrative data, and so on. Today, which institutions would claim to take full advantage of their digital assets? The answer, I would suggest, is few, if any, yet the need to do so to remain competitive is becoming increasingly recognized. Hence the interest in Big Data. To be fair, the value and opportunities associated with digital assets take time to be appreciated in any type of enterprise. For example, until relatively recently, universities and their academic health centers were analog—patient records on paper, courses taught with slides or overheads, course notes printed, research data on shelves in notebooks, admission applications kept in endless filing cabinets, and so on. Now all of that content is, or soon will be, purely digital. The problem is that, for the most part, these data are simply electronic versions of what was maintained in hard copy. Healthcare centers and research institutions are only now beginning to leverage the true power of the digital medium. This time-to-adopt is not new in business; the music industry, the book and newspaper industry, the manufacturing industry, etc—and now the scholarly publishing industry—all responded slowly to the new digital reality, and when change accelerated, old business practices died and new ones emerged. Becoming a Digital Enterprise represents a challenge for many institutions since their organizational structures make widespread data integration and analysis difficult. Research, clinical activities, hospital services, education, and administrative services are somewhat siloed, and, in many organizations, each silo maintains its own separate organizational (and sometimes duplicated) data and information infrastructure. Central services may provide computer networking and email accounts, but the needs of hospitals, clinics, schools, departments, colleges, commercial partners, government agencies, or whatever are different. Breaking down the silos takes vision, leadership, and significant resources. The advantages of the truly digital enterprise are hard to imagine; that is just the point. Having said that, I imagine that you, like me, have said many times, “if only I could….” Here is an example of one of those situations where the stakeholders in the Digital Enterprise are Jane, an MD, PhD student; Jack a graduate student in chemistry; and Joe Smith, a neuroscience professor. Jane scores extremely well in parts of her graduate online neurology class. Neurology professors, whose research profiles are online and well described, are automatically notified of Jane's potential based on a computer analysis of her scores against the background interests of the neuroscience professors. Consequently, Professor Smith interviews Jane and offers her a research rotation. During the rotation, she enters details of her experiments related to understanding a widespread neurodegenerative disease in an online laboratory notebook kept in a shared online research space—an institutional resource where stakeholders provide metadata, including access rights and provenance beyond that available in a commercial offering. According to Jane's preferences, the underlying computer system may automatically bring to her attention Jack, a graduate student in the chemistry department whose notebook shows that he is working on using bacteria for purposes of toxic waste cleanup. Why the connection? They reference the same gene a number of times in their notes, which is of interest to two very different disciplines—neurology and environmental sciences. In the analog academic health center, they would never have discovered each other, but thanks to the Digital Enterprise, pooled knowledge can lead to a distinct advantage. The collaboration results in the discovery of a homologous human gene product as a putative target in treating the neurodegenerative disorder. A new chemical entity is developed and patented. Accordingly, by automatically matching details of the innovation with biotech companies worldwide that might have potential interest, a licensee is found. The licensee hires Jack to continue working on the project. Jane joins Professor Smith's laboratory, and he hires another student using the revenue from the license. The research continues and leads to a federal grant award. The students are employed, further research is supported, and, in time, societal benefit arises from the technology. A hypothetical example perhaps, but there are no technical reasons why this example cannot be realized—the discovery and analysis tools exist, and data can be sufficiently well described to make the needed connections. Missing are the organizational infrastructure, data sharing policies, tangible incentives to those who share their research products before publication, and the leaders to make this hypothetical example a reality. Continued interest in Big Data is the catalyst that takes us to a new and exciting place that comes to be known as the Digital Enterprise. It is our role to do our part in creating the environment in which it can become a reality. None. Commissioned; internally peer reviewed. Philip E. Bourne |
J. Am. Medical Informatics Assoc. | 1 |
| 2014 | Ten Simple Rules for Approaching a New JobabstractAt some point in your professional career, you will be faced with a job interview. This may range from visiting a graduate school where you already have a placement should you want it, to interviewing for a very high-profile position in industry, government, or academia where there is significant competition for that job. Thinking both as a job applicant and a job interviewer about how I have approached job situations over the years before, during, and after the interview and how those situations have turned out, I can offer the following ten simple rules as you prepare. Where appropriate, I conclude a rule with an illustrative scenario for a junior- and/or senior-level position since while the general principles are universal, how they are applied depends somewhat on the seniority of the position. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2014 | Ten Simple Rules for Writing a PLOS Ten Simple Rules ArticleabstractWhen I read the title of this article I laughed out loud-how many times has that happened to you when reading professional articles?Laughter is good whatever the context.When I started the series in 2005, I had no idea it would be so successful.This article, which I had no part in writing, only adding commentary shown in italics, is in my mind a celebration of that success.My commentary is simply to provide a historical perspective to explain some aspects of why the collection is the way it is and, of course, to make a few personal observations which, after all, is what the collection is meant for.Thanks to HD and AL for making this happen and for including me as an author (see Rule 4) and to all those that have contributed over the past nine years. Harriet Dashnow, Andrew Lonsdale, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2014 | Ten Simple Rules for Better FiguresabstractInternational audience Nicolas P. Rougier, Michael Droettboom, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2014 | Towards Structural Systems Pharmacology to Study Complex Diseases and Personalized MedicineabstractGenome-Wide Association Studies (GWAS), whole genome sequencing, and high-throughput omics techniques have generated vast amounts of genotypic and molecular phenotypic data. However, these data have not yet been fully explored to improve the effectiveness and efficiency of drug discovery, which continues along a one-drug-one-target-one-disease paradigm. As a partial consequence, both the cost to launch a new drug and the attrition rate are increasing. Systems pharmacology and pharmacogenomics are emerging to exploit the available data and potentially reverse this trend, but, as we argue here, more is needed. To understand the impact of genetic, epigenetic, and environmental factors on drug action, we must study the structural energetics and dynamics of molecular interactions in the context of the whole human genome and interactome. Such an approach requires an integrative modeling framework for drug action that leverages advances in data-driven statistical modeling and mechanism-based multiscale modeling and transforms heterogeneous data from GWAS, high-throughput sequencing, structural genomics, functional genomics, and chemical genomics into unified knowledge. This is not a small task, but, as reviewed here, progress is being made towards the final goal of personalized medicines for the treatment of complex diseases. Lei Xie 0006, Xiaoxia Ge, Hepan Tan, Li Xie 0002, Yinliang Zhang, Thomas Hart, Philip E. Bourne |
PLoS Comput. Biol. | 8 |
| 2013 | The era of openabstractWe are definitely in the open era. What does that mean and what does the future hold? I will provide a practitioners perspective on these questions as someone involved in running widely used biological databases, a producer of open source software and as founding editor in chief of an open access journal from the Public Library of Science (PLOS). In a nutshell it means profound change in the way educate, collaborate, disseminate and comprehend. I look forward to an open dialog as to more details of what that really means drawing on some examples from my own experiences. Philip E. Bourne |
OpenSym | 1 |
| 2013 | The reaming of life: based on the 2010 Jim Gray eScience Award LectureabstractSUMMARY We are well into the era of data intensive‐digital scientific discovery, an era defined by Jim Gray as the Fourth Paradigm. From my own perspective of the life sciences, much has been accomplished, but there is much to do if we are to maximize our understanding of biological systems given the data we have today, let alone what is coming. In my 2010 Jim Gray eScience Award Lecture, I gave my own thoughts on what needs to be accomplished, and with an additional year of hindsight, I expand on that here. Copyright © 2012 John Wiley & Sons, Ltd. Philip E. Bourne |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | Learning How to Run a Lab: Interviews with Principal InvestigatorsabstractScience could benefit from adopting management approaches, techniques, and expertise created in other fields of professional teamwork. The closest field to science is possibly software engineering and development. In its structure and spirit, a small software company very much resembles a scientific lab: It is a small motivated group of highly-educated professionals creating nonmaterial products with teamwork centered around collaboration between individuals. Science could benefit from using frameworks for project management, process control, issue tracking, and time management, which have been undergoing development in IT for decades, such as modern approaches of agile software development describing collaboration of self-organizing cross-functional teams. For a young PI at the beginning of his or Theodore Alexandrov, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2013 | Let's Make Those Book Chapters Open Too!abstractAs authors, many of us have had less than satisfactory experiences in writing book chapters as part of a themed volume or textbook that, when published, are expensive, inaccessible, and cited and used infrequently through lack of availability [1]. It could even be argued we write them from some sense of obligation and need, but do not put our best science and efforts into them because we know they won't be read and hence cited. Putting that thought aside, let's just say that a lot of good science goes underutilized. Some finds its way into journal reviews and journals that specialize in such content, for example the Elsevier Current Opinions series, but much languishes. I, many of the PLOS editors, and PLOS management have long wanted this situation to change; well, now it has.
The value of themed hardcover volumes and textbooks was understandable in a purely print era, an era during which we frequented the library more, which is where these volumes resided, being too expensive for individuals to purchase. These volumes make no sense today in a digital, open-access world. We are proud to report that PLOS Computational Biology has taken the first steps to address this nonsense. Translational Bioinformatics, edited by Guest Editor Maricel Kann and Education Editor Fran Lewitter, is the first complete PLOS “book” that can be accessed online as individual chapters or downloaded as a complete volume [2]. An ePub version is now available too. The content is indexed, each chapter has a Digital Object Identifier (doi) assigned and hence is resolvable (i.e., uniquely findable), and is indexed in PubMed and available as full text from PubMed Central, like all PLOS content.
Translational Bioinformatics can be used as a reference guide or textbook, and includes exercises. After review, the authors have had their hard work rewarded through a PLOS citation and greater accessibility to their work. The book uses the PLOS collection feature to bring the content together into a single entity while retaining the individuality of each article. While we regard this as an important step forward, it does raise some questions.
The first question is, who pays? Obviously, we are strong proponents of open access, but also the first to admit there must be a business model if open-access content is to be persistent. Certainly, most of us have never made any money from writing specialized book chapters contributed to a volume, but would we pay a modest amount to have them published under an open–access license? This remains an open question at this time. PLOS met the cost of publishing Translational Bioinformatics, but if this approach to book chapters were to take off, someone will have to foot the bill. At this time, PLOS is interested in furthering this cause, and, as with all front matter within the Education Section of the journal, book chapters are not subject to publishing fees. If the demand becomes too great we will need to revisit this. Support for chapter content would seem an opportunity for a wonderful contribution by an individual philanthropist or foundation in furthering scientific dissemination.
The second question is, what quality of review do we require of chapter content? In general terms, solicited book chapters do not undergo the level of review found in a research article. Unless the content is terrible, the editors soliciting the material are hard pressed to reject it, having persuaded the authors to write it in the first place. Good editors, as we have here, will provide the level of review found in a research article. Moreover, with greater exposure and article-level metrics (ALMs) applied to each chapter, the content will rise or fall on its own merits and the end result will likely be higher quality content than we have traditionally seen from book chapters. With regard to PLOS Computational Biology specifically, we regard this book content as front matter in the Education Section, and as such it has had significant review. We will be revisiting the issue of review, and indeed all aspects of the scope of our support for chapters, as demand increases.
The third question is, what do we lose and gain in an online book? Of course there are the obvious issues, now long debated, regarding ebooks versus physical books, and there is no need to revisit that here. What is worth visiting are the specific issues surrounding a book publication by PLOS, an organization that is currently set up to operate as a journal publisher. PLOS has been wonderful in making this project happen and paving the way for more through the notion of collections. Collections do present challenges when used to represent a book. For example, the collection has no ISBN or other book-like identifier defining the citation; books have editions, whereas there is no notion of versioning in a collection. On the positive side, the collection can become a dynamic entity, such that new chapters (with their own journal-like citation) can be added to the collection at any time.
Open questions there might be, but an exciting time nevertheless, with PLOS continuing to push the envelope regarding scholarly communication. Over time we will sort this out, but in the meantime, enjoy Translational Bioinformatics, a new innovation in open-access publishing. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2013 | Ten Simple Rules for Cultivating Open Science and Collaborative R&DabstractHow can we address the complexity and cost of applying science to societal challenges?
Open science and collaborative R&D may help [1]–[3]. Open science has been described as “a research accelerator” [4]. Open science implies open access [5] but goes beyond it: “Imagine a connected online web of scientific knowledge that integrates and connects data, computer code, chains of scientific reasoning, descriptions of open problems, and beyond …. tightly integrated with a scientific social web that directs scientists' attention where it is most valuable, releasing enormous collaborative potential.” [1].
Open science and collaborative approaches are often described as open source, by analogy with open-source software such as the operating system Linux which powers Google and Amazon—collaboratively created software which is free to use and adapt, and popular for Internet infrastructure and scientific research [6], [7]. However, this use of “open source” is unclear. Some people use “open source” when a project's results are free to use, others when a project's process is highly collaborative [4].
It is clearer to classify open source and open science within a broader class of collaborative R&D, which can be defined as scalable collaboration (usually enabled by information technology) across organizational boundaries to solve R&D challenges [8].
Many approaches to open science and collaborative R&D have been tried [1], [9]. The Gene Wiki has created over 10,000 Wikipedia articles, and aims to provide one for every notable human gene [10]. The crowdsourcing platform InnoCentive has reportedly facilitated solutions to roughly half of the thousands of technical problems posed on the site, including many in life sciences such as the $1 million ALS Biomarker Prize [11]. Other examples include prizes (X-Prize [12]), scientific games (FoldIt [13]), and licensing schemes inspired by open-source software (BIOS [14]).
Collaborative R&D approaches vary in openness [15]. In some approaches, the R&D process and outputs are open to all—for example, open-science projects like the Gene Wiki described above. In other approaches which demonstrate what might be called controlled collaboration, there are strong controls on who contributes and benefits—for example, computational platforms like Collaborative Drug Discovery or InnoCentive that support both commercial and nonprofit research [9], [11].
Collaborative approaches can unleash innovation from unforeseen sources, as with crowdsourcing health technologies [11]–[13], [16]. They may help in global challenges like drug development [17], as with India's OSDD (Open Source Drug Discovery) project that recruited over 7,000 volunteers [16] and an open-source drug synthesis project that improved an existing drug without increasing its cost [18].
If you want to apply open science and collaborative R&D, what principles are useful? We suggest Ten Simple Rules for Cultivating Open Science and Collaborative R&D. We also offer eight conversational interviews exploring life experiences that led to these rules (Box 1).
Box 1. Conversations on Open Science and Collaborative R&D
Many commentators have considered challenges in translating open science and collaborative methods to biomedical research [2]–[4], [9], [17], [20], [24], [26], [28], [29]. How can protecting intellectual property be balanced with freeing researchers to build on previous knowledge? If R&D results are collaboratively created and freely available, who will take responsibility for costly clinical trials and quality control? What will be the Linux of open-source R&D?
To explore such challenges and convey life experiences in biomedical open science and collaborative R&D, we offer eight conversational interviews by the first author of this article as supplementary material. The conversations were done on behalf of the Results for Development Institute and are with:
Alph Bingham, cofounder of InnoCentive (Text S1)
Barry Bunin, CEO of Collaborative Drug Discovery (Text S2)
Leslie Chan, open access pioneer and director of Bioline International (Text S3)
Aled Edwards, director of the Structural Genomics Consortium (Text S4)
Benjamin Good, coleader of the Gene Wiki initiative (Text S5)
Bernard Munos, pharmaceutical innovation thought leader (Text S6)
Zakir Thomas, director of India's Open Source Drug Discovery (OSDD) project (Text S7)
Matt Todd, open science and drug development pioneer (Text S8) Hassan Masum, Aarthi Rao, Benjamin M. Good, Matthew H. Todd, Aled M. Edwards, Leslie Chan, Barry A. Bunin, Andrew I. Su, Zakir Thomas, Philip E. Bourne |
PLoS Comput. Biol. | 10 |
| 2012 | BioJava: an open-source framework for bioinformatics in 2012abstractUNLABELLED: BioJava is an open-source project for processing of biological data in the Java programming language. We have recently released a new version (3.0.5), which is a major update to the code base that greatly extends its functionality. RESULTS: BioJava now consists of several independent modules that provide state-of-the-art tools for protein structure comparison, pairwise and multiple sequence alignments, working with DNA and protein sequences, analysis of amino acid properties, detection of protein modifications and prediction of disordered regions in proteins as well as parsers for common file formats using a biologically meaningful data model. AVAILABILITY: BioJava is an open-source project distributed under the Lesser GPL (LGPL). BioJava can be downloaded from the BioJava website (http://www.biojava.org). BioJava requires Java 1.6 or higher. All inquiries should be directed to the BioJava mailing lists. Details are available at http://biojava.org/wiki/BioJava:MailingLists. Andreas Prlic, Andy Yates, Spencer Bliven, Peter W. Rose, Julius O. B. Jacobsen, Peter V. Troshin, Mark Chapman, Jianjiong Gao, Chuan Hock Koh, Sylvain Foisy, Richard C. G. Holland, Gediminas Rimsa, Michael L. Heuer, Hannes Brandstätter-Müller, Philip E. Bourne, Scooter Willis |
Bioinform. | 15 |
| 2012 | Seven Years; It's Time for a ChangeabstractAfter seven years as the Editor-in-Chief of PLOS Computational Biology, I have decided to step to the side. It's time to bring in new leadership and a new vision. As scientists we generally do not learn a lot of management skills (a mistake in my opinion), but if I have learnt two management skills it is the value of enablement and to start planning for your successor on day one. Well, I did not start on day one, but Ruth Nussinov has been the Deputy Editor-in-Chief since October 2008, and she is an outstanding scientist and editor ideally suited to take over as editor-in-chief. So please welcome Ruth to this leadership role. The journal is in very safe hands. The journal staff, editors, and Ruth have not seen the last of me, however—this is just too much fun. As Founding Editor-in-Chief, I will continue to be involved with the journal, with a focus on special projects, helping Ruth and the team where I can, and, of course, continue as an author of both research articles and front matter, such as the Ten Simple Rules series.
In changing roles, I would like to make a few personal comments about the evolution of the journal and where it might go next. When Steven Brenner, Michael Eisen, and I founded the journal, we had a vision for how it would fill a gap between journals supporting purely computational methods and the array of experimental journals with the odd, token computational paper [1]. Thanks to you, the readers and authors, that vision has been realized beyond what we imagined, and I am very proud of how the journal is a voice for our broad and important community and at the same time helps build that community. The appearance of the journal in June 2005 was timely since it both propelled our field of science at a time when it was being recognized as a critical part of the life sciences, and made a strong statement about the importance of open access.
I must confess that when we started planning for the journal in 2004, support for open access seemed like the right thing to do, but it was not a major driver for me. That changed when, early on in my tenure, I realized open access is critical to maximizing the rate of scientific discovery. Supporting open access, and the new forms of scholarly communication it fosters, enriched my career and brought me into contact with many amazing people I would not otherwise have met. Emphasizing that enrichment, of the 22 invited lectures I gave in 2011, 18 were evangelizing about the importance of open access and open science, and four were directly related to my science. In my opinion, we have yet to see open access reach its full potential, but we will [2], and our journal will be poised to play an ever-increasing role as an exemplar for what is possible. In short, PLOS and this journal will continue to foster change, which maximizes the accessibility and comprehension of science. I am proud to continue to be a part of that.
It only remains for me to thank and acknowledge a variety of people who have all been so critically important in the past seven years. First and foremost are the approximately 150 editors who have worked tirelessly to shape and ensure a high-quality product over the years. There is no journal without community-driven efforts, which offer limited reward beyond a job well done. I can't mention everyone, but I must call out Karl Friston, who between 2005 and 2010 worked tirelessly to make computational neuroscience such a rich part of the journal, and Steven Brenner, Simon Levin, and Sebastian Bonhoeffer, who, since the journal's inception, have always responded with good advice.
Thanks are also due to Mark Patterson, who, until late 2011, was Director of Publishing at PLOS. Mark was critical to the success of the journal and to PLOS as a whole. We first began talking about the journal in May 2004, and he listened to, refined, and contributed to a lot of crazy ideas that define what the journal is today. To Catherine Nancarrow, who managed the journal from 2005 to 2010. A nicer and more dedicated person will not be found. In the current era of 140 character snapshots, her emails conferred a quality, beauty, and caring that you just don't see anymore. To Evie Browne, who contributed so much as Publications Assistant and Publications Manager between 2006 and 2009. Her spreadsheet analyses of how the journal was doing were amazing. To Andy Collings, who in various roles from Publications Assistant to Editorial Manager from 2005 to 2012 just made all ideas work, however outside the box they were. To Fran Lewitter, who, as Education Editor since the beginning, has created an important community jewel. To Scott Markel, who has been a voice of reason and an interface to the International Society of Computational Biology (ISCB) over the years. Finally, to all the other staff who have contributed over the past seven years: Emily Stevenson, Johanna Dehlinger, Helen Budd, Sheran Basra, and Cecy Marden, and the amazing current staff of Laura Taylor, Clare Weaver, and Chris Hall, all so well led by Rosemary Dickin and Theo Bloom.
PLOS is a family of journals and a family of people; in short, a family organization whose goal is to disseminate science in the most open and useful way possible. It is a successful organization able to recruit the best and most dedicated staff, and to adjust to its growing success and the success of open access itself. The world of scientific publishing will never be the same again because of PLOS and the people who drive it. I am proud to continue to be a member of this amazing family. There is no other publisher like PLOS and no other journal like PLOS Computational Biology. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2012 | A Review of 2011 for PLoS Computational BiologyabstractIn 2011, during discussions at various conferences, as well as informally with authors, readers, reviewers, and editors, we were struck by one resonating theme: the view that PLoS Computational Biology has helped to create a sense of community amongst a broad group of scientists and educators. While the journal labels itself as a PLoS “community” journal, if that label has any true meaning, it must come from the community itself. We feel that, after six years, we are indeed serving the community well, but as always you can disagree at any time, either publicly with a comment in response to this article or by email (gro.solp@loibpmocsolp).
That service comes first and foremost from the research we publish, but also from our desire to educate, report on open-source software, provide a history of the field, capture the vision of our editors, and move beyond the boundaries of traditional publishing to inform people within and outside of our community. Before we take a look at developments in each of these areas, and what is to come in 2012, let us first review how we served the community in 2011.
According to Google Analytics, 2011 saw over 553,000 unique visitors to our website and more than two million article views (not including access statistics from PubMed Central). Visitors came from 211 countries/territories, which was undoubtedly helped by the fact that the journal is open access. India, Spain, Russia, and Iran each showed over a 40% increase in visitors from the previous year. From the journal website, the most accessed Research Article was “Effect of Promoter Architecture on the Cell-to-Cell Variability in Gene Expression” by Sanchez et al. [1], published in March 2011 (8,954 views at the time of writing); the most accessed article overall was “Ten Simple Rules for Building and Maintaining a Scientific Reputation” by Bourne and Barbour [2], published in June 2011 (15,255 views at the time of writing).
Also in 2011, 1,623 research articles from 57 countries were submitted, up 16% from 2010, and 384 were published (down 2% from 2010). Receiving more but publishing about the same number in real terms should reflect the increasing quality of our content. We are very grateful to our Associate Editors, Guest Editors, reviewers (a list of Guest Editors and reviewers from 2011 is available in Table S1), and, of course, our Deputy Editors – Patricia Babbitt, Joel Bader, Sebastian Bonhoeffer, Lyle J. Graham, Konrad Kording, Douglas Lauffenburger, Uwe Ohler, Nathan Price, Burkhard Rost, Olaf Sporns, Wyeth Wasserman, and Weixiong Zhang – for helping us to handle this growth. With this growth, we have not met our goal of reducing the times to first decision, even with the addition of new editors, but we will continue to work on this in 2012. Our median decision before review time in 2011 was 8 days, and our median decision after review time was 47 days.
A number of our Research Articles were featured in blogs and the popular press. Notably, Mitra Hartman's paper on the morphology of the rat vibrissal array [3] was covered extensively, including two videos, by National Public Radio and Science Bytes.
Our Software section was launched in August 2011, and we have so far published one article, with six more either accepted or under review. Uptake has been relatively slow, based on, we believe, the open source and stringent documentation requirements we have imposed. We believe it is better to publish only a few, but high-quality, software articles, and that this will highlight the lack of rigor of software otherwise in the field.
Our Education section has continued to flourish, in part because of the journal's relationship with the International Society for Computational Biology (ISCB). This year we introduced a collection, Bioinformatics: Starting Early, which takes the notion of biology as a computational science into secondary schools. We are hoping for more articles from those involved in secondary teaching in 2012. Open science removes all boundaries not only to reading the latest science, but also to contributing to that science. We have even seen secondary school students as authors and expect to see more in the future.
In July 2011 we began the Editors Outlook series, with five published [4]–[8] and more on the way. These mini-reviews already broach subjects from ontologies to genome organization, and from evolution to data and privacy. They speak to the breadth of our field and editorial board, and collectively will form a vision from our many expert editors of what is being, and will be, accomplished in the coming years.
That our journal is fully open access provides opportunities for maximizing the use and reuse of our scholarship; we intend to explore this further in 2012. Early in 2012 we will launch our first Topic Page on circular permutations in proteins. Wikipedia is a valuable resource for knowledge dissemination, yet Wikipedia pages are lacking in coverage of computational biology. In part this is because authors gain little career-based reward for creating Wikipedia pages. We aim to bridge the gap. Topic Page articles, which will be published in the journal and will each receive a PubMed identifier and DOI, will become the copy of record, thereby crediting the author(s). At the same time the Topic Page will be used to seed a Wikipedia article and become a living version of the same material–a viable option thanks to our Creative Commons license. Look for an announcement of this development in the new year, but in the interim if you have ideas for Topic Pages you would like to contribute, please do get in touch for further information (gro.solp@loibpmocsolp).
We are also contemplating a new article type: Data Pages. Data Pages would be brief publications about datasets, in which the data are not already well described in other papers yet are considered of great value to the community. Such brief publications would bring a traditional reward to the producers of these shared datasets. Which is more valuable: a dataset downloaded and used by 100 investigators, who in turn publish research based on these data, or a paper that is cited only by the authors who wrote it? Data Pages would, from our point of view, help to answer this question.
If you want to provide feedback on our plans for Data Pages later in 2012, please do so by commenting on this article. Feel free to comment in public or to us privately on anything we are doing, or ideas that you have for the future of the journal. After all, PLoS Computational Biology is a community journal, and if you have read this far, you should consider yourself an important part of our ever-broadening community. Rosemary Dickin, Chris James Hall, Laura K. Taylor, Andrew M. Collings, Ruth Nussinov, Philip E. Bourne |
PLoS Comput. Biol. | 6 |
| 2012 | Ten Simple Rules for Starting a CompanyabstractMany faculty, staff, and students at academic institutions think about starting companies at some point in their careers. As academic funding models change, and how academia views entrepreneurial activity changes, starting companies is likely to happen more frequently. Hence, it is worth considering Ten Simple Rules to contemplate when starting a company while in academia. There is a wealth of general information out there to help you, but that information is not aimed specifically towards computational biologists. What follows is a hybrid that is intended as a general quick review for anyone, intermingled with some specific advice for computational biologists.
By way of experience, we should say we have been involved in starting several companies: both in the biomedical sciences, dealing with biological software, computational biology services, and currently SciVee Inc. (http://www.scivee.tv) distributing scientific rich media, and outside it, ranging from the distribution of independent films, to a socially oriented dining club aimed at supporting local businesses, to a business supplying art quality photographic prints. None have been a great financial success, but all have been immense fun, an opportunity to meet interesting people, and an opportunity to think quite differently than when doing scientific research. Read that as a personal endorsement to go for it, even if starting a company is not yet a well-formed idea, and even though, as you will see below, the rules themselves might be cautionary and off-putting. Anthony C. Fletcher, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2012 | Ten Simple Rules To Commercialize Scientific ResearchabstractCommercializing scientific research or a breakthrough idea is really no different, in principle, from commercializing anything, except perhaps that it's more difficult in practice because of the steps required to turn basic research into something practical and because you are looking for a market for a product, rather than designing a product to fit an established, or obvious market.
Commercialization is different to starting and running a company, a broader endeavour and the subject of a previous Ten Simple Rules article [1]. Even so, commercialization can be a broad endeavour. For example, at one extreme, you could hand over your monoclonal antibody to Sigma to supply it on your behalf to other researchers who might find it useful while the company pays you a small royalty; on the other, you could be involved in developing Herceptin (anti-HER2 monoclonal antibody) from its origins as a mouse-specific antibody through to its use as an effective anti-breast cancer drug, in a process that took more than decade. Here we assume the former—others are carrying out that commercialization, which has its pluses and minuses—less work for you, but typically less control of the commercialization process.
Commercialization is a much studied subject, both by academics [2] and the business community [3]. All larger academic institutions generally have offices to promote and help scientists get research to market. Consequently, in this Ten Simple Rules article we won't deal with the details, but instead will concentrate on some of the key issues to consider when working with, or before and after working with, a specialized office. Anthony C. Fletcher, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2012 | Ten Simple Rules to Protect Your Intellectual PropertyabstractThe concepts that underpin the protection of ideas and inventions are not new; such laws have been around for several hundred years and are discussed under the broad heading of intellectual property (IP). IP is easily misunderstood, but at the same time most scientists encounter it at some point in their career, as it is a necessary feature in the commercialization of research.
The term intellectual property includes such concepts and rights as copyright, trademarks, industrial design rights, and patents. It is important to remember that IP is a tool to help your endeavours, and not a goal in itself. Having IP for its own sake is pointless. IP can be crucial in commercializing research and running a successful science-based business, but having a patent and having a successful patented product are two very different things.
Above all, IP can only work for you if you understand what it is, why you want it, and what you are going to do with it. These ten simple rules are intended to provide an overview of these issues; however, we must start with a warning. Laws relating to IP change all the time, they are complex, sometimes rather obscure, and are very different from country to country. For example, research surrounding methods of treatment by surgery and therapy and diagnostic methods are patentable in the United States, but specifically excluded from patentability in Europe [1]. However, these boundaries seem to be shifting in both the US and Europe. In short, we are dealing with a complex and changing subject and restrict ourselves here to the guiding principles. Mark Jolly, Anthony C. Fletcher, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2012 | Topic Pages: PLoS Computational Biology Meets WikipediaabstractWhile there has been much debate about the coverage and quality of Wikipedia (starting with an article in 2005 [1]), there is no doubt about its value (and increasing role) as a reference source and starting point for in-depth research. For example, within the biomedical sciences, there have been recent articles about the accuracy and completeness of drug information in Wikipedia [2], Wikipedia as a source of information in nursing care [3] and mental disorders [4], and making biological databases available through Wikipedia [5]. Shoshana J. Wodak, Daniel Mietchen, Andrew M. Collings, Robert B. Russell, Philip E. Bourne |
PLoS Comput. Biol. | 5 |
| 2011 | Cobweb: a Java applet for network exploration and visualisationabstractAbstract Summary: Cobweb is a Java applet for real-time network visualization; its strength lies in enabling the interactive exploration of networks. Therefore, it allows new nodes to be interactively added to a network by querying a database on a server. The network constantly rearranges to provide the most meaningful topological view. Availability: Cobweb is available under the GPLv3 and may be freely downloaded at http://bioinformatics.charite.de/cobweb. Contact: [email protected] Supplementary Information: Supplementary data are available at Bioinformatics online. Joachim von Eichborn, Philip E. Bourne, Robert Preissner |
Bioinform. | 2 |
| 2011 | Ten Simple Rules for Getting Ahead as a Computational Biologist in AcademiaabstractGetting a promotion or a new position are important parts of the scientific career process. Ironically, a committee whose membership has limited ability to truly judge your scholarly standing is often charged with making these decisions. Here are ten simple rules from my own experiences, in both getting promoted and serving on such committees, for how you might maximize your chances of getting ahead under such circumstances. The rules focus on what might be added to a CV, research statement, personal statement, or cover letter, depending on the format of the requested promotion materials. In part, the rules suggest that you educate the committee members, who have a range of expertise, on what they should find important in the promotion application provided by a computational biologist. Further, while some rules are generally applicable, the focus here is on promotion in an academic setting. Having said that, in such a setting teaching and community service are obviously important, but barely touched upon here. Rather, the focus is on how to maximize the appreciation of your research-related activities. As a final thought before we get started on the rules, this is not just about you, but an opportunity to educate a broad committee on what is important in our field. Use that opportunity well, for it will serve future generations of computational biologists. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2011 | Ten Simple Rules for Building and Maintaining a Scientific ReputationabstractWhile we cannot articulate exactly what defines the less quantitative side of a scientific reputation, we might be able to seed a discussion. We invite you to crowdsource a better description and path to achieving such a reputation by using the comments feature associated with this article.Consider yourself challenged to contribute. Philip E. Bourne, Virginia Barbour |
PLoS Comput. Biol. | 1 |
| 2011 | A Review of 2010 for PLoS Computational BiologyabstractComputational Biology celebrated its fifth anniversary in 2010, and all in our community, either as readers, authors, or editors, should take pride in what has been accomplished in such a short space of time.In the past year we received 1,403 new Research Articles, a 295% increase from our first year of operation in 2005-2006 and a 17% increase over 2009.Of the articles submitted in 2010, 875 (62%) were rejected, and 70% of these were before review.We have seen growth not only in submissions, but in readership as well.Currently, around 16,000 readers receive the electronic table of contents, a 14% increase over the previous year.We published 392 Research Articles this year, along with 23 ''front section'' articles (Reviews, Perspectives, Education), down from 33 in the previous year.Eighty Associate Editors handled the combined submissions, with a total of 26 new editors joining this past year and six departing.We are proud to say that virtually every editor we asked to join accepted, a testament to how our community values the journal.These editors worked with more than 180 guest editors and 1,800 reviewers to handle the submissions, and we are of course very grateful for their support (Table S1). Rosemary Dickin, Cecy Marden, Andrew M. Collings, Ruth Nussinov, Philip E. Bourne |
PLoS Comput. Biol. | 5 |
| 2011 | Teaching Bioinformatics at the Secondary School LevelabstractBioinformatics is now an integral part of biology and biological research. The field began with a few people from other disciplines teaching themselves and each other the techniques that are now considered commonplace. These pioneers then began graduate programs [1]–[3] to educate the next generation. Those early graduate students typically came as bench biologists or as computer scientists, both groups requiring significant time to “hybridize”. Not surprisingly, this then led to undergraduate majors in bioinformatics to better prepare students for graduate school and research careers in bioinformatics. In addition, teaching bioinformatics in undergraduate biology classes is also a priority [4], [5]. Through the Education section of PLoS Computational Biology we have tried to support this evolution through a collection of educational articles pertinent to the undergraduate level and beyond. It is only natural that we would take the next step [6].
We now introduce a subsection of the Education section with articles devoted to teaching bioinformatics in secondary schools that is derived from the work of the Education committee of the International Society for Computational Biology (ISCB), who identified a need to address the issue of incorporating bioinformatics into secondary school biology classes. They also recognized the interest among researchers to build and participate in outreach programs at the secondary school level given that many funding agencies worldwide encourage such a component in grant applications.
To move the ball forward on secondary school bioinformatics education, at ISCB's 2010 international conference, Intelligent Systems in Molecular Biology (ISMB), the ISCB Education committee organized a half-day tutorial aimed at secondary school biology and chemistry teachers in the Boston area interested in learning about bioinformatics and how to include it in their curricula. The tutorial also attracted researchers involved in organizing or formulating outreach programs in their community. The main focus of the ISMB tutorial was the presentation of lesson plans by a secondary school teacher (David Form, a biology teacher at Nashoba Regional High School, Bolton, Massachusetts) who has successfully incorporated bioinformatics into his courses for more than five years. His is one example of such an effort and is embraced in the Ten Simple Rules and its supplementary material found in this issue. Also in this issue we have an article by Suzanne Gallagher and colleagues on the experience of teaching secondary school level bioinformatics in Boulder, Colorado.
There are many examples of outreach efforts to high school students that we would like to feature in coming months, which incorporate bioinformatics into their programs (see Table 1).
Table 1
Examples of Online Resources and Outreach Programs.
There are many other examples of educators doing similar work in school districts worldwide. A recent issue of Briefings in Bioinformatics was dedicated to bioinformatics education [7] with a specific example of programs for secondary school students [8], [9].The ISCB Education committee is building a resource of information useful to secondary school teachers who would like to incorporate bioinformatics into their curriculum. In addition, the committee has begun to explore how to include bioinformatics in Advanced Placement courses and exams in the United States, which we also hope to feature in the Education section of the journal.
We encourage feedback of any form, including comments on this editorial, and hearing about your experience teaching bioinformatics to secondary school students. Fran Lewitter, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2011 | Drug Discovery Using Chemical Systems Biology: Weak Inhibition of Multiple Kinases May Contribute to the Anti-Cancer Effect of NelfinavirabstractNelfinavir is a potent HIV-protease inhibitor with pleiotropic effects in cancer cells. Experimental studies connect its anti-cancer effects to the suppression of the Akt signaling pathway, but the actual molecular targets remain unknown. Using a structural proteome-wide off-target pipeline, which integrates molecular dynamics simulation and MM/GBSA free energy calculations with ligand binding site comparison and biological network analysis, we identified putative human off-targets of Nelfinavir and analyzed the impact on the associated biological processes. Our results suggest that Nelfinavir is able to inhibit multiple members of the protein kinase-like superfamily, which are involved in the regulation of cellular processes vital for carcinogenesis and metastasis. The computational predictions are supported by kinase activity assays and are consistent with existing experimental and clinical evidence. This finding provides a molecular basis to explain the broad-spectrum anti-cancer effect of Nelfinavir and presents opportunities to optimize the drug as a targeted polypharmacology agent. Li Xie 0002, Thomas Evangelidis, Lei Xie 0006, Philip E. Bourne |
PLoS Comput. Biol. | 4 |
| 2010 | Pre-calculated protein structure alignments at the RCSB PDB websiteabstractSUMMARY: With the continuous growth of the RCSB Protein Data Bank (PDB), providing an up-to-date systematic structure comparison of all protein structures poses an ever growing challenge. Here, we present a comparison tool for calculating both 1D protein sequence and 3D protein structure alignments. This tool supports various applications at the RCSB PDB website. First, a structure alignment web service calculates pairwise alignments. Second, a stand-alone application runs alignments locally and visualizes the results. Third, pre-calculated 3D structure comparisons for the whole PDB are provided and updated on a weekly basis. These three applications allow users to discover novel relationships between proteins available either at the RCSB PDB or provided by the user. AVAILABILITY AND IMPLEMENTATION: A web user interface is available at http://www.rcsb.org/pdb/workbench/workbench.do. The source code is available under the LGPL license from http://www.biojava.org. A source bundle, prepared for local execution, is available from http://source.rcsb.org CONTACT: [email protected]; [email protected]. Andreas Prlic, Spencer Bliven, Peter W. Rose, Wolfgang Bluhm, Chris Bizon, Adam Godzik, Philip E. Bourne |
Bioinform. | 7 |
| 2010 | dConsensus: a tool for displaying domain assignments by multiple structure-based algorithms and for construction of a consensus assignmentabstractBACKGROUND: Partitioning of a protein into structural components, known as domains, is an important initial step in protein classification and for functional and evolutionary studies. While the systematic assignments of domains by human experts exist (CATH and SCOP), the introduction of high throughput technologies for structure determination threatens to overwhelm expert approaches. A variety of algorithmic methods have been developed to expedite this process, allowing almost instant structural decomposition into domains. The performance of algorithmic methods can approach 85% agreement on the number of domains with the consensus reached by experts. However, each algorithm takes a somewhat different conceptual approach, each with unique strengths and weaknesses. Currently there is no simple way to automatically compare assignments from different structure-based domain assignment methods, thereby providing a comprehensive understanding of possible structure partitioning as well as providing some insight into the tendencies of particular algorithms. Most importantly, a consensus assignment drawn from multiple assignment methods can provide a singular and presumably more accurate view. RESULTS: We introduce dConsensus http://pdomains.sdsc.edu/dConsensus; a web resource that displays the results of calculations from multiple algorithmic methods and generates a domain assignment consensus with an associated reliability score. Domain assignments from seven structure-based algorithms - PDP, PUU, DomainParser2, NCBI method, DHcL, DDomains and Dodis are available for analysis and comparison alongside assignments made by expert methods. The assignments are available for all protein chains in the Protein Data Bank (PDB). A consensus domain assignment is built by either allowing each algorithm to contribute equally (simple approach) or by weighting the contribution of each method by its prior performance and observed tendencies. An analysis of secondary structure around domain and fragment boundaries is also available for display and further analysis. CONCLUSION: dConsensus provides a comprehensive assignment of protein domains. For the first time, seven algorithmic methods are brought together with no need to access each method separately via a webserver or local copy of the software. This aggregation permits a consensus domain assignment to be computed. Comparison viewing of the consensus and choice methods provides the user with insights into the fundamental units of protein structure so important to the study of evolutionary and functional relationships. Kieran Alden, Stella Veretnik, Philip E. Bourne |
BMC Bioinform. | 3 |
| 2010 | Word add-in for ontology recognition: semantic enrichment of scientific literatureabstractBACKGROUND: In the current era of scientific research, efficient communication of information is paramount. As such, the nature of scholarly and scientific communication is changing; cyberinfrastructure is now absolutely necessary and new media are allowing information and knowledge to be more interactive and immediate. One approach to making knowledge more accessible is the addition of machine-readable semantic data to scholarly articles. RESULTS: The Word add-in presented here will assist authors in this effort by automatically recognizing and highlighting words or phrases that are likely information-rich, allowing authors to associate semantic data with those words or phrases, and to embed that data in the document as XML. The add-in and source code are publicly available at http://www.codeplex.com/UCSDBioLit. CONCLUSIONS: The Word add-in for ontology term recognition makes it possible for an author to add semantic data to a document as it is being written and it encodes these data using XML tags that are effectively a standard in life sciences literature. Allowing authors to mark-up their own work will help increase the amount and quality of machine-readable literature metadata. J. Lynn Fink, Pablo Fernicola, Rahul Chandran, Savas Parastatidis, Alex D. Wade, Oscar Naim, Greg B. Quinn, Philip E. Bourne |
BMC Bioinform. | 8 |
| 2010 | Integration of open access literature into the RCSB Protein Data Bank using BioLitabstractBACKGROUND: Biological data have traditionally been stored and made publicly available through a variety of on-line databases, whereas biological knowledge has traditionally been found in the printed literature. With journals now on-line and providing an increasing amount of open access content, often free of copyright restriction, this distinction between database and literature is blurring. To exploit this opportunity we present the integration of open access literature with the RCSB Protein Data Bank (PDB). RESULTS: BioLit provides an enhanced view of articles with markup of semantic data and links to biological databases, based on the content of the article. For example, words matching to existing biological ontologies are highlighted and database identifiers are linked to their database of origin. Among other functions, it identifies PDB IDs that are mentioned in the open access literature, by parsing the full text for all research articles in PubMed Central (PMC) and exposing the results as simple XML Web Services. Here, we integrate BioLit results with the RCSB PDB website by using these services to find PDB IDs that are mentioned in research articles and subsequently retrieving abstract, figures, and text excerpts for those articles. A new RCSB PDB literature view permits browsing through the figures and abstracts of the articles that mention a given structure. The BioLit Web Services that are providing the underlying data are publicly accessible. A client library is provided that supports querying these services (Java). CONCLUSIONS: The integration between literature and websites, as demonstrated here with the RCSB PDB, provides a broader view for how a given structure has been analyzed and used. This approach detects the mention of a PDB structure even if it is not formally cited in the paper. Other structures related through the same literature references can also be identified, possibly providing new scientific insight. To our knowledge this is the first time that database and literature have been integrated in this way and it speaks to the opportunities afforded by open and free access to both database and literature content. Andreas Prlic, Marco A. Martinez, Bojan Beran, Benjamin T. Yukich, Peter W. Rose, Philip E. Bourne, J. Lynn Fink |
BMC Bioinform. | 7 |
| 2010 | What Do I Want from the Publisher of the Future?abstractWhen I took on the role of Editor-in-Chief of this open-access journal, I began, for the first time, to think about scholarly communication beyond submitting my papers and getting them published.This thinking led to previous Perspectives [1-3], all of which shared an underlying themethere are many opportunities to achieve better dissemination and comprehension of our science, and as producers of that output I believe authors have a responsibility to see it used in the best possible way.No need to take my word regarding the opportunities that exist to improve scholarly communication and comprehension.I recommend reading ''Part 4: Scholarly Communication'' from the free online book the Fourth Paradigm: Data Intensive Scientific Discovery [4] (http://research.microsoft.com/en- Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2010 | Will Widgets and Semantic Tagging Change Computational Biology?abstractWe argue here, through the use of several examples from our work in support of structural biology, that the answer to the question posed by the title of this Perspective is a resounding yes. The discussion that follows is aimed primarily at those of the journal's readers who are biological resource developers and Web page developers interested in developing the richest possible Web pages. However, those of you who simply use biological resources might find this a helpful discussion in understanding what is on the horizon. Whatever your interest, please let us hear your opinion on the question posed by this Perspective through the associated comment feature. Philip E. Bourne, Bojan Beran, Chunxiao Bi, Wolfgang Bluhm, Roland L. Dunbrack Jr., Andreas Prlic, Greg B. Quinn, Peter W. Rose, Raship Shah, Wendy Tao, Brian D. Weitzner, Benjamin T. Yukich |
PLoS Comput. Biol. | 1 |
| 2010 | Drug Off-Target Effects Predicted Using Structural Analysis in the Context of a Metabolic Network ModelabstractRecent advances in structural bioinformatics have enabled the prediction of protein-drug off-targets based on their ligand binding sites. Concurrent developments in systems biology allow for prediction of the functional effects of system perturbations using large-scale network models. Integration of these two capabilities provides a framework for evaluating metabolic drug response phenotypes in silico. This combined approach was applied to investigate the hypertensive side effect of the cholesteryl ester transfer protein inhibitor torcetrapib in the context of human renal function. A metabolic kidney model was generated in which to simulate drug treatment. Causal drug off-targets were predicted that have previously been observed to impact renal function in gene-deficient patients and may play a role in the adverse side effects observed in clinical trials. Genetic risk factors for drug treatment were also predicted that correspond to both characterized and unknown renal metabolic disorders as well as cryptic genetic deficiencies that are not expected to exhibit a renal disorder phenotype except under drug treatment. This study represents a novel integration of structural and systems biology and a first step towards computational systems medicine. The methodology introduced herein has important implications for drug development and personalized medicine. Roger L. Chang, Li Xie 0002, Lei Xie 0006, Philip E. Bourne, Bernhard O. Palsson |
PLoS Comput. Biol. | 4 |
| 2010 | PLoS Computational Biology Conference Postcards from PSB 2010abstractDOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone. Ruchira S. Datta, Matthew W. Lux, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2010 | A Review of 2009 for PLoS Computational Biology
Rosemary Dickin, Cecy Marden, Catherine Nancarrow, Philip E. Bourne |
PLoS Comput. Biol. | 4 |
| 2010 | A Multidimensional Strategy to Detect Polypharmacological Targets in the Absence of Structural and Sequence HomologyabstractConventional drug design embraces the "one gene, one drug, one disease" philosophy. Polypharmacology, which focuses on multi-target drugs, has emerged as a new paradigm in drug discovery. The rational design of drugs that act via polypharmacological mechanisms can produce compounds that exhibit increased therapeutic potency and against which resistance is less likely to develop. Additionally, identifying multiple protein targets is also critical for side-effect prediction. One third of potential therapeutic compounds fail in clinical trials or are later removed from the market due to unacceptable side effects often caused by off-target binding. In the current work, we introduce a multidimensional strategy for the identification of secondary targets of known small-molecule inhibitors in the absence of global structural and sequence homology with the primary target protein. To demonstrate the utility of the strategy, we identify several targets of 4,5-dihydroxy-3-(1-naphthyldiazenyl)-2,7-naphthalenedisulfonic acid, a known micromolar inhibitor of Trypanosoma brucei RNA editing ligase 1. As it is capable of identifying potential secondary targets, the strategy described here may play a useful role in future efforts to reduce drug side effects and/or to increase polypharmacology. Jacob D. Durrant, Rommie E. Amaro, Lei Xie 0006, Michael D. Urbaniak, Michael A. J. Ferguson, Antti Haapalainen, Anne Marie Di Guilmi, Frank Wunder, Philip E. Bourne, James Andrew McCammon |
PLoS Comput. Biol. | 10 |
| 2010 | The Mycobacterium tuberculosis Drugome and Its Polypharmacological ImplicationsabstractWe report a computational approach that integrates structural bioinformatics, molecular modelling and systems biology to construct a drug-target network on a structural proteome-wide scale. The approach has been applied to the genome of Mycobacterium tuberculosis (M.tb), the causative agent of one of today's most widely spread infectious diseases. The resulting drug-target interaction network for all structurally characterized approved drugs bound to putative M.tb receptors, we refer to as the 'TB-drugome'. The TB-drugome reveals that approximately one-third of the drugs examined have the potential to be repositioned to treat tuberculosis and that many currently unexploited M.tb receptors may be chemically druggable and could serve as novel anti-tubercular targets. Furthermore, a detailed analysis of the TB-drugome has shed new light on the controversial issues surrounding drug-target networks [1]-[3]. Indeed, our results support the idea that drug-target networks are inherently modular, and further that any observed randomness is mainly caused by biased target coverage. The TB-drugome (http://funsite.sdsc.edu/drugome/TB) has the potential to be a valuable resource in the development of safe and efficient anti-tubercular drugs. More generally the methodology may be applied to other pathogens of interest with results improving as more of their structural proteomes are determined through the continued efforts of structural biology/genomics. Sarah L. Kinnings, Li Xie 0002, Kingston H. Fung, Richard M. Jackson, Lei Xie 0006, Philip E. Bourne |
PLoS Comput. Biol. | 6 |
| 2009 | A unified statistical model to support local sequence order independent similarity searching for ligand-binding sites and its application to genome-based drug discoveryabstractFunctional relationships between proteins that do not share global structure similarity can be established by detecting their ligand-binding-site similarity. For a large-scale comparison, it is critical to accurately and efficiently assess the statistical significance of this similarity. Here, we report an efficient statistical model that supports local sequence order independent ligand-binding-site similarity searching. Most existing statistical models only take into account the matching vertices between two sites that are defined by a fixed number of points. In reality, the boundary of the binding site is not known or is dependent on the bound ligand making these approaches limited. To address these shortcomings and to perform binding-site mapping on a genome-wide scale, we developed a sequence-order independent profile-profile alignment (SOIPPA) algorithm that is able to detect local similarity between unknown binding sites a priori. The SOIPPA scoring integrates geometric, evolutionary and physical information into a unified framework. However, this imposes a significant challenge in assessing the statistical significance of the similarity because the conventional probability model that is based on fixed-point matching cannot be applied. Here we find that scores for binding-site matching by SOIPPA follow an extreme value distribution (EVD). Benchmark studies show that the EVD model performs at least two-orders faster and is more accurate than the non-parametric statistical method in the previous SOIPPA version. Efficient statistical analysis makes it possible to apply SOIPPA to genome-based drug discovery. Consequently, we have applied the approach to the structural genome of Mycobacterium tuberculosis to construct a protein-ligand interaction network. The network reveals highly connected proteins, which represent suitable targets for promiscuous drugs. Lei Xie 0006, Li Xie 0002, Philip E. Bourne |
Bioinform. | 3 |
| 2009 | Ten Simple Rules for Chairing a Scientific SessionabstractChairing a session at a scientific conference is a thankless task.If you get it right, no one is likely to notice.But there are many ways to get it wrong and a little preparation goes a long way to making the session a success.Here are a few pointers that we have picked up over the years. Alex Bateman, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2009 | A Review of 2008 for PLoS Computational BiologyabstractComputational Biology; which saw 50% more submissions than in 2007 (900 full articles and 175 presubmission inquiries), more than 260 high-quality research articles published, and regular contributions of Editorials, Reviews and Perspectives, and Education and Society pages.This growth and maturity of content leaves no doubt that our Journal has become a leading reference for the field of computational biology and a trusted place to publish.Such success has come through the hard work of our Editors, not only from our Editorial Board but also from the anonymous reviewers and Guest Editors who expend so much time and energy in the assessment of submitted manuscripts (each averaging 2.8 reviews and 1 to 2 rounds of revisions), and from the attention to detail and care taken over the content.Peer review by external experts is essential to ensuring that the work published in PLoS Computational Biology is of the very highest quality, and we are grateful to all of our reviewers for their thoughtful and informed comments.Guest Editors are those who step in to edit one particular paper that describes work in an area of research that falls outside the expertise of the more than 50 volunteer Editors on our Board. Evie Browne, Rosemary Dickin, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2009 | Drug Discovery Using Chemical Systems Biology: Repositioning the Safe Medicine Comtan to Treat Multi-Drug and Extensively Drug Resistant TuberculosisabstractThe rise of multi-drug resistant (MDR) and extensively drug resistant (XDR) tuberculosis around the world, including in industrialized nations, poses a great threat to human health and defines a need to develop new, effective and inexpensive anti-tubercular agents. Previously we developed a chemical systems biology approach to identify off-targets of major pharmaceuticals on a proteome-wide scale. In this paper we further demonstrate the value of this approach through the discovery that existing commercially available drugs, prescribed for the treatment of Parkinson's disease, have the potential to treat MDR and XDR tuberculosis. These drugs, entacapone and tolcapone, are predicted to bind to the enzyme InhA and directly inhibit substrate binding. The prediction is validated by in vitro and InhA kinetic assays using tablets of Comtan, whose active component is entacapone. The minimal inhibition concentration (MIC(99)) of entacapone for Mycobacterium tuberculosis (M.tuberculosis) is approximately 260.0 microM, well below the toxicity concentration determined by an in vitro cytotoxicity model using a human neuroblastoma cell line. Moreover, kinetic assays indicate that Comtan inhibits InhA activity by 47.0% at an entacapone concentration of approximately 80 microM. Thus the active component in Comtan represents a promising lead compound for developing a new class of anti-tubercular therapeutics with excellent safety profiles. More generally, the protocol described in this paper can be included in a drug discovery pipeline in an effort to discover novel drug leads with desired safety profiles, and therefore accelerate the development of new drugs. Sarah L. Kinnings, Nina Liu, Nancy Buchmeier, Peter J. Tonge, Lei Xie 0006, Philip E. Bourne |
PLoS Comput. Biol. | 6 |
| 2009 | Sm/Lsm Genes Provide a Glimpse into the Early Evolution of the SpliceosomeabstractThe spliceosome, a sophisticated molecular machine involved in the removal of intervening sequences from the coding sections of eukaryotic genes, appeared and subsequently evolved rapidly during the early stages of eukaryotic evolution. The last eukaryotic common ancestor (LECA) had both complex spliceosomal machinery and some spliceosomal introns, yet little is known about the early stages of evolution of the spliceosomal apparatus. The Sm/Lsm family of proteins has been suggested as one of the earliest components of the emerging spliceosome and hence provides a first in-depth glimpse into the evolving spliceosomal apparatus. An analysis of 335 Sm and Sm-like genes from 80 species across all three kingdoms of life reveals two significant observations. First, the eukaryotic Sm/Lsm family underwent two rapid waves of duplication with subsequent divergence resulting in 14 distinct genes. Each wave resulted in a more sophisticated spliceosome, reflecting a possible jump in the complexity of the evolving eukaryotic cell. Second, an unusually high degree of conservation in intron positions is observed within individual orthologous Sm/Lsm genes and between some of the Sm/Lsm paralogs. This suggests that functional spliceosomal introns existed before the emergence of the complete Sm/Lsm family of proteins; hence, spliceosomal machinery with considerably fewer components than today's spliceosome was already functional. Stella Veretnik, Christopher Wills, Philippe Youkharibache, Ruben E. Valas, Philip E. Bourne |
PLoS Comput. Biol. | 5 |
| 2009 | Ten Simple Rules To Combine Teaching and ResearchabstractThe late Lindley J. Stiles famously made himself an advocate for teaching during his professorship at the University of Colorado: “If a better world is your aim, all must agree: The best should teach” (http://thebestshouldteach.org/). In fact, dispensing high-quality teaching and professional education is the primary goal of any university [1]. Thus, for most faculty positions in academia, teaching is a significant requirement of the job. Yet, the higher education programs offered to Ph.D. students do not necessarily incorporate any form of teaching exposure. We offer 10 simple rules that should help you to get prepared for the challenge of teaching while keeping some composure. Quentin Vicens, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2009 | Drug Discovery Using Chemical Systems Biology: Identification of the Protein-Ligand Binding Network To Explain the Side Effects of CETP InhibitorsabstractSystematic identification of protein-drug interaction networks is crucial to correlate complex modes of drug action to clinical indications. We introduce a novel computational strategy to identify protein-ligand binding profiles on a genome-wide scale and apply it to elucidating the molecular mechanisms associated with the adverse drug effects of Cholesteryl Ester Transfer Protein (CETP) inhibitors. CETP inhibitors are a new class of preventive therapies for the treatment of cardiovascular disease. However, clinical studies indicated that one CETP inhibitor, Torcetrapib, has deadly off-target effects as a result of hypertension, and hence it has been withdrawn from phase III clinical trials. We have identified a panel of off-targets for Torcetrapib and other CETP inhibitors from the human structural genome and map those targets to biological pathways via the literature. The predicted protein-ligand network is consistent with experimental results from multiple sources and reveals that the side-effect of CETP inhibitors is modulated through the combinatorial control of multiple interconnected pathways. Given that combinatorial control is a common phenomenon observed in many biological processes, our findings suggest that adverse drug effects might be minimized by fine-tuning multiple off-target interactions using single or multiple therapies. This work extends the scope of chemogenomics approaches and exemplifies the role that systems biology has in the future of drug discovery. Li Xie 0002, Jerry Li 0004, Lei Xie 0006, Philip E. Bourne |
PLoS Comput. Biol. | 4 |
| 2008 | ElliPro: a new structure-based tool for the prediction of antibody epitopesabstractBACKGROUND: Reliable prediction of antibody, or B-cell, epitopes remains challenging yet highly desirable for the design of vaccines and immunodiagnostics. A correlation between antigenicity, solvent accessibility, and flexibility in proteins was demonstrated. Subsequently, Thornton and colleagues proposed a method for identifying continuous epitopes in the protein regions protruding from the protein's globular surface. The aim of this work was to implement that method as a web-tool and evaluate its performance on discontinuous epitopes known from the structures of antibody-protein complexes. RESULTS: Here we present ElliPro, a web-tool that implements Thornton's method and, together with a residue clustering algorithm, the MODELLER program and the Jmol viewer, allows the prediction and visualization of antibody epitopes in a given protein sequence or structure. ElliPro has been tested on a benchmark dataset of discontinuous epitopes inferred from 3D structures of antibody-protein complexes. In comparison with six other structure-based methods that can be used for epitope prediction, ElliPro performed the best and gave an AUC value of 0.732, when the most significant prediction was considered for each protein. Since the rank of the best prediction was at most in the top three for more than 70% of proteins and never exceeded five, ElliPro is considered a useful research tool for identifying antibody epitopes in protein antigens. ElliPro is available at http://tools.immuneepitope.org/tools/ElliPro. CONCLUSION: The results from ElliPro suggest that further research on antibody epitopes considering more features that discriminate epitopes from non-epitopes may further improve predictions. As ElliPro is based on the geometrical properties of protein structure and does not require training, it might be more generally applied for predicting different types of protein-protein interactions. Julia V. Ponomarenko, Huynh-Hoa Bui, Nicholas Fusseder, Philip E. Bourne, Alessandro Sette, Björn Peters |
BMC Bioinform. | 5 |
| 2008 | I Am Not a Scientist, I Am a NumberabstractWe suspect many of our readers will be familiar with the cult TV show The Prisoner, in which actor Patrick McGoohan had his identity taken away by unknown assailants for unknown reasons, and his pleas of “I am not a number, I am a person” (http://www.youtube.com/watch?v=29JewlGsYxs&feature=related) were greeted with variants of “whatever you say, number six.” We would suggest that, as scientists, we are in a situation where the opposite will soon be true, at least for the purposes of scientific scholarship. Scientists will want to be assigned a number, or at least a unique identifier. Why?
Imagine a time when you and your complete scholarly output—papers, grant applications, blog posts, etc.—could be identified online and in perpetuity and returned in a variety of easy-to-digest ways. While ego comes into it as a driver to make this happen, measuring scientific career advancement is something that lacks good metrics in a digital world. Unless one has a truly unique name, applying such a metric is not possible now. Even with a unique name, what is the guarantee that all of our scholarly output will be captured by one source of that information? In the end, we as individuals are the only ones who reliably track our scholarly output. This situation is beginning to change, and, as we shall see, new metrics have the promise of much more than simply returning references to our collective life's work as currently described by research papers, research proceedings, books, and book chapters. Although even a complete and current resume generated on demand would be a big step, if it could be returned in a variety of formats for a variety of purposes. These complete resumes are something many of us spend endless hours generating.
The idea of having our scholarly output properly characterized is not out of reach, since the articles we write are already identified uniquely by a Digital Object Identifier (DOI; discussed further below). A book or journal is identified by an ISBN, and citations are identified by PubMed identifiers, and so on. The ideas discussed here simply take this identification process for individual publications and citations to the point of providing unique descriptors for each author and to uniquely identify all of each author's scholarly work. Philip E. Bourne, J. Lynn Fink |
PLoS Comput. Biol. | 1 |
| 2008 | Open Access: Taking Full Advantage of the ContentabstractThis Journal and the Public Library of Science (PLoS) at large are standard bearers of the full potential offered through open access publication, but what of you, the reader?For most of you, open access may imply free access to read the journals, but nothing more.There is a far greater potential, but, up to now, little to point to that highlights its tangible benefits.We would argue that, as yet, the full promise of open access has not been realized.There are few persistent applications that collectively use the full on-line corpus, which for the biosciences at least is maintained in PubMed Central (http://www.pubmedcen-tral.nih.gov/).In short, there are no ''killer apps.''Since this readership, beyond any other, would seem to have the ability to change this situation at least in the biosciences, we are issuing a call to action.While, first and foremost, open access implies downloading and reading full papers for free, additional possibilities exist depending on how the open access material is licensed. Philip E. Bourne, J. Lynn Fink, Mark Gerstein |
PLoS Comput. Biol. | 1 |
| 2008 | Ten Simple Rules for Organizing a Scientific MeetingabstractScientific meetings come in various flavors—from one-day focused workshops of 1–20 people to large-scale multiple-day meetings of 1,000 or more delegates, including keynotes, sessions, posters, social events, and so on. These ten rules are intended to provide insights into organizing meetings across the scale.
Scientific meetings are at the heart of a scientist's professional life since they provide an invaluable opportunity for learning, networking, and exploring new ideas. In addition, meetings should be enjoyable experiences that add exciting breaks to the usual routine in the laboratory. Being involved in organizing these meetings later in your career is a community responsibility. Being involved in the organization early in your career is a valuable learning experience [1]. First, it provides visibility and gets your name and face known in the community. Second, it is useful for developing essential skills in organization, management, team work, and financial responsibility, all of which are useful in your later career. Notwithstanding, it takes a lot of time, and agreeing to help organize a meeting should be considered in the context of your need to get your research done and so is also a lesson in time management. What follows are the experiences of graduate students in organizing scientific meetings with some editorial oversight from someone more senior (PEB) who has organized a number of major meetings over the years.
The International Society for Computational Biology (ISCB) Student Council [2] is an organization within the ISCB that caters to computational biologists early in their career. The ISCB Student Council provides activities and events to its members that facilitate their scientific development. From our experience in organizing the Student Council Symposium [3],[4], a meeting that so far has been held within the context of the ISMB [5],[6] and ECCB conferences, we have gained knowledge that is typically not part of an academic curriculum and which is embodied in the following ten rules. Manuel Corpas, Nils Gehlenborg, Sarath Chandra Janga, Philip E. Bourne |
PLoS Comput. Biol. | 4 |
| 2008 | Computational Biology Resources Lack Persistence and UsabilityabstractInnovation in computational biology research is predicated on the availability of published methods and computational resources. These resources facilitate the generation of new hypotheses and observations both on the part of the creators and the scientists who use them. These methods and resources include Web servers, databases, and software, both complex and simple, that implement a specific procedure or algorithm. Usually, a resource is maintained by the laboratory in which it was initially developed. We would assert that there is a growing level of frustration among scientists who attempt to use many of these resources and find that they no longer exist or are not properly maintained. Stella Veretnik, J. Lynn Fink, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2007 | Con-Struct Map: a comparative contact map analysis toolabstractUNLABELLED: Con-Struct Map is a graphical tool for the comparative study of protein structures. The tool detects potential conserved residue contacts shared by multiple protein structures by superimposing their contact maps according to a multiple structure alignment. In general, Con-Struct Map allows the study of structural changes resulting from, e.g. sequence substitutions, or alternatively, the study of conserved components of a structure framework across structurally aligned proteins. Specific applications include the study of sequence-structure relationship in distantly related proteins and the comparisons of wild type and mutant proteins. AVAILABILITY: http://pdbrs3.sdsc.edu/ConStructMap/viewer_argument_generator/singleArguments. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jo-Lan Chung, John E. Beaver, Eric D. Scheeff, Philip E. Bourne |
Bioinform. | 4 |
| 2007 | High-throughput identification of interacting protein-protein binding sitesabstractBACKGROUND: With the advent of increasing sequence and structural data, a number of methods have been proposed to locate putative protein binding sites from protein surfaces. Therefore, methods that are able to identify whether these binding sites interact are needed. RESULTS: We have developed a new method using a machine learning approach to detect if protein binding sites, once identified, interact with each other. The method exploits information relating to sequence and structural complementary across protein interfaces and has been tested on a non-redundant data set consisting of 584 homo-dimers and 198 hetero-dimers extracted from the PDB. Results indicate 87.4% of the interacting binding sites and 68.6% non-interacting binding sites were correctly identified. Furthermore, we built a pipeline that links this method to a modified version of our previously developed method that predicts the location of binding sites. CONCLUSION: We have demonstrated that this high-throughput pipeline is capable of identifying binding sites for proteins, their interacting binding sites and, ultimately, their binding partners on a large scale. Jo-Lan Chung, Wei Wang 0051, Philip E. Bourne |
BMC Bioinform. | 3 |
| 2007 | Identifying allosteric fluctuation transitions between different protein conformational states as applied to Cyclin Dependent Kinase 2abstractBACKGROUND: The mechanisms underlying protein function and associated conformational change are dominated by a series of local entropy fluctuations affecting the global structure yet are mediated by only a few key residues. Transitional Dynamic Analysis (TDA) is a new method to detect these changes in local protein flexibility between different conformations arising from, for example, ligand binding. Additionally, Positional Impact Vertex for Entropy Transfer (PIVET) uses TDA to identify important residue contact changes that have a large impact on global fluctuation. We demonstrate the utility of these methods for Cyclin-dependent kinase 2 (CDK2), a system with crystal structures of this protein in multiple functionally relevant conformations and experimental data revealing the importance of local fluctuation changes for protein function. RESULTS: TDA and PIVET successfully identified select residues that are responsible for conformation specific regional fluctuation in the activation cycle of Cyclin Dependent Kinase 2 (CDK2). The detected local changes in protein flexibility have been experimentally confirmed to be essential for the regulation and function of the kinase. The methodologies also highlighted possible errors in previous molecular dynamic simulations that need to be resolved in order to understand this key player in cell cycle regulation. Finally, the use of entropy compensation as a possible allosteric mechanism for protein function is reported for CDK2. CONCLUSION: The methodologies embodied in TDA and PIVET provide a quick approach to identify local fluctuation change important for protein function and residue contacts that contributes to these changes. Further, these approaches can be used to check for possible errors in protein dynamic simulations and have the potential to facilitate a better understanding of the contribution of entropy to protein allostery and function. Jenny Gu, Philip E. Bourne |
BMC Bioinform. | 2 |
| 2007 | A robust and efficient algorithm for the shape description of protein structures and its application in predicting ligand binding sitesabstractBACKGROUND: An accurate description of protein shape derived from protein structure is necessary to establish an understanding of protein-ligand interactions, which in turn will lead to improved methods for protein-ligand docking and binding site analysis. Most current shape descriptors characterize only the local properties of protein structure using an all-atom representation and are slow to compute. We need new shape descriptors that have the ability to capture both local and global structural information, are robust for application to models and low quality structures and are computationally efficient to permit high throughput analysis of protein structures. RESULTS: We introduce a new shape description that requires only the Calpha atoms to represent the protein structure, thus making it both fast and suitable for use on models and low quality structures. The notion of a geometric potential is introduced to quantitatively describe the shape of the structure. This geometric potential is dependent on both the global shape of the protein structure as well as the surrounding environment of each residue. When applying the geometric potential for binding site prediction, approximately 85% of known binding sites can be accurately identified with above 50% residue coverage and 80% specificity. Moreover, the algorithm is fast enough for proteome-scale applications. Proteins with fewer than 500 amino acids can be scanned in less than two seconds. CONCLUSION: The reduced representation of the protein structure combined with the geometric potential provides a fast, quantitative description of protein-ligand binding sites with potential for use in large-scale predictions, comparisons and analysis. Lei Xie 0006, Philip E. Bourne |
BMC Bioinform. | 2 |
| 2007 | Ten Simple Rules for Making Good Oral Presentationsabstractontinuing our ''Ten Simple Rules'' series [1][2][3][4][5], we consider here what it takes to make a good oral presentation.While the rules apply broadly across disciplines, they are certainly important from the perspective of this readership.Clear and logical delivery of your ideas and scientific results is an important component of a successful scientific career.Presentations encourage broader dissemination of your work and highlight work that may not receive attention in written form. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2007 | Developing Computational BiologyabstractScientific research is an international endeavor, and computational biology is no exception. Last year we were fortunate enough to attend conferences and visit laboratories in a number of different countries and developing economies. Some of those countries had well-established programs in computational biology, while others had fledgling efforts in place and big plans. Because of its cost structure, computational biology is a particularly promising field for emerging economies. Each country we visited had unique features in areas such as the education programs they offered, the types of research being undertaken, and the ways that research is funded. Each nation also has unique challenges, and many nations are developing so quickly that people who leave them for study abroad may find programs entirely different when they return.
Given these differences and rapid developments, we feel that there is much we can learn from each other. To this end we have begun a series of Perspective articles from computational biologists in a variety of countries, each of whom offers their personal perspectives on the history, current status, and future of computational biology in their region. We asked the authors to describe the specific challenges they have faced, their perceived strengths, as well as to discuss the institutions (government and private), opportunities, and difficulties of computational biology in their country as a whole.
This month we begin the series with a Perspective on the development of computational and genomic biology in Mexico from two of the leaders in the field. Dr. Rafael Palacios de la Lama is a foreign member of the US National Academy of Sciences and one of the world's experts on the genetics of Rhizobium and genome dynamics. Dr. Julio Collado-Vides is currently a professor and the Director of the Center for Genomic Sciences at the Universidad Nacional Autonoma de Mexico, a leader in understanding regulation of the complete Escherichia coli genome. In their article, we learn about the pioneering effort of a Mexican research group participating in the E. coli genome project. E. coli was the first among still few bacterial genomes with quality predictions of operons and upstream regulatory elements. Various other ambitious projects are under way in Mexico, such as determining the DNA sequence of the genome of Phaseolus vulgaris. To these ends, the authors educate an elite set of undergraduates in genomic sciences.
Beyond Mexico, in the coming months we will be touring a number of countries where computational biology is expanding.
The Intelligent Systems in Molecular Biology (ISMB) conference brought close to a thousand computational biologists to Fortaleza, Brazil, in the summer of 2006. Memories of that experience will revive as we read about the character and nature of research efforts that are being undertaken in Brazil.
We will learn also about the unique challenges of doing computational biology in Cuba, a country impacted by the US embargo, yet with a determined education system and a strong research emphasis. The passion of Cuban scientists is an inspiration, and their soup-to-nuts approach to research represents an interesting solution to their situation.
We can also expect perspectives from South Africa, Thailand, Argentina, and China, shedding light on current approaches in these countries to a young, rapidly evolving, and changing discipline.
We hope reading these accounts will inspire you to comment, or to write a Perspective on computational biology in your part of the world. If you are interested, please e-mail us at gro.solp@loibpmocsolp, and we will be happy to send you guidelines for a submission.
The pursuit of scientific endeavors around the world allows individuals and nations to capitalize on their potential. It furthers progress and mutual understanding where political and other means have reached their limits. We hope you will learn from the differing approaches and enjoy reading these viewpoints as much as we did. Philip E. Bourne, Steven E. Brenner |
PLoS Comput. Biol. | 1 |
| 2007 | Ten Simple Rules for a Good Poster PresentationabstractPosters are a key component of communicating your science and an important element in a successful scientific career. Posters, while delivering the same high-quality science, offer a different medium from either oral presentations [1] or published papers [2], and should be treated accordingly. Posters should be considered a snapshot of your work intended to engage colleagues in a dialog about the work, or, if you are not present, to be a summary that will encourage the reader to want to learn more. Many a lifelong collaboration [3] has begun in front of a poster board. Here are ten simple rules for maximizing the return on the time-consuming process of preparing and presenting an effective poster. Thomas C. Erren, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2007 | Ten Simple Rules for Doing Your Best Research, According to Hammingabstracthis editorial can be considered the preface to the ''Ten Simple Rules'' series [1][2][3][4][5][6][7].The rules presented here are somewhat philosophical and behavioural rather than concrete suggestions for how to tackle a particular scientific professional activity such as writing a paper or a grant.The thoughts presented are not our own; rather, we condense and annotate some excellent and timeless suggestions made by the mathematician Richard Hamming two decades ago on how to do ''first-class research'' [8].As far as we know, the transcript of the Bell Communications Research Colloquium Seminar provided by Dr. Kaiser [8] was never formally published, so that Dr. Hamming's thoughts are not as widely known as they deserve to be.By distilling these thoughts into something that can be thought of as ''Ten Simple Rules,'' we hope to bring these ideas to broader attention.Hamming's 1986 talk was remarkable.In ''You and Your Research,'' he addressed the question: How can scientists do great research, i.e., Nobel-Prize-type work?His insights were based on more than forty years of research as a pioneer of computer science and telecommunications who had the privilege of interacting with such luminaries as the physicists Richard Feynman, Enrico Fermi, Thomas C. Erren, Paul Cullen, Michael Erren, Philip E. Bourne |
PLoS Comput. Biol. | 4 |
| 2007 | Ten Simple Rules for Graduate StudentsabstractChoosing to go to graduate school is a major life decision. Whether you have already made that decision or are about to, now it is time to consider how best to be a successful graduate student. Here are some thoughts from someone who holds these memories fresh in her mind (JG) and from someone who has had a whole career to reflect back on the decisions made in graduate school, both good and bad (PEB). These thoughts taken together, from former student and mentor, represent experiences spanning some 25 or more years. For ease, these experiences are presented as ten simple rules, in approximate order of priority as defined by a number of graduate students we have consulted here in the US; but we hope the rules are more globally applicable, even though length, method of evaluation, and institutional structure of graduate education varies widely. These rules are intended as a companion to earlier editorials covering other areas of professional development [1–7]. Jenny Gu, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2007 | The International Society for Computational Biology 10th AnniversaryabstractPLoS Computational Biology is the official journal of the International Society for Computational Biology (ISCB), a partnership that was formed during the Journal's conception in 2005. With ISCB being the only international body representing computational biologists, it made perfect sense for PLoS Computational Biology to be closely affiliated. The Society had to take more of a chance than similar societies, choosing to step away from an existing financially beneficial subscription journal to align with an open access publication as a matter of principle. To our knowledge, ISCB was the first major international scientific society to do so.
Now, as PLoS Computational Biology reaches its two-year mark, ISCB simultaneously celebrates its tenth anniversary, having formed officially on June 18, 1997. We early presidents of ISCB reflect on the state of computational biology ten years ago, how far we have come since, and what thought-provoking future challenges might lie ahead with regard to innovations in publishing technologies. Lawrence Hunter, Russ B. Altman, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2007 | Ten Simple Rules for a Successful CollaborationabstractDOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone. Quentin Vicens, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2007 | In Silico Elucidation of the Molecular Mechanism Defining the Adverse Effect of Selective Estrogen Receptor ModulatorsabstractEarly identification of adverse effect of preclinical and commercial drugs is crucial in developing highly efficient therapeutics, since unexpected adverse drug effects account for one-third of all drug failures in drug development. To correlate protein-drug interactions at the molecule level with their clinical outcomes at the organism level, we have developed an integrated approach to studying protein-ligand interactions on a structural proteome-wide scale by combining protein functional site similarity search, small molecule screening, and protein-ligand binding affinity profile analysis. By applying this methodology, we have elucidated a possible molecular mechanism for the previously observed, but molecularly uncharacterized, side effect of selective estrogen receptor modulators (SERMs). The side effect involves the inhibition of the Sacroplasmic Reticulum Ca2+ ion channel ATPase protein (SERCA) transmembrane domain. The prediction provides molecular insight into reducing the adverse effect of SERMs and is supported by clinical and in vitro observations. The strategy used in this case study is being applied to discover off-targets for other commercially available pharmaceuticals. The process can be included in a drug discovery pipeline in an effort to optimize drug leads and reduce unwanted side effects. Lei Xie 0006, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2006 | Multipolar representation of protein structureabstractBACKGROUND: That the structure determines the function of proteins is a central paradigm in biology. However, protein functions are more directly related to cooperative effects at the residue and multi-residue scales. As such, current representations based on atomic coordinates can be considered inadequate. Bridging the gap between atomic-level structure and overall protein-level functionality requires parameterizations of the protein structure (and other physicochemical properties) in a quasi-continuous range, from a simple collection of unrelated amino acids coordinates to the highly synergistic organization of the whole protein entity, from a microscopic view in which each atom is completely resolved to a "macroscopic" description such as the one encoded in the three-dimensional protein shape. RESULTS: Here we propose such a parameterization and study its relationship to the standard Euclidian description based on amino acid representative coordinates. The representation uses multipoles associated with residue Calpha coordinates as shape descriptors. We demonstrate that the multipoles can be used for the quantitative description of the protein shape and for the comparison of protein structures at various levels of detail. Specifically, we construct a (dis)similarity measure in multipolar configuration space, and show how such a function can be used for the comparison of a pair of proteins. We then test the parameterization on a benchmark set of the protein kinase-like superfamily. We prove that, when the biologically relevant portions of the proteins are retained, it can robustly discriminate between the various families in the set in a way not possible through sequence or conventional structural representations alone. We then compare our representation with the Cartesian coordinate description and show that, as expected, the correlation with that representation increases as the level of detail, measured by the highest rank of multipoles used in the representation, approaches the dimensionality of the fold space. CONCLUSION: The results described here demonstrate how a granular description of the protein structure can be achieved using multipolar coefficients. The description has the additional advantage of being immediately generalizable for any residue-specific property therefore providing a unitary framework for the study and comparison of the spatial profile of various protein properties. Apostol Gramada, Philip E. Bourne |
BMC Bioinform. | 2 |
| 2006 | Application of protein structure alignments to iterated hidden Markov model protocols for structure predictionabstractBACKGROUND: One of the most powerful methods for the prediction of protein structure from sequence information alone is the iterative construction of profile-type models. Because profiles are built from sequence alignments, the sequences included in the alignment and the method used to align them will be important to the sensitivity of the resulting profile. The inclusion of highly diverse sequences will presumably produce a more powerful profile, but distantly related sequences can be difficult to align accurately using only sequence information. Therefore, it would be expected that the use of protein structure alignments to improve the selection and alignment of diverse sequence homologs might yield improved profiles. However, the actual utility of such an approach has remained unclear. RESULTS: We explored several iterative protocols for the generation of profile hidden Markov models. These protocols were tailored to allow the inclusion of protein structure alignments in the process, and were used for large-scale creation and benchmarking of structure alignment-enhanced models. We found that models using structure alignments did not provide an overall improvement over sequence-only models for superfamily-level structure predictions. However, the results also revealed that the structure alignment-enhanced models were complimentary to the sequence-only models, particularly at the edge of the "twilight zone". When the two sets of models were combined, they provided improved results over sequence-only models alone. In addition, we found that the beneficial effects of the structure alignment-enhanced models could not be realized if the structure-based alignments were replaced with sequence-based alignments. Our experiments with different iterative protocols for sequence-only models also suggested that simple protocol modifications were unable to yield equivalent improvements to those provided by the structure alignment-enhanced models. Finally, we found that models using structure alignments provided fold-level structure assignments that were superior to those produced by sequence-only models. CONCLUSION: When attempting to predict the structure of remote homologs, we advocate a combined approach in which both traditional models and models incorporating structure alignments are used. Eric D. Scheeff, Philip E. Bourne |
BMC Bioinform. | 2 |
| 2006 | Curation of complex, context-dependent immunological dataabstractBACKGROUND: The Immune Epitope Database and Analysis Resource (IEDB) is dedicated to capturing, housing and analyzing complex immune epitope related data http://www.immuneepitope.org. DESCRIPTION: To identify and extract relevant data from the scientific literature in an efficient and accurate manner, novel processes were developed for manual and semi-automated annotation. CONCLUSION: Formalized curation strategies enable the processing of a large volume of context-dependent data, which are now available to the scientific community in an accessible and transparent format. The experiences described herein are applicable to other databases housing complex biological data and requiring a high level of curation expertise. Randi Vita, Kerrie Vaughan, Laura Zarebski, Nima Salimi, Ward Fleri, Howard Grey, Muthu Sathiamurthy, John Mokili, Huynh-Hoa Bui, Philip E. Bourne, Julia V. Ponomarenko, Romulo de Castro Jr., Russell K. Chan, John Sidney, Stephen S. Wilson, Scott Stewart, Scott Way, Björn Peters, Alessandro Sette |
BMC Bioinform. | 10 |
| 2006 | One Year of PLoS Computational BiologyabstractThe June 2006 issue of PLoS Computational Biology marked one year of publication of the journal. While it is too early to formally assess the impact of the journal, it is worth reflecting on what has been achieved in the first year of publication. Twelve monthly issues actually reflect eighteen months of submissions, given that we have been considering papers since January of 2005. At the end of June 2006, 631 research articles have been submitted to the journal, 110 have been published, and 64 (48 new submissions and 16 revisions) are currently under review. We estimate that our acceptance rate is currently about 30%, increasing in recent months from less than 20% as authors become more familiar with the expectations of the journal and do not submit papers that have little chance of being published. We are now publishing about 15 research articles per month. Accompanying these research articles over the year have been three Editorials, six Reviews, and five Perspectives, and the Education section just published its first tutorial. Time for review averages 13.2 days and for acceptance to publication is five to six weeks, although authors have the option of having their manuscripts posted as soon as they are accepted. The journal is published in association with the International Society for Computational Biology (ISCB), an important relationship whereby one author of each accepted paper gets a free one-year membership to the Society. Articles about ISCB activities are also a regular feature in the journal. As the relationship between PLoS and ISCB matures in the coming years, we hope that this will provide a stimulus for other scholarly societies to explore and adopt open access publishing.
Submissions have been received from 41 countries (based on the location of the corresponding author). The top six countries submitting articles are the US (49.1%), the UK (5.1%), Japan and India (3.3% each), Netherlands (3.0%), and Israel (2.5%).
Interest in the journal can be gauged by the number of people who have signed up to receive an electronic alert of journal contents (eTOC) and by the number of downloads of journal articles. Currently 7,125 people are signed up to receive eTOCs, a number that grows by several hundred each month. Since the launch, there have been more than 250,000 article downloads, comprising 200,019 research articles, 23,408 Editorials, 14,331 Perspectives, 13,869 Reviews, and a number of hits for the Education Column and for Message from the ISCB. The top ten papers downloaded thus far are shown in Table 1.
Table 1
Top Ten Papers Downloaded from PLoS Computational Biology in Its First Year
Two of the top ten papers are in the area of neuroscience, and two others form part of the ongoing “Ten Rules” series of Editorials to aid our less experienced readers (see Table 1). Of the remainder, one is a Perspective discussing team versus individual science and the rest are in the realm of computational molecular biology. Demographics of materials downloaded show North America and Europe account for 60%–70% of the usage of the site, but there is significant usage from the developing world, a testament to open access.
Overall, we are well on the way to meeting the original editorial goal of the journal—to establish a high-quality knowledge resource serving a community interested in advancing our understanding of living systems through the use of computational techniques. While reported advances have been predominantly at the molecular level, there is a growing body of work being submitted that covers different levels of biological organization. This reflects our goal to publish great work involving computational analyses on all biological scales. We want to make connections between researchers who are using conceptually related approaches to tackle diverse issues in biology.
To put this goal in perspective, consider that in one year we have explored the RNA silencing pathway in issue 1(2), designed a nanotube using naturally occurring protein building blocks in issue 2(4), modeled the transition to quorum sensing in a population of Agrobacterium in issue 1(4), and understood more about the role of mechanical factors in the morphology of the primate cerebral cortex in issue 2(3), to name but a few articles. Not bad for the first year.
We aspire to have PLoS Computational Biology develop into an exemplar open access community journal which will provide a model for scholarly publishing of the future. The publication fee for the journal has recently increased from US$1,500 to US$2,000, and, along with the growth of the journal (in terms of submissions and published articles), the journal is moving steadily along a path towards financial sustainability. This is good news, and at the same time PLoS retains a fee waiver policy for those authors with insufficient funds, so money is never a factor in disseminating good science.
Over the course of the past year, our monthly submissions have increased to a record 55 in June. We are all tremendously gratified to see this strong community response, and it should be remembered that this support has been offered in the absence of an impact factor (perhaps not a good indicator for an open access journal, but that is another Editorial). That we have been able to maintain an efficient editorial and publishing service in the face of this growth is testament to the tremendous efforts of our Editorial Board. My thanks go to them, and also to the PLoS staff, notably Catherine Nancarrow, Emily Stevenson, and Mark Patterson who really are the ones who keep it all on track.
Join us in our second year by getting your work published in this fast-growing and diverse journal, which we hope will make those of us involved in Computational Biology proud of our collective efforts in this rapidly evolving field. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2006 | Ten Simple Rules for Getting GrantsabstractThis piece follows an earlier Editorial, ‘‘Ten Simple Rules for Getting Published’’ [1], which has generated significant interest, is well read, and continues to generate a variety of positive comments. That Editorial was aimed at students in the early stages of a life of scientific paper writing. This interest has prompted us to try to help scientists in making the next academic career step—becoming a young principal investigator. Leo Chalupa has joined us in putting together ten simple rules for getting grants, based on our many collective years of writing both successful and unsuccessful grants. While our grant writing efforts have been aimed mainly at United States government funding agencies, we believe the rules presented here are generic, transcending funding institutions and national boundaries. At the present time, US funding is frequently below 10% for a given grant program. Today, more than ever, we need all the help we can get in writing successful grant proposals. We hope you find these rules useful in reaching your research career goals. Philip E. Bourne, Leo M. Chalupa |
PLoS Comput. Biol. | 1 |
| 2006 | Ten Simple Rules for Selecting a Postdoctoral PositionabstractYou are a PhD candidate and your thesis defense is already in sight. You have decided you would like to continue with a postdoctoral position rather than moving into industry as the next step in your career (that decision should be the subject of another “Ten Simple Rules”). Further, you already have ideas for the type of research you wish to pursue and perhaps some ideas for specific projects. Here are ten simple rules to help you make the best decisions on a research project and the laboratory in which to carry it out. Philip E. Bourne, Iddo Friedberg |
PLoS Comput. Biol. | 1 |
| 2006 | Correction: Ten Simple Rules for Selecting a Postdoctoral PositionabstractIn the section called Rule 1: Select a Position That Excites You, the sentence, ''Just because the mentor is excited about the project does not mean that you will be six months into it.'', had the two words ''that' ' and ''you'' switched. Philip E. Bourne, Iddo Friedberg |
PLoS Comput. Biol. | 1 |
| 2006 | Ten Simple Rules for ReviewersabstractLast summer, the Student Council of the International Society for Computational Biology prompted an Editorial, “Ten Simple Rules for Getting Published” [1]. The interest in that piece (it has been downloaded 14,880 times thus far) prompted “Ten Simple Rules for Writing a Grant” [2]. With this third contribution, the “Ten Rules” series would seem to be established, and more rules for different audiences are in the making. Ten Simple Rules for Reviewers is based upon our years of experience as reviewers and as managers of the review process. Suggestions also came from PLoS staff and Editors and our research groups, the latter being new and fresh to the process of reviewing.
The rules for getting articles published included advice on becoming a reviewer early in your career. If you followed that advice, by working through your mentors who will ask you to review, you will then hopefully find these Ten Rules for Reviewers helpful. There is no magic formula for what constitutes a good or a bad paper—the majority of papers fall in between—so what do you look for as a reviewer? We would suggest, above all else, you are looking for what the journal you are reviewing for prides itself on. Scientific novelty—there is just too much “me-too” in scientific papers—is often the prerequisite, but not always. There is certainly a place for papers that, for example, support existing hypotheses, or provide a new or modified interpretation of an existing finding. After journal scope, it comes down to a well-presented argument and everything else described in “Ten Simple Rules for Getting Published” [1]. Once you know what to look for in a paper, the following simple reviewer guidelines we hope will be useful. Certainly (as with all PLoS Computational Biology material) we invite readers to use the PLoS eLetters feature to suggest their own rules and comments on this important subject. Philip E. Bourne, Alon Korngreen |
PLoS Comput. Biol. | 1 |
| 2006 | Biocurators: Contributors to the World of ScienceabstractComputational biology is a discipline built upon data (mostly free access), found in biological databases, and knowledge (mostly not free access), found in the literature. So important are these online sources of data that the discipline, and indeed this Journal, simply would not exist without them. Whether we are using the data in “browse mode”—doing a PubMed search, looking up a reaction in an enzymatic pathway, or in “compute mode”—analysis of a large dataset, we usually visit Web sites and download information without a second thought. Since our discipline is so dependent on the availability, extent, and quality of biological data, it is worth taking some time to think about the processes of data accessibility, annotation, and validation. These processes depend very much on biocurators—trained staff who ensure the information you are receiving is as complete and accurate as possible.
Biocurators can be considered the museum catalogers of the Internet age: they turn inert and unidentifiable objects (now virtual) into a powerful exhibit from which we can all marvel and learn. That would be a decent enough contribution to the world of science, but the task of the biocurator is even more extensive. Computational biologists do not expect to merely walk through the door, cast a casual eye over the exhibit, and exit wiser (although we frequently do); we also want to add our own data to the exhibit, plus pick and choose pieces of it to take home and create new exhibits of our own. Oh, and we would like to do all these things with minimal effort, please. We can be a pretty exacting bunch of customers, and it takes skills over and above a knowledge of biology to juggle the different needs of data submitters, information seekers, and power players.
“We pay homage to these special individuals who are dedicated to making our research endeavors a success.”
In this October issue, we pay homage to these special individuals who are dedicated to making our research endeavors a success. We do so through two Perspectives written by biocurators working with different types of biological data. The first is by biocurators from the Research Collaboratory for Structural Bioinformatics Protein Data Bank (PDB), a well-established biological resource of macromolecular structure data used by more than 10,000 individual scientists per day, and the second by biocurators of the Immune Epitope Database and Analysis Resource (IEDB), a new resource detailing known epitopes and their immunological outcomes. The PDB validates the quality and consistency of primary data submitted by structural biologists as a prerequisite to publication. The IEDB curates the published literature, extracting relevant facts about the epitopes discussed therein. As you read these two Perspectives, similarities and differences concerning the approaches will emerge. But more than anything, we hope you are struck by the level of professionalism and dedication that goes into helping to make the quality research articles that you read in this Journal and elsewhere.
These two articles are told from the perspective of the biocurators themselves. It is only two perspectives; we certainly encourage you to send eLetters with your own perspective on biocuration, either as a curator of a different type of information, or as a person whose information has been curated, or as a consumer of information that has been curated. If you are not moved to comment, at least give a thought to the person upon whose efforts your research may well depend. Philip E. Bourne, Johanna R. McEntyre |
PLoS Comput. Biol. | 1 |
| 2006 | Wiggle - Predicting Functionally Flexible Regions from Primary SequenceabstractThe Wiggle series are support vector machine-based predictors that identify regions of functional flexibility using only protein sequence information. Functionally flexible regions are defined as regions that can adopt different conformational states and are assumed to be necessary for bioactivity. Many advances have been made in understanding the relationship between protein sequence and structure. This work contributes to those efforts by making strides to understand the relationship between protein sequence and flexibility. A coarse-grained protein dynamic modeling approach was used to generate the dataset required for support vector machine training. We define our regions of interest based on the participation of residues in correlated large-scale fluctuations. Even with this structure-based approach to computationally define regions of functional flexibility, predictors successfully extract sequence-flexibility relationships that have been experimentally confirmed to be functionally important. Thus, a sequence-based tool to identify flexible regions important for protein function has been created. The ability to identify functional flexibility using a sequence based approach complements structure-based definitions and will be especially useful for the large majority of proteins with unknown structures. The methodology offers promise to identify structural genomics targets amenable to crystallization and the possibility to engineer more flexible or rigid regions within proteins to modify their bioactivity. Jenny Gu, Michael Gribskov, Philip E. Bourne |
PLoS Comput. Biol. | 3 |
| 2006 | ISMB 2006
Goran Neshich, Philip E. Bourne, Søren Brunak |
PLoS Comput. Biol. | 2 |
| 2005 | The Molecular Biology Toolkit (MBT): a modular platform for developing molecular visualization applicationsabstractBACKGROUND: The large amount of data that are currently produced in the biological sciences can no longer be explored and visualized efficiently with traditional, specialized software. Instead, new capabilities are needed that offer flexibility, rapid application development and deployment as standalone applications or available through the Web. RESULTS: We describe a new software toolkit--the Molecular Biology Toolkit (MBT; http://mbt.sdsc.edu)--that enables fast development of applications for protein analysis and visualization. The toolkit is written in Java, thus offering platform-independence and Internet delivery capabilities. Several applications of the toolkit are introduced to illustrate the functionality that can be achieved. CONCLUSIONS: The MBT provides a well-organized assortment of core classes that provide a uniform data model for the description of biological structures and automate most common tasks associated with the development of applications in the molecular sciences (data loading, derivation of typical structural information, visualization of sequence and standard structural entities). John L. Moreland, Apostol Gramada, Oleksandr V. Buzko, Philip E. Bourne |
BMC Bioinform. | 5 |
| 2005 | Will a Biological Database Be Different from a Biological Journal?abstractCitation: Bourne P (2005) Will a biological database be different from a biological journal? PLoS Comp Biol 1(3): e34. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2005 | Ten Simple Rules for Getting PublishedabstractBiology asked me to present my thoughts on getting published in the field of computational biology at the Intelligent Systems in Molecular Biology conference held in Detroit in late June of 2005.Close to 200 bright young souls (and a few not so young) crammed into a small room for what proved to be a wonderful interchange among a group of whom approximately one-half had yet to publish their first paper.The advice I gave that day I have modified and present as ten rules for getting published. Philip E. Bourne |
PLoS Comput. Biol. | 1 |
| 2005 | PLoS Computational Biology: A New Community JournalabstractDOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone. Philip E. Bourne, Steven E. Brenner, Michael B. Eisen |
PLoS Comput. Biol. | 1 |
| 2005 | Structural Evolution of the Protein Kinase-Like SuperfamilyabstractThe protein kinase family is large and important, but it is only one family in a larger superfamily of homologous kinases that phosphorylate a variety of substrates and play important roles in all three superkingdoms of life. We used a carefully constructed structural alignment of selected kinases as the basis for a study of the structural evolution of the protein kinase-like superfamily. The comparison of structures revealed a "universal core" domain consisting only of regions required for ATP binding and the phosphotransfer reaction. Remarkably, even within the universal core some kinase structures display notable changes, while still retaining essential activity. Hence, the protein kinase-like superfamily has undergone substantial structural and sequence revision over long evolutionary timescales. We constructed a phylogenetic tree for the superfamily using a novel approach that allowed for the combination of sequence and structure information into a unified quantitative analysis. When considered against the backdrop of species distribution and other metrics, our tree provides a compelling scenario for the development of the various kinase families from a shared common ancestor. We propose that most of the so-called "atypical kinases" are not intermittently derived from protein kinases, but rather diverged early in evolution to form a distinct phyletic group. Within the atypical kinases, the aminoglycoside and choline kinase families appear to share the closest relationship. These two families in turn appear to be the most closely related to the protein kinase family. In addition, our analysis suggests that the actin-fragmin kinase, an atypical protein kinase, is more closely related to the phosphoinositide-3 kinase family than to the protein kinase family. The two most divergent families, alpha-kinases and phosphatidylinositol phosphate kinases (PIPKs), appear to have distinct evolutionary histories. While the PIPKs probably have an evolutionary relationship with the rest of the kinase superfamily, the relationship appears to be very distant (and perhaps indirect). Conversely, the alpha-kinases appear to be an exception to the scenario of early divergence for the atypical kinases: they apparently arose relatively recently in eukaryotes. We present possible scenarios for the derivation of the alpha-kinases from an extant kinase fold. Eric D. Scheeff, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2005 | Functional Coverage of the Human Genome by Existing Structures, Structural Genomics Targets, and Homology ModelsabstractThe bias in protein structure and function space resulting from experimental limitations and targeting of particular functional classes of proteins by structural biologists has long been recognized, but never continuously quantified. Using the Enzyme Commission and the Gene Ontology classifications as a reference frame, and integrating structure data from the Protein Data Bank (PDB), target sequences from the structural genomics projects, structure homology derived from the SUPERFAMILY database, and genome annotations from Ensembl and NCBI, we provide a quantified view, both at the domain and whole-protein levels, of the current and projected coverage of protein structure and function space relative to the human genome. Protein structures currently provide at least one domain that covers 37% of the functional classes identified in the genome; whole structure coverage exists for 25% of the genome. If all the structural genomics targets were solved (twice the current number of structures in the PDB), it is estimated that structures of one domain would cover 69% of the functional classes identified and complete structure coverage would be 44%. Homology models from existing experimental structures extend the 37% coverage to 56% of the genome as single domains and 25% to 31% for complete structures. Coverage from homology models is not evenly distributed by protein family, reflecting differing degrees of sequence and structure divergence within families. While these data provide coverage, conversely, they also systematically highlight functional classes of proteins for which structures should be determined. Current key functional families without structure representation are highlighted here; updated information on the "most wanted list" that should be solved is available on a weekly basis from http://function.rcsb.org:8080/pdb/function_distribution/index.html. Lei Xie 0006, Philip E. Bourne |
PLoS Comput. Biol. | 2 |
| 2004 | The Future of Bioinformatics
Philip E. Bourne |
APBC | 1 |
| 2004 | The Protein Data Bank and lessons in data managementabstractThe Protein Data Bank (PDB) is a widely used biological database of macromolecular structures with a long history. This history is treated as lessons learned and is used to highlight what are believed to be the best practices important to developers of biological databases today. While the focus is on data quality, data representation and the information technology to support these data, the non-data and technology issues cannot be ignored. The role of the human factor in the form of users, collaborators, scientific society and ad hoc committees is also included. Philip E. Bourne, John D. Westbrook, Helen M. Berman |
Briefings Bioinform. | 1 |
| 2004 | Statistically rigorous automated protein annotationabstractMOTIVATION: Assignment of putative protein functional annotation by comparative analysis using pre-defined experimental annotations is performed routinely by molecular biologists. The number and statistical significance of these assignments remains a challenge in this era of high-throughput proteomics. A combined statistical method that enables robust, automated protein annotation by reliably expanding existing annotation sets is described. An existing clustering scheme, based on relevant experimental information (e.g. sequence identity, keywords or gene expression data) is required. The method assigns new proteins to these clusters with a measure of reliability. It can also provide human reviewers with a reliability score for both new and previously classified proteins. RESULTS: A dataset of 27 000 annotated Protein Data Bank (PDB) polypeptide chains (of 36 000 chains currently in the PDB) was generated from 23 000 chains classified a priori. AVAILABILITY: PDB annotations and sample software implementation are freely accessible on the Web at http://pmr.sdsc.edu/go Werner G. Krebs, Philip E. Bourne |
Bioinform. | 2 |
| 2004 | A case study of high-throughput biological data processing on parallel platformsabstractMOTIVATION: Analysis of large biological data sets using a variety of parallel processor computer architectures is a common task in bioinformatics. The efficiency of the analysis can be significantly improved by properly handling redundancy present in these data combined with taking advantage of the unique features of these compute architectures. RESULTS: We describe a generalized approach to this analysis, but present specific results using the program CEPAR, an efficient implementation of the Combinatorial Extension algorithm in a massively parallel (PAR) mode for finding pairwise protein structure similarities and aligning protein structures from the Protein Data Bank. CEPAR design and implementation are described and results provided for the efficiency of the algorithm when run on a large number of processors. AVAILABILITY: Source code is available by contacting one of the authors. Dmitry Pekurovsky, Ilya N. Shindyalov, Philip E. Bourne |
Bioinform. | 3 |
| 2003 | BioEditor - Simplifying Macromolecular Structure AnnotationabstractSUMMARY: BioEditor is an application to enable scientists and educators to prepare and present structure annotations containing formatted text, graphics, sequence data, and interactive molecular views. It is intended to bridge the gap between printed journal articles and Internet presentation formats. BioEditor is relevant in the era of structural genomics, where annotation and publication could become the rate determining step in structure determination. AVAILABILITY: BioEditor is available at http://bioeditor.sdsc.edu. The Web site includes the latest version of the software for Microsoft Windows, including documentation, the opportunity to submit bug reports and suggestions, example documentaries prepared with BioEditor and a repository where users can submit documentaries for posting to the site. Paul A. Craig, David S. Goodsell, Philip E. Bourne |
Bioinform. | 4 |
| 2002 | An ontology driven architecture for derived representations of macromolecular structureabstractUNLABELLED: An object metamodel based on a standard scientific ontology has been developed and used to generate a CORBA interface, an SQL schema and an XML representation for macromolecular structure (MMS) data. In addition to the interface and schema definitions, the metamodel was also used to generate the core elements of a CORBA reference server and a JDBC database loader. The Java source code which implements this metamodel, the CORBA server, database loader and XML converter along with detailed documentation and code examples are available as part of the OpenMMS toolkit. AVAILABILITY: http://openmms.sdsc.edu CONTACT: [email protected] Douglas S. Greer, John D. Westbrook, Philip E. Bourne |
Bioinform. | 3 |
| 2000 | ISMB-2000: Bioinformatics enters a new millenniumabstractThe huge success of the 8th International Conference on Intelligent Systems for Molecular Biology, ISMB-2000 (http://ismb2000.sdsc.edu),indicates that bioinformatics and computational biology have truly come of age.The ISMB-2000 meeting attracted over 1200 attendees from 44 countries and registration had to be closed due to lack of space.We believe the incredible demand for this conference reflects the growing acceptance of bioinformatics as a scientific discipline, as well as the still growing need for trained professionals to deal with the growth in biological data.The completion of the human genome marks a unique point in time for biology.The reductionist paradigm will soon have been pushed to its ultimate limit, in which we will have identified every gene, every message, and every protein in the cell.The challenge now is to build synthetic and predictive models that will enable us to understand the complexity of living organisms as holistic systems.ISMB-2000 placed special emphasis on knowledge discovery from the modeling and simulation of complex biological systems.This included, but was not limited to, interpretation of large-scale gene expression data, whole genome comparative analysis, and mathematical modeling of biochemical pathways.This seemed appropriate given the completion of the human genome and thus the continual reference to the post-genomic era.Bioinformatics remains a notably international discipline.Over 140 papers were received, with approximately 60% coming from the US, 30% from Europe and 10% from other parts of the world.Forty-two papers were selected for oral presentation.In addition keynote addresses were given by Philip E. Bourne, Michael Gribskov |
Bioinform. | 1 |
| 2000 | STAR/mmCIF: An ontology for macromolecular structureabstractMOTIVATION: Crystallographers were motivated 10 years ago to develop a simple and consistent data representation for the exchange and archiving of data associated with the crystallographic experiment and the final structure. As this process evolved (and the data grew at near exponential rates) came the recognition that this representation should also facilitate the automated management of the data and, with the aid of additional software for verification and validation, provide improved consistency and accuracy and hence improved scientific inquiry. This realization led to a new Dictionary Definition Language (DDL) and an extensive dictionary based on this DDL for describing macromolecular structure. In broad terms this could be considered an ontology. An important feature in the development of the ontology was the endorsement and ongoing maintenance and support of the International Union of Crystallography (IUCr). While the description of macromolecular structure and the x-ray crystallographic experiment used to derive it represent explicit data, the ontology is extensible and applicable to other less well-characterized data domains. RESULTS: Details of the DDL, the dictionaries that have been developed, and software for reading and using this ontology are presented. AVAILABILITY: Extensive documentation, software tools and the DDL and dictionaries are available from http://ndbserver.rutgers.edu/mmcif and associated mirror sites. CONTACT: Bourne: [email protected] and Westbrook:[email protected] John D. Westbrook, Philip E. Bourne |
Bioinform. | 2 |
| 1999 | An analysis of the Protein Data Bank in search of temporal and global trendsabstractMOTIVATION: Biological databases, with their rapidly expanding contents, are indispensable tools in the quest to understand more about biological function. However, a serious user of a database that comprises a large collection of data, collected over a long period, will likely be struck by the inconsistency in reporting individual items of data. This paper takes a critical look at the Protein Data Bank (PDB) to explore the seriousness of the problem in one particular data set and to explore the implications to those actively engaged in comparative analysis of these data. RESULTS: Averaged over the complete corpus, the stereochemical quality of atomic models has, in the past few years, moved towards ideal values. At the same time, there are inconsistencies in how data are reported. Water content is not reported consistently and the percent of data collected when reporting the high-resolution shell varies, detracting from the value of resolution as a yardstick for assessing the quality of a structure. A more detailed analysis of these inconsistencies is hampered by the lack of machine-readable experimental data. To the user of macromolecular structure data, this suggests that structural details beyond the standard quality measures of resolution and R value should be considered when using coordinate sets for further derivation or in inferring biological function. To the curators of the PDB, this suggests the need to capture more of the experimental data associated with the experiment in a way that permits straightforward parsing. Helge Weissig, Philip E. Bourne |
Bioinform. | 2 |
| 1997 | Code Generation through Annotation of Macromolecular Structure Data
John Biggs, Calton Pu, Philip E. Bourne |
ISMB | 3 |
| 1997 | Protein data representation and query using optimized data decompositionabstractMotivation: To provide data management tools to maintain and query efficiently experimental and derived protein data with the goal of providing new insights into structure-function relationships. The tools should be portable, extensible, and accessible locally, or via the World Wide Web, providing data that would not otherwise be available. Results: The initial phase of the work, the data representation and query of all available macromolecular structure data, including real-time access to complex property patterns based on the amino acid sequence, is reported. Protein structure data taken from the Protein Data Bank (PDB) are decomposed into native and derived elementary properties, and represented as compact indexed objects minimizing storage requirements and query time for select types of query. In addition, collections of indices representing a particular property are maintained and can be queried for specific property patterns found across the whole database. The approach is proving applicable to a wide variety of data available on specific protein families. Availability: Three resources an available using this approach, (i) The query of basic structural components and property patterns of the complete PDB is available via the World Wide Web at the URL http://www.sdsc.edu/moose. (ii) WPDB, a PC-based compressed macromolecular structure database and loader with a Microsoft Windows inteiface, is available fivm ftp:llftpsdsc.edu/publsdsdbiologyl WPDB/. (iii) A data base supporting real-time three-dimensional substructure searching will be reported elsewhere. Source code is available by contacting the authors. Contact: E-mail:{shindyal.bourne}@sdsc.edu Ilya N. Shindyalov, Philip E. Bourne |
Comput. Appl. Biosci. | 2 |
| 1994 | Design and Application of a C++ Macromolecular Class Library
Weider Chang, Ilya N. Shindyalov, Calton Pu, Philip E. Bourne |
ISMB | 4 |
| 1994 | Design and application of PDBlib, a C++ macromolecular class libraryabstractPDBlib is an extensible object-oriented class library written in C++ for representing the three-dimensional structure of biological macromolecules. The software design strategy, features of many of the 129 classes currently distributed with the library, and two sample applications which use the library are described. Version 1.0 of the library represents the structural features of proteins, DNA, RNA and complexes thereof, at a level of detail on a par with that which can be parsed from a Protein Data Bank (PDB) entry. However, the memory-resident representation of the macromolecule is independent of the PDB entry and can be obtained from other sources, e.g. relational and object-oriented databases. PDBlib classes are organized into four categories: (i) classes that model the macromolecule; (ii) classes that enhance the extensibility of the library; (iii) classes that provide navigation facilities of the object-oriented macromolecular structure representation; and (iv) a class that loads a PDB file into the memory-resident object-oriented representation. A number of general-purpose procedures that return features of this representation and that are relevant to all biological disciplines are included in (i). The library has been used to develop PDBtool, a prototype structure verification tool, and PDBview, a structure rendering tool that requires no specialized graphics hardware and software. Current work centers on making the macromolecular structures represented by PDBlib persistent using a commercial object-oriented database and providing an additional class library, MMQLlib, to query those structures. Weider Chang, Ilya N. Shindyalov, Calton Pu, Philip E. Bourne |
Comput. Appl. Biosci. | 4 |
| 1994 | An algorithm based on graph theory for the assembly of contigs in physical mapping of DNAabstractAn algorithm is described for mapping DNA contigs based on an interval graph (IG) representation. In general terms, the input to the algorithm is a set of binary overlapping relations among finite intervals spread along a real line, from which the algorithm generates sets of ordered overlapping fragments spanning that line. The implications of a more general case of the IG, called a probe interval graph (PIG), in which only a subset of cosmids are used as probes, are also discussed. In the specific case of cosmids hybridizing to regions of a YAC, the algorithm takes cross-hybridization information using the cosmids as probes, and orders them along the YAC; if gaps exist due to insufficient coverage of cosmid contigs along the length of the YAC, repetitive use of the algorithm generates sets of ordered overlapping fragments. Both the IG and the PIG can expose problems caused by false overlaps, such as hybridizations due to repetitive elements. The algorithm, has been coded in C; CPU time is essentially linear with respect to the number of cosmids analyzed. Results are presented for the application of a PIG to cosmid contig assembly along a human chromosome 13-specific YAC. An alignment of 67 cosmids spanning a YAC took 0.28 seconds of CPU time on a Convex 220 computer. Peisen Zhang, Eric A. Schon, Stuart G. Fischer, Eftihia Cayanis, Janie Weiss, Susan Kistler, Philip E. Bourne |
Comput. Appl. Biosci. | 7 |