Allegra Via

dblp:09/4212 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
0since 2021 · last 2020
0000-0002-3398-5462ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 7 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 77% Computing education · 23%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computing education
bioinformatics training
0.212015
The GOBLET training portal: a global repository of bioinformatics training materials, courses and trainers · Bioinform. 2015
Bioinformatics and computational biology
drug discovery
0.212013
TiPs: a database of therapeutic targets in pathogens and associated tools · Bioinform. 2013
Bioinformatics and computational biology › drug discovery › target identification
therapeutic target identification
0.212013
TiPs: a database of therapeutic targets in pathogens and associated tools · Bioinform. 2013
Bioinformatics and computational biology › protein analysis
protein-protein interaction
0.112006
A novel structure-based encoding for machine-learning applied to the inference of SH3 domain specificity · Bioinform. 2006
Bioinformatics and computational biology › protein analysis › protein bioinformatics
protein annotation
0.112005
Seq2Struct: a resource for establishing sequence-structure links · Bioinform. 2005
Bioinformatics and computational biology › structural bioinformatics
protein structure database
0.112005
Seq2Struct: a resource for establishing sequence-structure links · Bioinform. 2005
Bioinformatics and computational biology
database and knowledge base
0.012013
TiPs: a database of therapeutic targets in pathogens and associated tools · Bioinform. 2013

Methods — techniques the papers use, named apart from their topics

portal development · 0.2widget integration · 0.2web services · 0.2sequence-structure integration · 0.1machine learning · 0.1sequence alignment · 0.1
YearPublicationVenuePosition
2020 Ten simple rules for making training materials FAIR
abstract
Everything we do today is becoming more and more reliant on the use of computers. The field of biology is no exception; but most biologists receive little or no formal preparation for the increasingly computational aspects of their discipline. In consequence, informal training courses are often needed to plug the gaps; and the demand for such training is growing worldwide. To meet this demand, some training programs are being expanded, and new ones are being developed. Key to both scenarios is the creation of new course materials. Rather than starting from scratch, however, it's sometimes possible to repurpose materials that already exist. Yet finding suitable materials online can be difficult: They're often widely scattered across the internet or hidden in their home institutions, with no systematic way to find them. This is a common problem for all digital objects. The scientific community has attempted to address this issue by developing a set of rules (which have been called the Findable, Accessible, Interoperable and Reusable [FAIR] principles) to make such objects more findable and reusable. Here, we show how to apply these rules to help make training materials easier to find, (re)use, and adapt, for the benefit of all.
Leyla Jael Castro, Bérénice Batut, Melissa L. Burke, Mateusz Kuzak, Fotis E. Psomopoulos, Ricardo Arcila, Terri K. Attwood, Niall Beard, Denise Carvalho-Silva, Alexandros C. Dimopoulos, Victoria Dominguez Del Angel, Michel Dumontier, Kim T. Gurwitz, Roland Krause, Peter McQuilton, Loredana Le Pera, Sarah L. Morgan, Päivi Rauste, Allegra Via, Pascal Kahlem, Gabriella Rustici, Celia W. G. van Gelder, Patricia M. Palagi
PLoS Comput. Biol.19
2020 A framework to assess the quality and impact of bioinformatics training across ELIXIR
abstract
ELIXIR is a pan-European intergovernmental organisation for life science that aims to coordinate bioinformatics resources in a single infrastructure across Europe; bioinformatics training is central to its strategy, which aims to develop a training community that spans all ELIXIR member states. In an evidence-based approach for strengthening bioinformatics training programmes across Europe, the ELIXIR Training Platform, led by the ELIXIR EXCELERATE Quality and Impact Assessment Subtask in collaboration with the ELIXIR Training Coordinators Group, has implemented an assessment strategy to measure quality and impact of its entire training portfolio. Here, we present ELIXIR's framework for assessing training quality and impact, which includes the following: specifying assessment aims, determining what data to collect in order to address these aims, and our strategy for centralised data collection to allow for ELIXIR-wide analyses. In addition, we present an overview of the ELIXIR training data collected over the past 4 years. We highlight the importance of a coordinated and consistent data collection approach and the relevance of defining specific metrics and answer scales for consortium-wide analyses as well as for comparison of data across iterations of the same course.
Kim T. Gurwitz, Prakash Singh Gaur, Louisa J. Bellis, Lee D. Larcombe, Eva Alloza, Balint Laszlo Balint, Alexander Botzki, Jure Dimec, Victoria Dominguez Del Angel, Pedro L. Fernandes, Eija Korpelainen, Roland Krause, Mateusz Kuzak, Loredana Le Pera, Brane Leskosek, Jessica M. Lindvall, Diana Marek, Paula Andrea Martínez, Tuur Muyldermans, Ståle Nygård, Patricia M. Palagi, Hedi Peterson, Fotis E. Psomopoulos, Vojtech Spiwok, Celia W. G. van Gelder, Allegra Via, Marko Vidak, Daniel Wibberg, Sarah L. Morgan, Gabriella Rustici
PLoS Comput. Biol.26
2019 A new pan-European Train-the-Trainer programme for bioinformatics: pilot results on feasibility, utility and sustainability of learning
abstract
Demand for training life scientists in bioinformatics methods, tools and resources and computational approaches is urgent and growing. To meet this demand, new trainers must be prepared with effective teaching practices for delivering short hands-on training sessions-a specific type of education that is not typically part of professional preparation of life scientists in many countries. A new Train-the-Trainer (TtT) programme was created by adapting existing models, using input from experienced trainers and experts in bioinformatics, and from educational and cognitive sciences. This programme was piloted across Europe from May 2016 to January 2017. Preparation included drafting the training materials, organizing sessions to pilot them and studying this paradigm for its potential to support the development and delivery of future bioinformatics training by participants. Seven pilot TtT sessions were carried out, and this manuscript describes the results of the pilot year. Lessons learned include (i) support is required for logistics, so that new instructors can focus on their teaching; (ii) institutions must provide incentives to include training opportunities for those who want/need to become new or better instructors; (iii) formal evaluation of the TtT materials is now a priority; (iv) a strategy is needed to recruit, train and certify new instructor trainers (faculty); and (v) future evaluations must assess utility. Additionally, defining a flexible but rigorous and reliable process of TtT 'certification' may incentivize participants and will be considered in future.
Allegra Via, Terri K. Attwood, Pedro L. Fernandes, Sarah L. Morgan, Maria Victoria Schneider, Patricia M. Palagi, Gabriella Rustici, Rochelle E. Tractenberg
Briefings Bioinform.1
2018 Lesson Development for Open Source Software Best Practices Adoption
abstract
The "ELIXIR Training Platform" is partnering with The Carpentries (Software and Data Carpentry) to train life science researchers in computing and data management skills. The "ELIXIR Software development best practices" group, which is part of the ELIXIR Tools Platform, has proposed "Four simple recommendations to encourage best practices in research software" aiming to help researchers and developers to adopt Open Source Software (OSS) practices and thus improve the quality and sustainability of research software. In order to encourage researchers and developers to adopt the four recommendations (4OSS) and build FAIR software, we are developing specific and practical training materials, taking advantage of the Carpentries approach and experience in training material development and maintenance.
Mateusz Kuzak, Jennifer L. Harrow, Rafael C. Jiménez, Paula Andrea Martínez, Fotis E. Psomopoulos, Radka Svobodová Vareková, Allegra Via
eScience7
2015 The GOBLET training portal: a global repository of bioinformatics training materials, courses and trainers
abstract
SUMMARY: Rapid technological advances have led to an explosion of biomedical data in recent years. The pace of change has inspired new collaborative approaches for sharing materials and resources to help train life scientists both in the use of cutting-edge bioinformatics tools and databases and in how to analyse and interpret large datasets. A prototype platform for sharing such training resources was recently created by the Bioinformatics Training Network (BTN). Building on this work, we have created a centralized portal for sharing training materials and courses, including a catalogue of trainers and course organizers, and an announcement service for training events. For course organizers, the portal provides opportunities to promote their training events; for trainers, the portal offers an environment for sharing materials, for gaining visibility for their work and promoting their skills; for trainees, it offers a convenient one-stop shop for finding suitable training resources and identifying relevant training events and activities locally and worldwide. AVAILABILITY AND IMPLEMENTATION: http://mygoblet.org/training-portal.
Manuel Corpas, Rafael C. Jiménez, Erik Bongcam-Rudloff, Aidan Budd, Michelle D. Brazas, Pedro L. Fernandes, Bruno A. Gaëta, Celia W. G. van Gelder, Eija Korpelainen, Fran Lewitter, Annette McGrath, Daniel MacLean, Patricia M. Palagi, Kristian Rother, Jan Taylor, Allegra Via, Mick Watson, Maria Victoria Schneider, Terri K. Attwood
Bioinform.16
2013 Best practices in bioinformatics training for life scientists
abstract
The mountains of data thrusting from the new landscape of modern high-throughput biology are irrevocably changing biomedical research and creating a near-insatiable demand for training in data management and manipulation and data mining and analysis. Among life scientists, from clinicians to environmental researchers, a common theme is the need not just to use, and gain familiarity with, bioinformatics tools and resources but also to understand their underlying fundamental theoretical and practical concepts. Providing bioinformatics training to empower life scientists to handle and analyse their data efficiently, and progress their research, is a challenge across the globe. Delivering good training goes beyond traditional lectures and resource-centric demos, using interactivity, problem-solving exercises and cooperative learning to substantially enhance training quality and learning outcomes. In this context, this article discusses various pragmatic criteria for identifying training needs and learning objectives, for selecting suitable trainees and trainers, for developing and maintaining training skills and evaluating training quality. Adherence to these criteria may help not only to guide course organizers and trainers on the path towards bioinformatics training excellence but, importantly, also to improve the training experience for life scientists.
Allegra Via, Thomas Blicher, Erik Bongcam-Rudloff, Michelle D. Brazas, Catherine Brooksbank, Aidan Budd, Javier De Las Rivas, Jacqueline Dreyer, Pedro L. Fernandes, Celia W. G. van Gelder, Joachim Jacob, Rafael C. Jiménez, Jane E. Loveland, Federico Morán, Nicola J. Mulder, Tommi H. Nyrönen, Kristian Rother, Maria Victoria Schneider, Terri K. Attwood
Briefings Bioinform.1
2013 iAnn: an event sharing platform for the life sciences
abstract
SUMMARY: We present iAnn, an open source community-driven platform for dissemination of life science events, such as courses, conferences and workshops. iAnn allows automatic visualisation and integration of customised event reports. A central repository lies at the core of the platform: curators add submitted events, and these are subsequently accessed via web services. Thus, once an iAnn widget is incorporated into a website, it permanently shows timely relevant information as if it were native to the remote site. At the same time, announcements submitted to the repository are automatically disseminated to all portals that query the system. To facilitate the visualization of announcements, iAnn provides powerful filtering options and views, integrated in Google Maps and Google Calendar. All iAnn widgets are freely available. AVAILABILITY: http://iann.pro/iannviewer CONTACT: [email protected].
Rafael C. Jiménez, Juan P. Albar, Jong Bhak, Marie-Claude Blatter, Thomas Blicher, Michelle D. Brazas, Catherine Brooksbank, Aidan Budd, Javier De Las Rivas, Jacqueline Dreyer, Marc A. van Driel, Michael J. Dunn, Pedro L. Fernandes, Celia W. G. van Gelder, Henning Hermjakob, Vassilios Ioannidis, David Phillip Judge, Pascal Kahlem, Eija Korpelainen, Hans-Joachim Kraus, Jane E. Loveland, Christine Mayer, Jennifer McDowall, Federico Morán, Nicola J. Mulder, Tommi H. Nyrönen, Kristian Rother, Gustavo A. Salazar, Reinhard Schneider 0002, Allegra Via, Jose M. Villaveces, Maria Victoria Schneider, Terri K. Attwood, Manuel Corpas
Bioinform.30
2013 TiPs: a database of therapeutic targets in pathogens and associated tools
abstract
MOTIVATION: The need for new drugs and new targets is particularly compelling in an era that is witnessing an alarming increase of drug resistance in human pathogens. The identification of new targets of known drugs is a promising approach, which has proven successful in several cases. Here, we describe a database that includes information on 5153 putative drug-target pairs for 150 human pathogens derived from available drug-target crystallographic complexes. AVAILABILITY AND IMPLEMENTATION: The TiPs database is freely available at http://biocomputing.it/tips. CONTACT: [email protected] or [email protected].
Rosalba Lepore, Anna Tramontano, Allegra Via
Bioinform.3
2012 Bioinformatics Training Network (BTN): a community resource for bioinformatics trainers
abstract
Funding bodies are increasingly recognizing the need to provide graduates and researchers with access to short intensive courses in a variety of disciplines, in order both to improve the general skills base and to provide solid foundations on which researchers may build their careers. In response to the development of 'high-throughput biology', the need for training in the field of bioinformatics, in particular, is seeing a resurgence: it has been defined as a key priority by many Institutions and research programmes and is now an important component of many grant proposals. Nevertheless, when it comes to planning and preparing to meet such training needs, tension arises between the reward structures that predominate in the scientific community which compel individuals to publish or perish, and the time that must be devoted to the design, delivery and maintenance of high-quality training materials. Conversely, there is much relevant teaching material and training expertise available worldwide that, were it properly organized, could be exploited by anyone who needs to provide training or needs to set up a new course. To do this, however, the materials would have to be centralized in a database and clearly tagged in relation to target audiences, learning objectives, etc. Ideally, they would also be peer reviewed, and easily and efficiently accessible for downloading. Here, we present the Bioinformatics Training Network (BTN), a new enterprise that has been initiated to address these needs and review it, respectively, to similar initiatives and collections.
Maria Victoria Schneider, Peter Walter, Marie-Claude Blatter, James Watson, Michelle D. Brazas, Kristian Rother, Aidan Budd, Allegra Via, Celia W. G. van Gelder, Joachim Jacob, Pedro L. Fernandes, Tommi H. Nyrönen, Javier De Las Rivas, Thomas Blicher, Rafael C. Jiménez, Jane E. Loveland, Jennifer McDowall, Philip Jones, Brendan W. Vaughan, Rodrigo Lopez, Terri K. Attwood, Catherine Brooksbank
Briefings Bioinform.8
2011 Ten Simple Rules for Developing a Short Bioinformatics Training Course
abstract
[No abstract available]
Allegra Via, Javier De Las Rivas, Terri K. Attwood, David Landsman, Michelle D. Brazas, Jack A. M. Leunissen, Anna Tramontano, Maria Victoria Schneider
PLoS Comput. Biol.1
2010 Bioinformatics training: a review of challenges, actions and support requirements
abstract
As bioinformatics becomes increasingly central to research in the molecular life sciences, the need to train non-bioinformaticians to make the most of bioinformatics resources is growing. Here, we review the key challenges and pitfalls to providing effective training for users of bioinformatics services, and discuss successful training strategies shared by a diverse set of bioinformatics trainers. We also identify steps that trainers in bioinformatics could take together to advance the state of the art in current training practices. The ideas presented in this article derive from the first Trainer Networking Session held under the auspices of the EU-funded SLING Integrating Activity, which took place in November 2009.
Maria Victoria Schneider, James Watson, Terri K. Attwood, Kristian Rother, Aidan Budd, Jennifer McDowall, Allegra Via, Pedro L. Fernandes, Tommi H. Nyrönen, Thomas Blicher, Philip Jones, Marie-Claude Blatter, Javier De Las Rivas, David Phillip Judge, Wouter van der Gool, Catherine Brooksbank
Briefings Bioinform.7
2009 A structure filter for the Eukaryotic Linear Motif Resource
abstract
BACKGROUND: Many proteins are highly modular, being assembled from globular domains and segments of natively disordered polypeptides. Linear motifs, short sequence modules functioning independently of protein tertiary structure, are most abundant in natively disordered polypeptides but are also found in accessible parts of globular domains, such as exposed loops. The prediction of novel occurrences of known linear motifs attempts the difficult task of distinguishing functional matches from stochastically occurring non-functional matches. Although functionality can only be confirmed experimentally, confidence in a putative motif is increased if a motif exhibits attributes associated with functional instances such as occurrence in the correct taxonomic range, cellular compartment, conservation in homologues and accessibility to interacting partners. Several tools now use these attributes to classify putative motifs based on confidence of functionality. RESULTS: Current methods assessing motif accessibility do not consider much of the information available, either predicting accessibility from primary sequence or regarding any motif occurring in a globular region as low confidence. We present a method considering accessibility and secondary structural context derived from experimentally solved protein structures to rectify this situation. Putatively functional motif occurrences are mapped onto a representative domain, given that a high quality reference SCOP domain structure is available for the protein itself or a close relative. Candidate motifs can then be scored for solvent-accessibility and secondary structure context. The scores are calibrated on a benchmark set of experimentally verified motif instances compared with a set of random matches. A combined score yields 3-fold enrichment for functional motifs assigned to high confidence classifications and 2.5-fold enrichment for random motifs assigned to low confidence classifications. The structure filter is implemented as a pipeline with both a graphical interface via the ELM resource http://elm.eu.org/ and through a Web Service protocol. CONCLUSION: New occurrences of known linear motifs require experimental validation as the bioinformatics tools currently have limited reliability. The ELM structure filter will aid users assessing candidate motifs presenting in globular structural regions. Most importantly, it will help users to decide whether to expend their valuable time and resources on experimental testing of interesting motif candidates.
Allegra Via, Cathryn M. Gould, Christine Gemünd, Toby J. Gibson, Manuela Helmer-Citterich
BMC Bioinform.1
2008 FunClust: a web server for the identification of structural motifs in a set of non-homologous protein structures
abstract
BACKGROUND: The occurrence of very similar structural motifs brought about by different parts of non homologous proteins is often indicative of a common function. Indeed, relatively small local structures can mediate binding to a common partner, be it a protein, a nucleic acid, a cofactor or a substrate. While it is relatively easy to identify short amino acid or nucleotide sequence motifs in a given set of proteins or genes, and many methods do exist for this purpose, much more challenging is the identification of common local substructures, especially if they are formed by non consecutive residues in the sequence. RESULTS: Here we describe a publicly available tool, able to identify common structural motifs shared by different non homologous proteins in an unsupervised mode. The motifs can be as short as three residues and need not to be contiguous or even present in the same order in the sequence. Users can submit a set of protein structures deemed or not to share a common function (e.g. they bind similar ligands, or share a common epitope). The server finds and lists structural motifs composed of three or more spatially well conserved residues shared by at least three of the submitted structures. The method uses a local structural comparison algorithm to identify subsets of similar amino acids between each pair of input protein chains and a clustering procedure to group similarities shared among different structure pairs. CONCLUSIONS: FunClust is fast, completely sequence independent, and does not need an a priori knowledge of the motif to be found. The output consists of a list of aligned structural matches displayed in both tabular and graphical form. We show here examples of its usefulness by searching for the largest common structural motifs in test sets of non homologous proteins and showing that the identified motifs correspond to a known common functional feature.
Gabriele Ausiello, Pier Federico Gherardini, Paolo Marcatili, Anna Tramontano, Allegra Via, Manuela Helmer-Citterich
BMC Bioinform.5
2007 Local comparison of protein structures highlights cases of convergent evolution in analogous functional sites
abstract
BACKGROUND: We performed an exhaustive search for local structural similarities in an ensemble of non-redundant protein functional sites. With the purpose of finding new examples of convergent evolution, we selected only those matching sites composed of structural regions whose residue order is inverted in the relative protein sequences. RESULTS: A novel case of local analogy was detected between members of the ABC transporter and of the HprK/P families in their ATP binding site. This case cannot be derived by events of circular permutation since the residues of one of the region pairs are located in reverse order in the sequence of the two protein families. One of the analogous binding sites, the one identified in HprK/P, is known to also bind pyrophosphate, which is used as preferred energy source in its kinase and phosphorylase activity. CONCLUSION: The discovery of this striking molecular similarity, also associated to a functional similarity, may help in suggesting new experiments aimed at a deeper understanding of members of the ABC transporter family known to be involved in many serious human diseases.
Gabriele Ausiello, Daniele Peluso, Allegra Via, Manuela Helmer-Citterich
BMC Bioinform.3
2007 False occurrences of functional motifs in protein sequences highlight evolutionary constraints
abstract
BACKGROUND: False occurrences of functional motifs in protein sequences can be considered as random events due solely to the sequence composition of a proteome. Here we use a numerical approach to investigate the random appearance of functional motifs with the aim of addressing biological questions such as: How are organisms protected from undesirable occurrences of motifs otherwise selected for their functionality? Has the random appearance of functional motifs in protein sequences been affected during evolution? RESULTS: Here we analyse the occurrence of functional motifs in random sequences and compare it to that observed in biological proteomes; the behaviour of random motifs is also studied. Most motifs exhibit a number of false positives significantly similar to the number of times they appear in randomized proteomes (=expected number of false positives). Interestingly, about 3% of the analysed motifs show a different kind of behaviour and appear in biological proteomes less than they do in random sequences. In some of these cases, a mechanism of evolutionary negative selection is apparent; this helps to prevent unwanted functionalities which could interfere with cellular mechanisms. CONCLUSION: Our thorough statistical and biological analysis showed that there are several mechanisms and evolutionary constraints both of which affect the appearance of functional motifs in protein sequences.
Allegra Via, Pier Federico Gherardini, Enrico Ferraro, Gabriele Ausiello, Gianpaolo Scalia Tomba, Manuela Helmer-Citterich
BMC Bioinform.1
2006 A novel structure-based encoding for machine-learning applied to the inference of SH3 domain specificity
abstract
MOTIVATION: Unravelling the rules underlying protein-protein and protein-ligand interactions is a crucial step in understanding cell machinery. Peptide recognition modules (PRMs) are globular protein domains which focus their binding targets on short protein sequences and play a key role in the frame of protein-protein interactions. High-throughput techniques permit the whole proteome scanning of each domain, but they are characterized by a high incidence of false positives. In this context, there is a pressing need for the development of in silico experiments to validate experimental results and of computational tools for the inference of domain-peptide interactions. RESULTS: We focused on the SH3 domain family and developed a machine-learning approach for inferring interaction specificity. SH3 domains are well-studied PRMs which typically bind proline-rich short sequences characterized by the PxxP consensus. The binding information is known to be held in the conformation of the domain surface and in the short sequence of the peptide. Our method relies on interaction data from high-throughput techniques and benefits from the integration of sequence and structure data of the interacting partners. Here, we propose a novel encoding technique aimed at representing binding information on the basis of the domain-peptide contact residues in complexes of known structure. Remarkably, the new encoding requires few variables to represent an interaction, thus avoiding the 'curse of dimension'. Our results display an accuracy >90% in detecting new binders of known SH3 domains, thus outperforming neural models on standard binary encodings, profile methods and recent statistical predictors. The method, moreover, shows a generalization capability, inferring specificity of unknown SH3 domains displaying some degree of similarity with the known data.
Enrico Ferraro, Allegra Via, Gabriele Ausiello, Manuela Helmer-Citterich
Bioinform.2
2005 Seq2Struct: a resource for establishing sequence-structure links
abstract
UNLABELLED: Several methods for establishing cross-links between Protein Data Bank (PDB) structures or Structural Classification of Proteins (SCOP) domains and Swiss-Prot + TrEMBL sequences (or vice versa) rely on database annotations. Alternatively, sequence alignment procedures can be used. In this study, we describe Seq2Struct, a web resource for the identification of sequence-structure links. The resource consists of an exhaustive collection of annotated links between Swiss-Prot + TrEMBL and PDB + SCOP database entries. Links are based on pre-established highly reliable thresholds and stored in a relational database, which has been enhanced using annotations derived from Swiss-Prot, PDB, SCOP, GOA and DSSP databases. The Seq2Struct database contents, supported by a WWW web interface, can be queried both online and downloaded. AVAILABILITY: The Seq2Struct resource, with related documentation, is available at http://surface.bio.uniroma2.it/seq2struct/ CONTACT: [email protected].
Allegra Via, Andreas Zanzoni, Manuela Helmer-Citterich
Bioinform.1
2005 Query3d: a new method for high-throughput analysis of functional residues in protein structures
abstract
BACKGROUND: The identification of local similarities between two protein structures can provide clues of a common function. Many different methods exist for searching for similar subsets of residues in proteins of known structure. However, the lack of functional and structural information on single residues, together with the low level of integration of this information in comparison methods, is a limitation that prevents these methods from being fully exploited in high-throughput analyses. RESULTS: Here we describe Query3d, a program that is both a structural DBMS (Database Management System) and a local comparison method. The method conserves a copy of all the residues of the Protein Data Bank annotated with a variety of functional and structural information. New annotations can be easily added from a variety of methods and known databases. The algorithm makes it possible to create complex queries based on the residues' function and then to compare only subsets of the selected residues. Functional information is also essential to speed up the comparison and the analysis of the results. CONCLUSION: With Query3d, users can easily obtain statistics on how many and which residues share certain properties in all proteins of known structure. At the same time, the method also finds their structural neighbours in the whole PDB. Programs and data can be accessed through the PdbFun web interface.
Gabriele Ausiello, Allegra Via, Manuela Helmer-Citterich
BMC Bioinform.2
2005 A neural strategy for the inference of SH3 domain-peptide interaction specificity
abstract
BACKGROUND: The SH3 domain family is one of the most representative and widely studied cases of so-called Peptide Recognition Modules (PRM). The polyproline II motif PxxP that generally characterizes its ligands does not reflect the complex interaction spectrum of the over 1500 different SH3 domains, and the requirement of a more refined knowledge of their specificity implies the setting up of appropriate experimental and theoretical strategies. Due to the limitations of the current technology for peptide synthesis, several experimental high-throughput approaches have been devised to elucidate protein-protein interaction mechanisms. Such approaches can rely on and take advantage of computational techniques, such as regular expressions or position specific scoring matrices (PSSMs) to pre-process entire proteomes in the search for putative SH3 targets. In this regard, a reliable inference methodology to be used for reducing the sequence space of putative binding peptides represents a valuable support for molecular and cellular biologists. RESULTS: Using as benchmark the peptide sequences obtained from in vitro binding experiments, we set up a neural network model that performs better than PSSM in the detection of SH3 domain interactors. In particular our model is more precise in its predictions, even if its performance can vary among different SH3 domains and is strongly dependent on the number of binding peptides in the benchmark. CONCLUSION: We show that a neural network can be more effective than standard methods in SH3 domain specificity detection. Neural classifiers identify general SH3 domain binders and domain-specific interactors from a PxxP peptide population, provided that there are a sufficient proportion of true positives in the training sets. This capability can also improve peptide selection for library definition in array experiments. Further advances can be achieved, including properly encoded domain sequences and structural information as input for a global neural network.
Enrico Ferraro, Allegra Via, Gabriele Ausiello, Manuela Helmer-Citterich
BMC Bioinform.2
2004 Phospho.ELM: A database of experimentally verified phosphorylation sites in eukaryotic proteins
abstract
BACKGROUND: Post-translational phosphorylation is one of the most common protein modifications. Phosphoserine, threonine and tyrosine residues play critical roles in the regulation of many cellular processes. The fast growing number of research reports on protein phosphorylation points to a general need for an accurate database dedicated to phosphorylation to provide easily retrievable information on phosphoproteins. DESCRIPTION: Phospho.ELM http://phospho.elm.eu.org is a new resource containing experimentally verified phosphorylation sites manually curated from the literature and is developed as part of the ELM (Eukaryotic Linear Motif) resource. Phospho.ELM constitutes the largest searchable collection of phosphorylation sites available to the research community. The Phospho.ELM entries store information about substrate proteins with the exact positions of residues known to be phosphorylated by cellular kinases. Additional annotation includes literature references, subcellular compartment, tissue distribution, and information about the signaling pathways involved as well as links to the molecular interaction database MINT. Phospho.ELM version 2.0 contains 1703 phosphorylation site instances for 556 phosphorylated proteins. CONCLUSION: Phospho.ELM will be a valuable tool both for molecular biologists working on protein phosphorylation sites and for bioinformaticians developing computational predictions on the specificity of phosphorylation reactions.
Francesca Diella, Scott Cameron, Christine Gemünd, Rune Linding, Allegra Via, Bernhard Küster, Thomas Sicheritz-Pontén, Nikolaj Blom, Toby J. Gibson
BMC Bioinform.5
2004 A structural study for the optimisation of functional motifs encoded in protein sequences
abstract
BACKGROUND: A large number of PROSITE patterns select false positives and/or miss known true positives. It is possible that--at least in some cases--the weak specificity and/or sensitivity of a pattern is due to the fact that one, or maybe more, functional and/or structural key residues are not represented in the pattern. Multiple sequence alignments are commonly used to build functional sequence patterns. If residues structurally conserved in proteins sharing a function cannot be aligned in a multiple sequence alignment, they are likely to be missed in a standard pattern construction procedure. RESULTS: Here we present a new procedure aimed at improving the sensitivity and/ or specificity of poorly-performing patterns. The procedure can be summarised as follows: 1. residues structurally conserved in different proteins, that are true positives for a pattern, are identified by means of a computational technique and by visual inspection. 2. the sequence positions of the structurally conserved residues falling outside the pattern are used to build extended sequence patterns. 3. the extended patterns are optimised on the SWISS-PROT database for their sensitivity and specificity. The method was applied to eight PROSITE patterns. Whenever structurally conserved residues are found in the surface region close to the pattern (seven out of eight cases), the addition of information inferred from structural analysis is shown to improve pattern selectivity and in some cases selectivity and sensitivity as well. In some of the cases considered the procedure allowed the identification of functionally interesting residues, whose biological role is also discussed. CONCLUSION: Our method can be applied to any type of functional motif or pattern (not only PROSITE ones) which is not able to select all and only the true positive hits and for which at least two true positive structures are available. The computational technique for the identification of structurally conserved residues is already available on request and will be soon accessible on our web server. The procedure is intended for the use of pattern database curators and of scientists interested in a specific protein family for which no specific or selective patterns are yet available.
Allegra Via, Manuela Helmer-Citterich
BMC Bioinform.1