Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Anne-Lise Veuthey

dblp:23/1243 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
interpretable classification
0.212014
Motifs tree: a new method for predicting post-translational modifications · Bioinform. 2014
Bioinformatics and computational biology › sequence analysis › motif discovery
pattern discovery
0.212014
Motifs tree: a new method for predicting post-translational modifications · Bioinform. 2014
Bioinformatics and computational biology › proteomics
post-translational modification prediction
0.212014
Motifs tree: a new method for predicting post-translational modifications · Bioinform. 2014

Methods — techniques the papers use, named apart from their topics

pattern discovery · 0.2genetic algorithm · 0.2c4.5 · 0.2web-based search · 0.1
YearPublicationVenuePosition
2014 Motifs tree: a new method for predicting post-translational modifications
abstract
MOTIVATION: Post-translational modifications (PTMs) are important steps in the maturation of proteins. Several models exist to predict specific PTMs, from manually detected patterns to machine learning methods. On one hand, the manual detection of patterns does not provide the most efficient classifiers and requires an important workload, and on the other hand, models built by machine learning methods are hard to interpret and do not increase biological knowledge. Therefore, we developed a novel method based on patterns discovery and decision trees to predict PTMs. The proposed algorithm builds a decision tree, by coupling the C4.5 algorithm with genetic algorithms, producing high-performance white box classifiers. Our method was tested on the initiator methionine cleavage (IMC) and N(α)-terminal acetylation (N-Ac), two of the most common PTMs. RESULTS: The resulting classifiers perform well when compared with existing models. On a set of eukaryotic proteins, they display a cross-validated Matthews correlation coefficient of 0.83 (IMC) and 0.65 (N-Ac). When used to predict potential substrates of N-terminal acetyltransferaseB and N-terminal acetyltransferaseC, our classifiers display better performance than the state of the art. Moreover, we present an analysis of the model predicting IMC for Homo sapiens proteins and demonstrate that we are able to extract experimentally known facts without prior knowledge. Those results validate the fact that our method produces white box models. AVAILABILITY AND IMPLEMENTATION: Predictors for IMC and N-Ac and all datasets are freely available at http://terminus.unige.ch/.
Christophe Charpilloz, Anne-Lise Veuthey, Bastien Chopard, Jean-Luc Falcone
Bioinform.2
2013 Application of text-mining for updating protein post-translational modification annotation in UniProtKB
abstract
BACKGROUND: The annotation of protein post-translational modifications (PTMs) is an important task of UniProtKB curators and, with continuing improvements in experimental methodology, an ever greater number of articles are being published on this topic. To help curators cope with this growing body of information we have developed a system which extracts information from the scientific literature for the most frequently annotated PTMs in UniProtKB. RESULTS: The procedure uses a pattern-matching and rule-based approach to extract sentences with information on the type and site of modification. A ranked list of protein candidates for the modification is also provided. For PTM extraction, precision varies from 57% to 94%, and recall from 75% to 95%, according to the type of modification. The procedure was used to track new publications on PTMs and to recover potential supporting evidence for phosphorylation sites annotated based on the results of large scale proteomics experiments. CONCLUSIONS: The information retrieval and extraction method we have developed in this study forms the basis of a simple tool for the manual curation of protein post-translational modifications in UniProtKB/Swiss-Prot. Our work demonstrates that even simple text-mining tools can be effectively adapted for database curation tasks, providing that a thorough understanding of the working process and requirements are first obtained. This system can be accessed at http://eagl.unige.ch/PTM/.
Anne-Lise Veuthey, Alan J. Bridge, Julien Gobeill, Patrick Ruch, Johanna R. McEntyre, Lydie Bougueleret, Ioannis Xenarios
BMC Bioinform.1
2010 Easy retrieval of single amino-acid polymorphisms and phenotype information using SwissVar
abstract
SUMMARY: The SwissVar portal provides access to a comprehensive collection of single amino acid polymorphisms and diseases in the UniProtKB/Swiss-Prot database via a unique search engine. In particular, it gives direct access to the newly improved Swiss-Prot variant pages. The key strength of this portal is that it provides a possibility to query for similar diseases, as well as the underlying protein products and the molecular details of each variant. In the context of the recently proposed molecular view on diseases, the SwissVar portal should be in a unique position to provide valuable information for researchers and to advance research in this area. AVAILABILITY: The SwissVar portal is available at www.expasy.org/swissvar CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anaïs Mottaz, Fabrice P. A. David, Anne-Lise Veuthey, Yum Lina Yip
Bioinform.3
2008 Gene Ontology density estimation and discourse analysis for automatic GeneRiF extraction
abstract
BACKGROUND: This paper describes and evaluates a sentence selection engine that extracts a GeneRiF (Gene Reference into Functions) as defined in ENTREZ-Gene based on a MEDLINE record. Inputs for this task include both a gene and a pointer to a MEDLINE reference. In the suggested approach we merge two independent sentence extraction strategies. The first proposed strategy (LASt) uses argumentative features, inspired by discourse-analysis models. The second extraction scheme (GOEx) uses an automatic text categorizer to estimate the density of Gene Ontology categories in every sentence; thus providing a full ranking of all possible candidate GeneRiFs. A combination of the two approaches is proposed, which also aims at reducing the size of the selected segment by filtering out non-content bearing rhetorical phrases. RESULTS: Based on the TREC-2003 Genomics collection for GeneRiF identification, the LASt extraction strategy is already competitive (52.78%). When used in a combined approach, the extraction task clearly shows improvement, achieving a Dice score of over 57% (+10%). CONCLUSIONS: Argumentative representation levels and conceptual density estimation using Gene Ontology contents appear complementary for functional annotation in proteomics.
Julien Gobeill, Imad Tbahriti, Frédéric Ehrler, Anaïs Mottaz, Anne-Lise Veuthey, Patrick Ruch
BMC Bioinform.5
2008 Mapping proteins to disease terminologies: from UniProt to MeSH
abstract
BACKGROUND: Although the UniProt KnowledgeBase is not a medical-oriented database, it contains information on more than 2,000 human proteins involved in pathologies. However, these annotations are not standardized, which impairs the interoperability between biological and clinical resources. In order to make these data easily accessible to clinical researchers, we have developed a procedure to link diseases described in the UniProtKB/Swiss-Prot entries to the MeSH disease terminology. RESULTS: We mapped disease names extracted either from the UniProtKB/Swiss-Prot entry comment lines or from the corresponding OMIM entry to the MeSH. Different methods were assessed on a benchmark set of 200 disease names manually mapped to MeSH terms. The performance of the retained procedure in term of precision and recall was 86% and 64% respectively. Using the same procedure, more than 3,000 disease names in Swiss-Prot were mapped to MeSH with comparable efficiency. CONCLUSIONS: This study is a first attempt to link proteins in UniProtKB to the medical resources. The indexing we provided will help clinicians and researchers navigate from diseases to genes and from genes to diseases in an efficient way. The mapping is available at: http://research.isb-sib.ch/unimed.
Anaïs Mottaz, Yum Lina Yip, Patrick Ruch, Anne-Lise Veuthey
BMC Bioinform.4
2005 Latent Argumentative Pruning for Compact MEDLINE Indexing
Patrick Ruch, Robert H. Baud, Johan Marty, Antoine Geissbühler, Imad Tbahriti, Anne-Lise Veuthey
AIME6
2005 GPSDB: a new database for synonyms expansion of gene and protein names
abstract
UNLABELLED: We present a new database, GPSDB (Gene and Protein Synonyms DataBase) which collects gene/protein names, in a species specific way, from 14 main biological resources. A web-based search interface gives access to the database: given a gene/protein name, it retrieves all synonyms for this entity and queries Medline with a set of user-selected terms. AVAILABILITY: GPSDB is freely available from http://biomint.oefai.at/ CONTACT: [email protected].
Violaine Pillet, Marc Zehnder, Alexander K. Seewald, Anne-Lise Veuthey, Johann Petrak
Bioinform.4