Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Julio Collado-Vides

dblp:28/5591 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
0since 2021 · last 2019
0000-0001-8780-7664ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-authorArtificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Automata and formal languages · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › ontology
ontology development
0.412019
MCO: towards an ontology and unified vocabulary for a framework-based annotation of microbial growth conditions · Bioinform. 2019
Bioinformatics and computational biology › protein structure prediction › template-based modeling
homology modeling
0.112007
TFmodeller: comparative modelling of protein-DNA complexes · Bioinform. 2007
Bioinformatics and computational biology
protein structure prediction
0.112007
TFmodeller: comparative modelling of protein-DNA complexes · Bioinform. 2007
Bioinformatics and computational biology › gene regulation
regulatory genomics
0.031998
Prediction of transcriptional regulatory sites in the complete genome sequence of Escherichia coli K-12 · Bioinform. 1998
Syntactic recognition of regulatory regions in Escherichia coli · Comput. Appl. Biosci. 1996
The search for a grammatical theory of gene regulation is formally justified by showing the inadequacy of context-free grammars · Comput. Appl. Biosci. 1991
Bioinformatics and computational biology
functional genomics
0.012002
A powerful non-homology method for the prediction of operons in prokaryotes · ISMB 2002
Bioinformatics and computational biology › genome annotation
operon prediction
0.012002
A powerful non-homology method for the prediction of operons in prokaryotes · ISMB 2002
Bioinformatics and computational biology
comparative genomics
0.012002
A powerful non-homology method for the prediction of operons in prokaryotes · ISMB 2002
Automata and formal languages
grammar formalisms
0.011991
The search for a grammatical theory of gene regulation is formally justified by showing the inadequacy of context-free grammars · Comput. Appl. Biosci. 1991
Bioinformatics and computational biology
genome annotation
0.011998
Prediction of transcriptional regulatory sites in the complete genome sequence of Escherichia coli K-12 · Bioinform. 1998

Methods — techniques the papers use, named apart from their topics

ontology curation · 0.4homology modelling · 0.1evolutionary contact matrix · 0.1weight matrix · 0.0log likelihood · 0.0string search · 0.0prolog parsing · 0.0formal grammar · 0.0
YearPublicationVenuePosition
2019 Towards the Prokaryotic Regulation Ontology: An Ontological Model to Infer Gene Regulation Physiology from Mechanisms in Bacteria
abstract
Here we present a formal ontological model that explicitly represents regulatory interactions among the main objects involved in transcriptional regulation in bacteria. These formal relations allow the inference of gene regulation physiology from gene regulation mechanisms. The automatically instantiated classes can be used to assist in the mechanistic interpretation of gene expression experiments done at the physiological level, such as RNA-seq. This is the first step to develop a more comprehensive ontology focused on prokaryotic gene regulation. The ontology is available at https://github.com/prokaryotic-regulation-ontology
Citlalli Mejía, Julio Collado-Vides
KEOD2
2019 MCO: towards an ontology and unified vocabulary for a framework-based annotation of microbial growth conditions
abstract
MOTIVATION: A major component in increasing our understanding of the biology of an organism is the mapping of its genotypic potential into its phenotypic expression profiles. This mapping is executed by the machinery of gene regulation, which is essentially studied by changes in growth conditions. Although many efforts have been made to systematize the annotation of experimental conditions in microbiology, the available annotations are not based on a consistent and controlled vocabulary, making difficult the identification of biologically meaningful comparisons of knowledge derived from different experiments or laboratories. RESULTS: We curated terms related to experimental conditions that affect gene expression in Escherichia coli K-12. Since this is the best-studied microorganism, the collected terms are the seed for the Microbial Conditions Ontology (MCO), a controlled and structured vocabulary that can be expanded to annotate microbial conditions in general. Moreover, we developed an annotation framework to describe experimental conditions, providing the foundation to identify regulatory networks that operate under particular conditions. AVAILABILITY AND IMPLEMENTATION: As far as we know, MCO is the first ontology for growth conditions of any bacterial organism, and it is available at http://regulondb.ccg.unam.mx and https://github.com/microbial-conditions-ontology. Furthermore, we will disseminate MCO throughout the Open Biological and Biomedical Ontology (OBO) Foundry in order to set a standard for the annotation of gene expression data. This will enable comparison of data from diverse data sources. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Víctor H. Tierrafría, Citlalli Mejía, J. M. Camacho-Zaragoza, Heladia Salgado, K. Alquicira, Chu Ishida, Socorro Gama-Castro, Julio Collado-Vides
Bioinform.8
2008 Prediction of TF target sites based on atomistic models of protein-DNA complexes
abstract
BACKGROUND: The specific recognition of genomic cis-regulatory elements by transcription factors (TFs) plays an essential role in the regulation of coordinated gene expression. Studying the mechanisms determining binding specificity in protein-DNA interactions is thus an important goal. Most current approaches for modeling TF specific recognition rely on the knowledge of large sets of cognate target sites and consider only the information contained in their primary sequence. RESULTS: Here we describe a structure-based methodology for predicting sequence motifs starting from the coordinates of a TF-DNA complex. Our algorithm combines information regarding the direct and indirect readout of DNA into an atomistic statistical model, which is used to estimate the interaction potential. We first measure the ability of our method to correctly estimate the binding specificities of eight prokaryotic and eukaryotic TFs that belong to different structural superfamilies. Secondly, the method is applied to two homology models, finding that sampling of interface side-chain rotamers remarkably improves the results. Thirdly, the algorithm is compared with a reference structural method based on contact counts, obtaining comparable predictions for the experimental complexes and more accurate sequence motifs for the homology models. CONCLUSION: Our results demonstrate that atomic-detail structural information can be feasibly used to predict TF binding sites. The computational method presented here is universal and might be applied to other systems involving protein-DNA recognition.
Vladimir Espinosa Angarica, Abel González-Pérez, Ana Tereza Ribeiro de Vasconcelos, Julio Collado-Vides, Bruno Contreras-Moreira
BMC Bioinform.4
2007 NLP-Based Curation of Bacterial Regulatory Networks
Carlos Rodríguez Penagos, Heladia Salgado, Irma Martínez-Flores, Julio Collado-Vides
CICLing4
2007 TFmodeller: comparative modelling of protein-DNA complexes
abstract
UNLABELLED: Interactions between proteins and DNA molecules lie at the core of the fundamental cellular processes such as transcriptional regulation. Some of these interactions have been experimentally described at atomic scale, but the molecular details of many others remain to be discovered. TFmodeller exploits the current knowledge about protein-DNA interfaces contained in the Protein Data Bank and uses it to model similar interfaces related by homology. Results are emailed to the user and include an evolutionary contact matrix, a schematic representation of the putative binding interface and atomic coordinates of the modelled complex. The library of complexes used by TFmodeller is updated on a weekly basis and is available for download. AVAILABILITY: TFmodeller and its web service interface are free for academic users at http://www.ccg.unam.mx/tfmodeller.
Bruno Contreras-Moreira, Pierre-Alain Branger, Julio Collado-Vides
Bioinform.3
2007 Automatic reconstruction of a bacterial regulatory network using Natural Language Processing
abstract
BACKGROUND: Manual curation of biological databases, an expensive and labor-intensive process, is essential for high quality integrated data. In this paper we report the implementation of a state-of-the-art Natural Language Processing system that creates computer-readable networks of regulatory interactions directly from different collections of abstracts and full-text papers. Our major aim is to understand how automatic annotation using Text-Mining techniques can complement manual curation of biological databases. We implemented a rule-based system to generate networks from different sets of documents dealing with regulation in Escherichia coli K-12. RESULTS: Performance evaluation is based on the most comprehensive transcriptional regulation database for any organism, the manually-curated RegulonDB, 45% of which we were able to recreate automatically. From our automated analysis we were also able to find some new interactions from papers not already curated, or that were missed in the manual filtering and review of the literature. We also put forward a novel Regulatory Interaction Markup Language better suited than SBML for simultaneously representing data of interest for biologists and text miners. CONCLUSION: Manual curation of the output of automatic processing of text is a good way to complement a more detailed review of the literature, either for validating the results of what has been already annotated, or for discovering facts and information that might have been overlooked at the triage or curation stages.
Carlos Rodríguez Penagos, Heladia Salgado, Irma Martínez-Flores, Julio Collado-Vides
BMC Bioinform.4
2007 Development of Genomic Sciences in Mexico: A Good Start and a Long Way to Go
abstract
DOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone.
Rafael Palacios, Julio Collado-Vides
PLoS Comput. Biol.2
2007 Metabolic Reconstruction and Modeling of Nitrogen Fixation in Rhizobium etli
abstract
Rhizobiaceas are bacteria that fix nitrogen during symbiosis with plants. This symbiotic relationship is crucial for the nitrogen cycle, and understanding symbiotic mechanisms is a scientific challenge with direct applications in agronomy and plant development. Rhizobium etli is a bacteria which provides legumes with ammonia (among other chemical compounds), thereby stimulating plant growth. A genome-scale approach, integrating the biochemical information available for R. etli, constitutes an important step toward understanding the symbiotic relationship and its possible improvement. In this work we present a genome-scale metabolic reconstruction (iOR363) for R. etli CFN42, which includes 387 metabolic and transport reactions across 26 metabolic pathways. This model was used to analyze the physiological capabilities of R. etli during stages of nitrogen fixation. To study the physiological capacities in silico, an objective function was formulated to simulate symbiotic nitrogen fixation. Flux balance analysis (FBA) was performed, and the predicted active metabolic pathways agreed qualitatively with experimental observations. In addition, predictions for the effects of gene deletions during nitrogen fixation in Rhizobia in silico also agreed with reported experimental data. Overall, we present some evidence supporting that FBA of the reconstructed metabolic network for R. etli provides results that are in agreement with physiological observations. Thus, as for other organisms, the reconstructed genome-scale metabolic network provides an important framework which allows us to compare model predictions with experimental measurements and eventually generate hypotheses on ways to improve nitrogen fixation.
Osbaldo Resendis-Antonio, Jennifer L. Reed, Sergio Encarnación-Guevara, Julio Collado-Vides, Bernhard O. Palsson
PLoS Comput. Biol.4
2006 The comprehensive updated regulatory network of Escherichia coli K-12
abstract
BACKGROUND: Escherichia coli is the model organism for which our knowledge of its regulatory network is the most extensive. Over the last few years, our project has been collecting and curating the literature concerning E. coli transcription initiation and operons, providing in both the RegulonDB and EcoCyc databases the largest electronically encoded network available. A paper published recently by Ma et al. (2004) showed several differences in the versions of the network present in these two databases. Discrepancies have been corrected, annotations from this and other groups (Shen-Orr et al., 2002) have been added, making the RegulonDB and EcoCyc databases the largest comprehensive and constantly curated regulatory network of E. coli K-12. RESULTS: Several groups have been using these curated data as part of their bioinformatics and systems biology projects, in combination with external data obtained from other sources, thus enlarging the dataset initially obtained from either RegulonDB or EcoCyc of the E. coli K12 regulatory network. We kindly obtained from the groups of Uri Alon and Hong-Wu Ma the interactions they have added to enrich their public versions of the E. coli regulatory network. These were used to search for original references and curate them with the same standards we use regularly, adding in several cases the original references (instead of reviews or missing references), as well as adding the corresponding experimental evidence codes. We also corrected all discrepancies in the two databases available as explained below. CONCLUSION: One hundred and fifty new interactions have been added to our databases as a result of this specific curation effort, in addition to those added as a result of our continuous curation work. RegulonDB gene names are now based on those of EcoCyc to avoid confusion due to gene names and synonyms, and the public releases of RegulonDB and EcoCyc are henceforth synchronized to avoid confusion due to different versions. Public flat files are available providing direct access to the regulatory network interactions thus avoiding errors due to differences in database modelling and representation. The regulatory network available in RegulonDB and EcoCyc is the most comprehensive and regularly updated electronically-encoded regulatory network of E. coli K-12.
Heladia Salgado, Alberto Santos-Zavaleta, Socorro Gama-Castro, Martín Peralta-Gil, Mónica I. Peñaloza-Spínola, Agustino Martínez-Antonio, Peter D. Karp, Julio Collado-Vides
BMC Bioinform.8
2002 A powerful non-homology method for the prediction of operons in prokaryotes
abstract
Abstract Motivation: The prediction of the transcription unit organization of genomes is an important clue in the inference of functional relationships of genes, the interpretation and evaluation of transcriptome experiments, and the overall inference of the regulatory networks governing the expression of genes in response to the environment. Though several methods have been devised to predict operons, most need a high characterization of the genome analysed. Log-likelihoods derived from inter-genic distance distributions work surprisingly well to predict operons in Escherichia coli and are available for any genome as soon as the gene sets are predicted. Results: Here we provide evidence that the very same method is applicable to any prokaryotic genome. First, the method has the same efficiency when evaluated using a collection of experimentally known operons of Bacillus subtilis. Second, operons among most if not all prokaryotes seem to have the same tendencies to keep short distances between their genes, the most frequent distances being the overlaps of four and one base pairs. The universality of this structural feature allows us to predict the organization of transcription units in all prokaryotes. Third, predicted operons contain a higher proportion of genes with related phylogenetic profiles and conservation of adjacency than predicted borders of transcription units. Supplementary information: Additional materials and graphs, are available at: http://www.cifn.unam.mx/moreno/pub/TUpredictions/ Contact: [email protected] Keywords: functional genomics; comparative genomics; operon prediction; operon structure. *To whom correspondence should be addressed.
Gabriel Moreno-Hagelsieb, Julio Collado-Vides
ISMB2
1998 Prediction of transcriptional regulatory sites in the complete genome sequence of Escherichia coli K-12
abstract
MOTIVATION: As one of the best-characterized free-living organisms, Escherichia coli and its recently completed genomic sequence offer a special opportunity to exploit systematically the variety of regulatory data available in the literature in order to make a comprehensive set of regulatory predictions in the whole genome. RESULTS: The complete genome sequence of E.coli was analyzed for the binding of transcriptional regulators upstream of coding sequences. The biological information contained in RegulonDB (Huerta, A.M. et al., Nucleic Acids Res.,26,55-60, 1998) for 56 different transcriptional proteins was the support to implement a stringent strategy combining string search and weight matrices. We estimate that our search included representatives of 15-25% of the total number of regulatory binding proteins in E.coli. This search was performed on the set of 4288 putative regulatory regions, each 450 bp long. Within the regions with predicted sites, 89% are regulated by one protein and 81% involve only one site. These numbers are reasonably consistent with the distribution of experimental regulatory sites. Regulatory sites are found in 603 regions corresponding to 16% of operon regions and 10% of intra-operonic regions. Additional evidence gives stronger support to some of these predictions, including the position of the site, biological consistency with the function of the downstream gene, as well as genetic evidence for the regulatory interaction. The predictions described here were incorporated into the map presented in the paper describing the complete E.coli genome (Blattner,F.R. et al., Science, 277, 1453-1461, 1997). AVAILABILITY: The complete set of predictions in GenBank format is available at the url: http://www. cifn.unam.mx/Computational_Biology/E.coli-predictions CONTACT: [email protected], [email protected]
Denis Thieffry, Heladia Salgado, Araceli M. Huerta, Julio Collado-Vides
Bioinform.4
1996 Syntactic recognition of regulatory regions in Escherichia coli
abstract
MOTIVATION: One of the most common methodologies to identify cis-regulatory sites in regulatory regions in the DNA is that of weight matrices, as testified by several articles in this issue. An alternative to strengthen the computational predictions in regulatory regions is to develop methods that incorporate more biological properties present in such DNA regions. The grammatical implementation presented in this paper provides a concrete example in this direction. RESULTS: On the basis of the analysis of an exhaustive collection of regulatory regions in Escherichia coli, a grammatical model for the regulatory regions of sigma 70 promoters has been developed. The terminal symbols of the grammar represent individual sites for the binding of activator and repressor proteins, and include the precise position of sites in relation to transcription initiation. Combining these symbols, the grammar generates a large number of different sentences, each of which can be searched for matching against a collection of regulatory regions by means of weight matrices specific for each set of sites for individual proteins. On the basis of this grammatical model, a Prolog syntactic recognizer is presented here. Specific subgrammars for ArgR, LexA and TyrR were implemented. When parsing a collection of 128 sigma 70 promoter regions, the syntactic recognizer produces a much lower number of false-positive sites than the standard search using weight matrices.
David A. Rosenblueth, Denis Thieffry, Araceli M. Huerta, Heladia Salgado, Julio Collado-Vides
Comput. Appl. Biosci.5
1991 The search for a grammatical theory of gene regulation is formally justified by showing the inadequacy of context-free grammars
abstract
No one questions the important practical contributions of computer sciences to molecular biology. It may well be that one day theoretical contributions also will become useful. One example of this type of interdisciplinary research is the attempt to construct a grammatical theory of the regulation of gene expression. In this paper, I demonstrate that context-free grammars are inadequate for the description of regulatory properties coded in the DNA. This result is supported by data available in the literature that show changes in the specificity of the recognition between regulatory proteins and their DNA targets. This result is an important limitation for the use of statistical approaches such as information theory as a source of inspiration for a theory of gene regulation. Additionally, such a demonstration gives formal justification to the search for more elaborate grammatical models in the study of gene regulation. Some basic proposals for such grammatical approach have been presented previously.
Julio Collado-Vides
Comput. Appl. Biosci.1